> ## Content Index
> Fetch the complete content index at: https://www.antoinebuteau.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Daily Digest - 2026-10-03
- URL: https://www.antoinebuteau.com/daily-digest-2026-10-03/
- Published: 2026-10-04T11:48:32.000Z
- Updated: 2026-10-04T11:48:32.000Z
- Description: How Uber scaled the Model Context Protocol across eight hundred internal servers and five thousand tools using a centralized gateway, automated IDL crawlers, and runtime discovery.
- Author: Antoine Buteau
- Tags: Digest

## In this digest

1. [Designing MCP Gateway Uber's MCP Management Platform](#digest-item-1)
2. [TypeSafe’s Jev: The Model That Does Not Talk — and Charges $0 for Output](#digest-item-2)
3. [The Art of Doing Financial Engineering](#digest-item-3)
4. [The Amazon Trap](#digest-item-4)
5. [Evals: how to know whether an AI system actually works](#digest-item-5)
6. [1/ Somewhere in your company today, an expert rejected a...](#digest-item-6)
7. [We’re going to need default hard budget caps on pretty much everything](#digest-item-7)
8. [Agents have compressed parts of a software estimate](#digest-item-8)
9. [Cash Machines](#digest-item-9)
10. [Agentic Assistants, AI Safety Discourse, Nuclear vs. Renewables](#digest-item-10)
11. [What’s 🔥 in AI/Infra/VC #518](#digest-item-11)
12. [You're Probably Sleeping On Computer Use](#digest-item-12)
13. [Claude-shaped science](#digest-item-13)
14. [The Slop Superhighway](#digest-item-14)
15. [How dots took over my Codex](#digest-item-15)

## Themes from yesterday

- **Decomposing intelligence into cheap decisions and selective deliberation:** High-volume routing and classification tasks are moving away from expensive generative frontier models toward zero-output-cost decision models like TypeSafe Jev and incremental discovery platforms like Uber's MCP Gateway. This approach lowers inference bills while keeping deterministic code in charge of execution.
- **The transition from reactive prompts to proactive, multi-channel agents:** Systems like OpenAI Dots, Meta Muse, and vision-based computer-use tools maintain persistent state across chat apps, mobile calls, and desktops. Rather than waiting for one-off prompts, they monitor services in the background, run tasks in parallel, and operate legacy software directly.
- **The economics of verification and enterprise data sovereignty:** As generating code and text becomes cheap, enterprise value centers on expert verification and structured evaluation. Companies are increasingly running their own eval pipelines and using open-weight models to keep decision logs and expert overrides from being captured by frontier model providers.
- **Infrastructure bottlenecks across capital, power, and machine protocols:** Deploying autonomous agents at scale is straining physical and operational infrastructure. This pressure has led to private debt financing for multi-gigawatt data centers, faster nuclear regulatory approvals, and an urgent need for default hard spending caps and native machine-to-machine communication protocols.

## 1\. **Designing MCP Gateway Uber's MCP Management Platform** — Uber Engineering

- Why read: How Uber scaled the Model Context Protocol across eight hundred internal servers and five thousand tools using a centralized gateway, automated IDL crawlers, and runtime discovery.
- Summary: To prevent fragmented tooling and bloated context windows across hundreds of engineering teams, Uber built MCP Gateway as a unified routing, discovery, and governance layer between AI agents and internal microservices. The platform runs a Cadence-powered AutoCrawler to continuously scan Protobuf and Thrift IDL registries, generating agent-ready tool definitions that service owners can review and enable without writing custom code. At runtime, the Proxy Gateway data plane delegates calls to Uber's Muttley service mesh, translating MCP requests into gRPC, HTTP, or TChannel calls while enforcing access rules and redacting PII. To control token usage, Uber introduced Omni MCP for intent-based incremental tool discovery alongside Response Projection, which trims API payloads by returning only the JSON fields requested by the model. For developer workflows, Uber added Code Mode to its internal CLI so coding agents can discover tools and pipe structured outputs directly to the local filesystem instead of consuming context window space.
- [Read more](https://twitter.com/UberEng/status/2106071967619322330/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

## 2\. **TypeSafe’s Jev: The Model That Does Not Talk — and Charges $0 for Output** — Decoding Discontinuity

- Why read: How zero-output-cost decision models split routine workflow routing from expensive frontier deliberation, shifting pricing leverage back to application developers.
- Summary: TypeSafe launched Jev, a fast decision model that evaluates structured questions against input context and returns predefined choices with calibrated probabilities instead of generating text. Because it drops the token decoding phase entirely, Jev charges forty-two cents per ten million input tokens and zero dollars for output, allowing applications to run high-volume logic branches at near-zero marginal cost. This structure splits machine intelligence into cheap, bounded decisions for routine forks and premium deliberation reserved for open-ended reasoning and synthesis. Although benchmarks show questions must be decomposed to maintain high accuracy and good calibration, platforms like Vercel and LangChain integrated the model within days to build lightweight front-end routers. Placing specialized decision models in front of frontier APIs keeps workflow control inside deterministic code while cutting inference bills.
- [Read more](https://www.decodingdiscontinuity.com/p/typesafe-jev-decision-models-ai-economics?ref=antoinebuteau.com)

## 3\. **The Art of Doing Financial Engineering** — netinterest.co

- Why read: How Wall Street is using structured debt to cover a six-trillion-dollar funding gap for AI data centers, borrowing techniques from the nineteenth-century railroad boom.
- Summary: Global AI infrastructure capital expenditures are projected to reach eight trillion dollars by 2030, creating a six-trillion-dollar financing gap that public corporate bond markets cannot absorb alone. Echoing the railroad boom of the late nineteenth century, financial institutions are assembling complex asset-backed debt and private credit structures to fund massive power and compute projects. Meta's five-gigawatt Hyperion data center in Louisiana shows how hyperscalers turn to specialized banking syndicates to keep multi-billion-dollar commitments from overloading corporate balance sheets. These bespoke vehicles allow physical construction to proceed quickly, but they also disperse credit exposure across private funds and rating agencies. The accumulation of off-balance-sheet leverage shows how heavily the physical AI buildout now relies on structured finance.
- [Read more](https://www.netinterest.co/p/the-art-of-doing-financial-engineering?ref=antoinebuteau.com)

## 4\. **The Amazon Trap** — Mike Vernal

- Why read: Why Amazon is blocking autonomous shopping agents to protect its eighty-billion-dollar ad business, and how a premium subscription could solve the innovator's dilemma.
- Summary: Amazon has started blocking horizontal consumer agents like Meta Muse and Instinct to stop automated price checks from bypassing sponsored product ads. Because sponsored search generates over eighty billion dollars in high-margin annual revenue, letting third-party agents checkout across competing storefronts threatens Amazon's primary discovery business. However, blocking agentic traffic creates an opening for rivals like Walmart and Shopify, which are actively partnering with assistant platforms to capture automated purchase volume. Former Benchmark partner Mike Vernal suggests Amazon introduce a Prime+ subscription tier that charges for agent access directly, replacing lost ad impressions with predictable subscription fees. The conflict highlights how consumer loyalty is shifting toward horizontal assistants that hold personal context rather than individual retail storefronts.
- [Read more](https://twitter.com/mvernal/status/2106396307476951105/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

## 5\. **Evals: how to know whether an AI system actually works** — Sergii Makarevych

- Why read: A practical framework for evaluating language model applications, from error analysis and code rubrics to statistical confidence intervals and agent reliability metrics.
- Summary: Moving generative AI into production requires replacing subjective vibe checks with structured evaluation suites built on curated test sets and verifiable gold answers. Teams should start with manual error analysis on raw model outputs to categorize specific failure modes, then deploy deterministic code checks before adding model-based judges. When using LLMs as judges, teams must measure grader agreement using Cohen's kappa, decompose prompts into binary questions that require quoted evidence, and correct for length and position biases. Evaluating multi-step agents requires tracking success rates across repeated runs rather than relying on a single passing run, while auditing tool trajectories to catch infinite loops. Release gates should also enforce statistical confidence intervals, since small sample sizes create detection floors that hide serious regressions.
- [Read more](https://twitter.com/sermakarevich/status/2106453816757354947/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

## 6\. **1/ Somewhere in your company today, an expert rejected a...** — Christian Catalini

- Why read: Why enterprise defensibility is shifting from execution to verification, and how companies can prevent frontier labs from absorbing their expert decision data.
- Summary: As generative models drive the marginal cost of producing code, text, and analysis toward zero, a company's economic moat shifts from execution to verification. Corporate hierarchies have historically operated as verification engines where experienced managers evaluate risk and catch subtle domain errors that automated workflows miss. Adopting third-party AI assistants without safeguards risks leaking a firm's most valuable asset: the logs of expert overrides and corrections that reveal how tacit judgment is applied. Frontier model providers can absorb these interaction traces to automate internal workflows, turning customer companies into commoditized wrappers reliant on rented models. To protect long-term defensibility, companies must run their own evaluation pipelines and use open-weight models to keep decision telemetry within corporate boundaries.
- [Read more](https://twitter.com/ccatalini/status/2106064313207398433/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

## 7\. **We’re going to need default hard budget caps on pretty much everything** — Simon Willison

- Why read: Why autonomous software agents make default hard spend limits necessary across cloud providers and API platforms.
- Summary: As autonomous coding agents and background automations take on broader tasks, pay-as-you-go APIs and cloud infrastructure need default hard spending limits instead of passive email warnings. Unattended agent loops can quickly run up thousands of dollars in compute, storage, or inference charges overnight when hitting unexpected bugs. Cloud providers are beginning to address this risk, with Amazon Web Services adding project pause limits and Google Cloud rolling out service-level spend caps. While engineering teams historically avoided hard limits to prevent service interruptions, unexpected billing spikes are far more damaging than brief downtime. Operators should audit their cloud accounts to set hard cutoff thresholds and configure agents to work only within capped environments.
- [Read more](https://simonwillison.net/2026/Oct/3/default-hard-budget-caps/?ref=antoinebuteau.com)

## 8\. **Agents have compressed parts of a software estimate** — Prashant Mittal

- Why read: How AI agents compress raw implementation while leaving review, security, and sign-offs unchanged, requiring revised estimates and P85 commitments.
- Summary: While AI coding agents accelerate raw implementation, they leave adjacent stages like project scoping, security audits, and user acceptance testing uncompressed. Because code review and verification quickly become the primary development bottleneck, realized productivity gains depend heavily on harness quality, stack familiarity, and output verifiability. Engineering managers should stop relying on pre-AI historical baselines and track agent-assisted hours separately to rebuild project benchmarks. To guard against delivery delays, teams should replace single-point estimates with eighty-fifth percentile probability targets. Maintaining a strict separation between objective technical estimates, executive goals, and contractual delivery terms prevents teams from passing unmanageable risk downstream.
- [Read more](https://twitter.com/prashant%5Fmit/status/2106380224367849518/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

## 9\. **Cash Machines** — Howard Lerman

- Why read: How small software teams use AI automation to reach tens of millions in revenue and free cash flow without large headcounts.
- Summary: Founders facing pressure to pursue top-line growth at all costs have a practical alternative in building lean, highly profitable software companies. Scaling enterprise software to fifty million dollars in annual recurring revenue once required hiring hundreds of sales reps, customer success managers, and support engineers. Today, automated workflows and AI agents allow teams of twenty to thirty-six people to manage high-volume sales funnels and customer onboarding. Operating at eighty percent gross margins, a fifty-million-dollar ARR business can generate more than thirty million dollars in annual free cash flow. This model gives founders substantial personal liquidity and independence without needing speculative multi-billion-dollar public exits.
- [Read more](https://twitter.com/howard/status/2106399893711368199/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

## 10\. **Agentic Assistants, AI Safety Discourse, Nuclear vs. Renewables** — Contrary Research

- Why read: Key developments across consumer assistants, safety regulations, and nuclear power approvals for AI data centers.
- Summary: The consumer assistant market expanded this week as OpenAI launched Dots, startup Instinct closed a one-billion-dollar round at a ten-billion-dollar valuation, and Meta Muse reached three million weekly active users. Deployment problems emerged as well, with Meta issuing emergency privacy patches after Muse shared private user addresses on public marketplace listings. Regulatory oversight grew stricter as major AI labs signed a voluntary White House agreement on self-improving models while the Federal Trade Commission opened consumer protection inquiries into leading providers. To supply power for upcoming data centers, the Nuclear Regulatory Commission accelerated approval for commercial small modular reactors ahead of schedule. In hardware, AMD completed an eight-billion-dollar acquisition of World Labs, establishing spatial intelligence and physical foundation models as key semiconductor battlegrounds.
- [Read more](https://substack.com/app-link/post?post%5Fid=218190220&publication%5Fid=1511474&ref=antoinebuteau.com)

## 11\. **What’s 🔥 in AI/Infra/VC #518** — Ed Sim from What's Hot 🔥 in AI/Infra/VC

- Why read: Anthropic's hundred-million-dollar enterprise training push, heavy investment in physical world models, and multi-model routing inside OpenAI Codex.
- Summary: Anthropic committed one hundred million dollars to train ten thousand forward deployed engineers across consulting firms and financial institutions, aiming to close the enterprise deployment gap. Placing specialized deployment engineers directly into client workflows helps bridge custom data pipelines, but it also increases long-term platform lock-in. Concurrently, physical AI and spatial world models saw heavy capital investment, marked by AMD's eight-billion-dollar buyout of World Labs and large funding rounds for General Intuition and Field AI. Infrastructure providers are also shifting commercial terms, with OpenAI enabling enterprise customers to run competing open-source models inside Codex against existing contract commitments. Across software and robotics, investors are focusing on platforms that convert raw model capabilities into dependable operational execution.
- [Read more](https://substack.com/app-link/post?post%5Fid=216461425&publication%5Fid=13300&ref=antoinebuteau.com)

## 12\. **You're Probably Sleeping On Computer Use** — Laura Entis

- Why read: How operators use OS-level automation and computer-use models to handle repetitive work across slides, file audits, and legacy software.
- Summary: Operating-system-level computer use has shifted from an experimental demo into a dependable tool for knowledge workers handling repetitive desktop chores. By letting models click, scroll, and type directly across standard desktop applications, workers can bypass missing APIs to automate tasks in legacy web portals. Teams are using computer-use agents to check hundreds of proof links across PDF documents and format complete presentation decks from design templates. Unlike brittle scripts, vision-driven agents interact dynamically with software like Google Slides, Blender, and web support chats without requiring custom integration code. Product operators can evaluate recurring manual bottlenecks to determine which tasks can be handed off to background desktop agents.
- [Read more](https://every.to/context-window/you-re-probably-sleeping-on-computer-use?ref=antoinebuteau.com)

## 13\. **Claude-shaped science** — Anthropic

- Why read: How a Harvard physicist used Claude to calculate unsolved particle physics integrals and resolve decades-old equations in genetics and ecology.
- Summary: Harvard physics professor Matthew Schwartz built BootLoops, an open-source harness designed to apply language models to quantitative scientific problems that fit their specific capabilities. Rather than using models for open-ended conceptual brainstorming, the system focuses on computational tasks like Feynman integrals that can be checked with high-precision numerical verification. Drawing on its cross-disciplinary training, Claude recognized that mathematical techniques developed for particle scattering amplitudes could solve long-standing differential equations in evolutionary biology and ecology. Working with domain specialists who steered the model away from mathematically valid but unhelpful results, the project solved twenty-year-old ecological biodiversity equations. The workflow shows that AI delivers the greatest scientific value when paired with expert direction and deterministic verification harnesses rather than running autonomously.
- [Read more](https://www.anthropic.com/research/claude-shaped-science?ref=antoinebuteau.com)

## 14\. **The Slop Superhighway** — Rhys

- Why read: Why routing automated agent traffic through human communication channels breaks workflows, and what dedicated machine protocols need to solve.
- Summary: As developers deploy personal agents to resolve software bugs, book reservations, and file support tickets, human communication channels are becoming overloaded with automated inbound requests. Forcing maintainers and support staff to manually review machine-generated pull requests and tickets slows response times and causes burnout across open-source ecosystems. The root problem is the lack of dedicated machine-to-machine communication protocols, which forces human users to act as intermediaries shuttling data between opposing agents. Establishing a machine-native communication protocol would allow software services to ingest structured edge-case reproductions directly into automated test loops. Platform architects need to build separate agent-facing APIs that isolate high-volume automated traffic from spaces reserved for human collaboration.
- [Read more](https://twitter.com/RhysSullivan/status/2106488797638861123/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

## 15\. **How dots took over my Codex** — dominik kundel

- Why read: How persistent personal agents maintain context across Slack, mobile, and cloud environments to coordinate engineering work in the background.
- Summary: OpenAI's Dots marks a transition from reactive, prompt-driven coding sessions toward always-on personal agents that coordinate engineering workflows across multiple platforms. By maintaining persistent context across Slack, mobile voice calls, and desktop chats, a dot lets developers switch devices without re-explaining project state. The agent manages multiple background tasks in parallel, routing front-end edits to local machines while dispatching heavy back-end builds to cloud environments. Its event system monitors pull request deployments and system outages in connected channels, alerting the developer only when manual intervention is needed. Delegating coordination to autonomous personal agents reduces context switching and saves developer time.
- [Read more](https://twitter.com/dkundel/status/2106141525306397099/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)