> ## Content Index
> Fetch the complete content index at: https://www.antoinebuteau.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Daily Digest - 2026-10-09
- URL: https://www.antoinebuteau.com/daily-digest-2026-10-09/
- Published: 2026-10-10T11:32:47.000Z
- Updated: 2026-10-10T11:32:47.000Z
- Description: Hex cut chart-editing errors from 21% to 3% by combining an 1,800-case evaluation suite with automated hill-climbing. Hex built its Quick Edits feature by having coding agents test changes against 1,800 test fixtures.
- Author: Antoine Buteau
- Tags: Digest

## In this digest

1. [Hill-climbing with evals: how we cut AI errors 7x](#digest-item-1)
2. [What's old is new again: Sidecars for Modal Sandboxes](#digest-item-2)
3. [Building for a workload that is mostly idle: inside DigitalOcean Managed Agents](#digest-item-3)
4. [How we built a personal data scientist for every Notion employee](#digest-item-4)
5. [Scaling AI in Legal: Building Uber's Redlining Agent](#digest-item-5)
6. [Why I tried to kill token billing (and why we kept it)](#digest-item-6)
7. [The Agent P&L Framework](#digest-item-7)
8. [Clouded Judgement 10.9.26 - Cache Is the New Egress](#digest-item-8)
9. [Skill Bundles: Dissecting the New Application Layer](#digest-item-9)
10. [Against the Personal Agents Theory of Everything](#digest-item-10)
11. [Assorted things I've been mulling over regarding agents in the...](#digest-item-11)
12. [Why AlphaFold Didn't Solve Protein Folding — Pushmeet Kohli, Google DeepMind & Sal Candido, Biohub](#digest-item-12)
13. [The Future(s) of Compute](#digest-item-13)
14. [Agentic Primitives 101](#digest-item-14)
15. [Teach the AI to Disagree With You](#digest-item-15)

## Themes from yesterday

- Deterministic harnesses and validators drive reliability: Programmatic validators, strict output schemas, and runtime sidecars cut production bugs far more reliably than endless prompt tweaking.
- Hardware-assisted isolation and split planes underpin modern agent infrastructure: Running bursty, stateful agent workloads economically requires Firecracker microVMs, on-host container bridges, and dedicated streaming data planes.
- Enterprise monetization is moving past commodity tokens toward output metrics and cache economics: Software vendors are shifting toward unified credit systems and agent P&L accounting, while prompt cache read rates now drive day-to-day operating expenses and provider lock-in.
- Shared workflow factories are superseding fragmented personal assistants: Teams are replacing individual prompt tinkering with centralized workflow factories and typed skill bundles that turn internal company playbooks into reusable systems.

## 1\. **Hill-climbing with evals: how we cut AI errors 7x** — David Wilson

- Why read: Hex cut chart-editing errors from 21% to 3% by combining an 1,800-case evaluation suite with automated hill-climbing.
- Summary: Hex built its Quick Edits feature by having coding agents test changes against 1,800 test fixtures. Instead of relying on prompt engineering, the team got more than half of its accuracy gains from an output validator that fixes formatting, rounds numbers, and enforces domain rules. They prioritized preventing wrong edits over avoiding handoffs to their larger agent, since bad edits quickly erode user trust. To avoid overfitting, Hex kept 40 percent of the test cases as a holdout set and ran each test ten times to filter out noise. Teams can apply this playbook by pairing cheaper models with automated eval loops and deterministic validators instead of constantly tweaking prompts.
- [Read more](https://twitter.com/daviddbwilson/status/2108226292630077883/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

## 2\. **What's old is new again: Sidecars for Modal Sandboxes** — Adam Azzam

- Why read: Modal added sidecar containers to its sandboxes, creating a secure boundary on the same host without the latency of remote control planes.
- Summary: Sandboxed coding agents often run untrusted code alongside credentials and test harnesses, which creates risks of prompt injection and credential leaks. Moving control logic to external services protects secrets, but calling across remote networks on every tool execution adds substantial latency. Modal addresses this by running trusted sidecars next to the execution sandbox on the same host, isolated by gVisor or virtual machine boundaries. Inter-container traffic flows across a local network bridge, running three times faster than remote sandboxes while maintaining separate outbound network rules. Platform teams can use this setup to inject API keys, proxy traffic, and audit agent activity without exposing credentials to untrusted code.
- [Read more](https://twitter.com/AAAzzam/status/2108278467796078614/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

## 3\. **Building for a workload that is mostly idle: inside DigitalOcean Managed Agents** — Salman Paracha

- Why read: DigitalOcean explains how hardware microVMs and decoupled control and data planes make bursty, mostly idle agent workloads affordable to run.
- Summary: Software engineering agents spend long stretches waiting for user inputs and model turns, yet they must keep language servers and active processes loaded in memory. DigitalOcean handles this pattern with Firecracker microVMs that snapshot disk and memory, letting the platform pause instances quickly without dropping state. When restoring a session, the system streams memory pages on demand through userfaultfd, waking paused instances in under 2.5 seconds to match warm instance speeds. The infrastructure separates low-volume lifecycle actions from high-volume event streaming, routing logs through session-keyed Kafka topics and resumable server-sent events. An Action Gateway checks permissions and injects API credentials from the outside so third-party tokens never enter the execution sandbox.
- [Read more](https://twitter.com/salman%5Fparacha/status/2108601051633061965/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

## 4\. **How we built a personal data scientist for every Notion employee** — David Thomas

- Why read: Notion cut ad-hoc data requests by 80% with an internal agent that pairs natural language Snowflake queries with workspace documentation.
- Summary: Notion rolled out Data Scout so employees in sales, product, and customer experience could query Snowflake directly. Generating valid SQL proved relatively simple, but getting business logic right required feeding organizational context into the retrieval pipeline. To stop the agent from misinterpreting metrics or missing changes in tracking history, the team indexed official data schemas, metric definitions, and company procedures stored in Notion pages. Every query runs under the requesting employee's Snowflake credentials, preserving permissions and row-level security without manual permission syncing. The setup has allowed internal teams to build scheduled reporting agents and domain helpers while freeing data engineers for strategic initiatives.
- [Read more](https://www.notion.com/blog/how-we-built-a-personal-data-scientist-for-every-notion-employee?ref=antoinebuteau.com)

## 5\. **Scaling AI in Legal: Building Uber's Redlining Agent** — Uber

- Why read: Uber explains how four iterations of a Word-based contract redlining agent cut negotiation review times by 20% while reaching 91% accuracy.
- Summary: Uber built its Legal Redlining Agent directly inside Microsoft Word so attorneys could review contracts without switching tools. The first version used standard retrieval-augmented generation against legal playbooks, but basic vector search led to inconsistent advice and the wrong negotiation tone. The team fixed this by logging every lawyer edit and applying exponential decay weighting to pull in relevant recent precedents at runtime. Uber also put in-house attorneys in charge of prompt design and added a single-pass reflection check for tone to avoid multi-turn latency. Replacing basic text generation with structured drafting workflows helped the system consistently produce compromise clauses that follow company policy.
- [Read more](https://www.uber.com/in/en/blog/building-ubers-redlining-agent/?ref=antoinebuteau.com)

## 6\. **Why I tried to kill token billing (and why we kept it)** — Scott Woody

- Why read: Metronome co-founder Scott Woody explains why billing customers for tokens damages software margins, arguing instead for unified credits that lead toward output-based pricing.
- Summary: Invoicing customers directly for token usage forces software companies to justify markups on commodity inference. As underlying model costs fall, buyers push back on margins or bypass the software to call APIs directly. To protect margins, vendors are turning to unified credit systems that let users spend from a shared balance across different product features while hiding model routing behind the scenes. Woody notes that pure outcome-based pricing is difficult for most products because buyers and sellers will argue over whether human effort or automated tools drove the result. He recommends billing against countable units of output, like completed data enrichments or resolved support tickets, which turn automated work into concrete business deliverables.
- [Read more](https://stripe.com/blog/where-pricing-is-headed?ref=antoinebuteau.com)

## 7\. **The Agent P&L Framework** — Good Better Best by PricingSaaS

- Why read: Databricks pricing specialist Manu Mehra presents a framework to calculate net value, cost per verified outcome, and return on investment for enterprise agents.
- Summary: Software vendors face a choice between commoditized token billing and difficult outcome-based contracts that require complex attribution. Manu Mehra suggests creating an explicit profit and loss statement for each agent workflow. On the expense side, the model factors in infrastructure, monitoring, evaluations, and compliance rather than tracking only token costs. On the return side, it tallies saved labor hours, faster turnaround times, and prevented operational errors. Tracking rework rates is essential, since workflows that take twenty prompts to correct will lose users no matter how cheap the model inference is. This accounting helps product teams demonstrate clear financial return and price their tools sustainably without getting bogged down in attribution debates.
- [Read more](https://read.readwise.io/read/01m4g9ers7qbjw86w208syy2wt?ref=antoinebuteau.com)

## 8\. **Clouded Judgement 10.9.26 - Cache Is the New Egress** — Clouded Judgement by Jamin Ball

- Why read: Altimeter partner Jamin Ball compares prompt cache pricing to cloud egress fees, showing how cache read costs determine real agent expenses and lock teams into providers.
- Summary: While headline token prices get most of the attention, cache read rates determine actual operating costs for multi-turn agents. Agent loops repeatedly pass tool definitions and conversation history, meaning cached tokens can account for up to 85 percent of all input volume. Even when competing frontier models post identical sticker prices, large gaps in cache read rates create wide swings in running costs across real sessions. Switching providers in the middle of a workflow forces teams to rebuild context caches, creating financial friction much like cloud egress fees. Teams building long-running background agents need to evaluate cache pricing models rather than baseline token costs when choosing providers.
- [Read more](https://substack.com/app-link/post?post%5Fid=219484083&publication%5Fid=56878&ref=antoinebuteau.com)

## 9\. **Skill Bundles: Dissecting the New Application Layer** — Josh Rosen

- Why read: An architectural breakdown of skill bundles from OpenAI, Microsoft, and Stripe, exploring how standardizing domain workflows turns agent instructions into a software layer.
- Summary: Tech companies are packaging domain workflows into collections of skills that guide agents through complex professional tasks. These bundles center around specific capabilities, such as financial modeling or systems diagnostics, rather than traditional software menus. More advanced setups use routing skills to dispatch requests to specialized agents and rely on typed schemas, such as JSON manifests, to validate outputs programmatically. Even so, because skills remain natural-language prompts processed by probabilistic models, teams still need rigid execution harnesses to guarantee compliance. Sharing these bundles also publishes internal operational workflows, turning company processes and playbooks into readable public specifications.
- [Read more](https://twitter.com/JoshARosen/status/2108276264763048104/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

## 10\. **Against the Personal Agents Theory of Everything** — Nathan Baschez

- Why read: Notion product designer Nathan Baschez argues against relying primarily on individual AI assistants, showing why centralized workflow factories produce better business results.
- Summary: Giving every employee an individual AI assistant leads workers to recreate prompts and tool setups for the same routine work. This fragmented setup produces uneven quality and keeps companies from compounding operational improvements. Baschez argues organizations should build shared software pipelines for recurring tasks like client onboarding, month-end financial closes, and slide deck creation. Because dedicated domain experts maintain these centralized pipelines, they receive the engineering work needed for higher accuracy, lower failure rates, and stronger security. Employees can then trigger these pipelines directly or have their personal assistants dispatch work to them without managing the underlying machinery.
- [Read more](https://www.gettheleverage.com/p/against-the-personal-agents-theory?ref=antoinebuteau.com)

## 11\. **Assorted things I've been mulling over regarding agents in the...** — Patrick Collison

- Why read: Stripe CEO Patrick Collison examines how consumer buying agents could erode traditional price discrimination and reward genuine product quality.
- Summary: Businesses have long leaned on consumer inattention and limited research time to profit from subscription inertia, upsells, and price discrimination. As personal agents take over shopping, buyers can run exhaustive price and feature comparisons, undermining retail placement fees and cross-subsidies. Collison suggests this dynamic will act as an indirect subsidy for quality, reducing the advantage of legacy distribution channels while rewarding better engineering. Verified data and objective durability metrics will become crucial signals for algorithmic shoppers. Founders should expect that building measurably better products will offer stronger protection as software agents replace human browsing.
- [Read more](https://twitter.com/patrickc/status/2108706907859116280/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

## 12\. **Why AlphaFold Didn't Solve Protein Folding — Pushmeet Kohli, Google DeepMind & Sal Candido, Biohub** — Latent.Space

- Why read: Research leaders from DeepMind and Biohub explain why predicting static protein structures is only an opening step toward modeling dynamic living cells.
- Summary: AlphaFold solved static protein structure prediction, but modeling living biology requires much more work. Pushmeet Kohli and Sal Candido explain that static snapshots miss conformational shifts, intrinsically disordered proteins, and molecular interactions inside living cells. Pouring more compute into existing public databases will not produce breakthrough biological intelligence without new experimental measurements. The researchers note that training protein models on noisy metagenomic data can still improve design work, highlighting the need to find true scaling laws for biological datasets. Building predictive models of entire cells will require wet labs and machine learning researchers to collaborate on generating custom training data.
- [Read more](https://substack.com/app-link/post?post%5Fid=219666029&publication%5Fid=1084089&ref=antoinebuteau.com)

## 13\. **The Future(s) of Compute** — Aishwarya Mahesh

- Why read: An analysis applying commodity market mechanics to GPU futures, showing why hardware differences and erratic price benchmarks make compute hard to hedge.
- Summary: Major exchanges like CME and ICE are designing GPU futures to turn raw compute into a tradeable commodity. Applying lessons from agriculture and energy markets shows why standardizing compute is difficult. A GPU hour is not an interchangeable asset, since interconnect bandwidth, cluster topology, regional latency, and contract length heavily affect real-world training performance. Different pricing indices produce conflicting benchmarks, with overlapping spot indices showing a correlation of just 0.17 for the same hardware. Until market participants agree on delivery grades and price adjustments, financial derivatives will not reliably hedge enterprise infrastructure costs.
- [Read more](https://aishwaryamahesh.substack.com/p/the-futures-of-compute)

## 14\. **Agentic Primitives 101** — Charlie Guo

- Why read: A practical guide to the core infrastructure primitives needed to run multi-hour, stateful agent sessions without losing context.
- Summary: Early agent deployments struggled because typical cloud setups expect quick, stateless requests rather than multi-hour workflows. Charlie Guo outlines the core building blocks required for persistent agent runs. Environments need isolated sandboxes with shell and filesystem access so models can run tests and inspect intermediate results safely. To prevent context overflow and degraded reasoning, teams should use structured scratchpads, periodic state compaction, and external memory stores. Adding human approval gates and automatic checkpoint recovery helps ensure temporary errors do not wipe out hours of progress.
- [Read more](https://www.ignorance.ai/p/agentic-primitives-101?ref=antoinebuteau.com)

## 15\. **Teach the AI to Disagree With You** — Austin Johnsen from Artificial Diligence

- Why read: A tactical guide to fighting model sycophancy through adversarial prompts and multi-model debate setups.
- Summary: Most foundation models are tuned to be polite and accommodating, which often leads them to validate shaky business plans with corporate buzzwords. Founders who rely on this agreeable feedback risk missing critical strategic and operational flaws. Austin Johnsen outlines how to configure models as skeptical critics that challenge core assumptions and push back on weak logic. Running multi-agent reviews where models from different providers critique each other's reasoning surfaces blind spots that a single prompt will miss. While an adversarial review adds friction during the planning stage, stress-testing proposals early prevents costly mistakes before pitches and launches.
- [Read more](https://substack.com/app-link/post?post%5Fid=219523785&publication%5Fid=1180691&ref=antoinebuteau.com)