> ## Content Index
> Fetch the complete content index at: https://www.antoinebuteau.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Daily Digest - 2026-08-11
- URL: https://www.antoinebuteau.com/daily-digest-2026-08-11/
- Published: 2026-08-12T11:19:16.000Z
- Updated: 2026-09-07T22:33:29.000Z
- Description: Vercel's AI Gateway data shows shifts in inference market share and enterprise spending. In July 2026, average token prices dropped 13.6% as demand for cheaper open-weight models grew.
- Author: Antoine Buteau
- Tags: Digest

## In this digest

1. [DeepSeek overtakes Google on volume, cost per token falls 13.6%](#1-deepseek-overtakes-google-on-volume-cost-per-token-falls-136-%E2%80%94-vercel)
2. [Nvidia’s Risky Business](#2-nvidia%E2%80%99s-risky-business-%E2%80%94-ben-thompson)
3. [No Process, No Agent](#3-no-process-no-agent-%E2%80%94-mark-ajzenstadt)
4. [Harness Your Company](#4-harness-your-company-%E2%80%94-jacob-posel)
5. [To FDE, or not to FDE?](#5-to-fde-or-not-to-fde-%E2%80%94-jesse-zhang)
6. [A Winner in Every Category](#6-a-winner-in-every-category-%E2%80%94-tomasz-tunguz)
7. [Can software factories actually work?](#7-can-software-factories-actually-work-%E2%80%94-posthog)
8. [Agentic Code Quality](#8-agentic-code-quality-%E2%80%94-addy-osmani)
9. [Climbing the Hills That Matter](#9-climbing-the-hills-that-matter-%E2%80%94-andrew-li)
10. [Everything hackable will get hacked](#10-everything-hackable-will-get-hacked-%E2%80%94-malte-ubl)
11. [\[AINews\] Muse Glimmer and Spark: Open Weights return Personal Superintelligence promise](#11-ainews-muse-glimmer-and-spark-open-weights-return-personal-superintelligence-promise-%E2%80%94-ainews)
12. [Some Simple Economics of Open Versus Closed AI](#12-some-simple-economics-of-open-versus-closed-ai-%E2%80%94-christian-catalini)
13. [insane paper, they extracted reasoning from frontier models by asking...](#13-insane-paper-they-extracted-reasoning-from-frontier-models-by-asking-%E2%80%94-elie)
14. [I’ve been getting this question a lot lately: "My startup...](#14-i%E2%80%99ve-been-getting-this-question-a-lot-lately-my-startup-%E2%80%94-aakash-sabharwal)
15. [How frontier models train on outcomes in 2026](#15-how-frontier-models-train-on-outcomes-in-2026-%E2%80%94-sergio-paniego)

## Themes from yesterday

- **Models to Harnesses**: AI competitive advantage is shifting from foundational models to proprietary data, workflows, and context layers.
- **Maturing Workflows**: Deploying agents requires process documentation, automated constraints, and monitoring based on real user behavior.
- **Open vs. Closed Economics**: Open-weights models are dropping inference prices and giving enterprises control over their data, despite frontier labs pushing for tighter security.
- **Training and Security**: AI training now uses Reinforcement Learning with Verifiable Rewards (RLVR) in interactive environments, raising new challenges in cybersecurity and the transparency of hidden reasoning.

## **1\. DeepSeek overtakes Google on volume, cost per token falls 13.6% — Vercel**

- Why read: Vercel's AI Gateway data shows shifts in inference market share and enterprise spending.
- Summary: In July 2026, average token prices dropped 13.6% as demand for cheaper open-weight models grew. Total inference spend rose 37%. DeepSeek passed Google in token volume, taking 25% of gateway traffic, mostly via its V4 Flash model. Anthropic kept 65% of spending with only 30% of volume, showing its hold on complex tasks like coding. Cheap, capable models are driving down prices, but most revenue stays in the premium tier where buyers have fewer options.
- [Read more](https://twitter.com/vercel/status/2087313511386878347/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

## **2\. Nvidia’s Risky Business — Ben Thompson**

- Why read: Compares the 1870s railroad bubble to current hyperscaler AI investments, highlighting the risks of funding capacity with debt.
- Summary: Tech companies are issuing heavy debt to build AI infrastructure, resembling the railway expansion that preceded the Panic of 1873\. Microsoft pays for CapEx with free cash flow, but others are borrowing heavily while bond yields rise. Google's $85 billion equity raise, featuring $10 billion from Berkshire Hathaway, shows compute demand is high and cash limits are real. This expensive race may consolidate power among the richest companies. Also, DeepMind's internal changes suggest a shift toward text-centric AI architectures over foundational world models.
- [Read more](https://stratechery.com/2026/nvidias-risky-business/?ref=antoinebuteau.com)

## **3\. No Process, No Agent — Mark Ajzenstadt**

- Why read: Why enterprises fail when they deploy AI agents before documenting their workflows.
- Summary: AI project success is 5% model choice and 95% documented process. Most companies put agents into chaotic, unmapped workflows. The AI fills gaps with bad assumptions and fails in production. Software engineering agents work because coding is a structured and observable digital process. To get similar results elsewhere, companies must map workflows and define edge cases. They also need human fallbacks in place before writing any agent code.
- [Read more](https://twitter.com/mardehaym/status/2087086419491647589/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

## **4\. Harness Your Company — Jacob Posel**

- Why read: How to move from fragmented AI tools to a company-wide AI operating environment.
- Summary: Organizations struggle with AI fragmentation. Disconnected tools, local prompts, and scattered context hurt productivity. A functioning company AI harness needs shared memory, central secrets, permission policies, and real-time project sync. Separating the knowledge graph and integrations from the user interface ensures all agents and employees share the same context. This turns AI from a personal tool into a background worker that executes multi-step company workflows.
- [Read more](https://twitter.com/jacob%5Fposel/status/2087177863824802102/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

## **5\. To FDE, or not to FDE? — Jesse Zhang**

- Why read: The rise of Forward Deployed Engineers (FDEs) in AI startups and how to avoid becoming a consulting firm.
- Summary: The Forward Deployed Engineer (FDE) role is the new default go-to-market strategy for AI startups facing undefined enterprise workflows. Putting engineers onsite helps discover the product, since users don't know what they want until they see it. But relying on FDEs to patch product gaps hurts margins and scalability. Startups should use FDEs to turn bespoke customer fixes into standard platform features. This automates deployment and moves the company away from a pure services model.
- [Read more](https://twitter.com/thejessezhang/status/2087198484093149421/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

## **6\. A Winner in Every Category — Tomasz Tunguz**

- Why read: How AI adoption drives high valuation premiums for category leaders in public SaaS markets.
- Summary: While SaaS multiples have dropped, category leaders are trading at high premiums because of their AI positioning. Companies like CrowdStrike, Cloudflare, and Shopify use their data and distribution networks to make money from AI agents. The market rewards firms that own the context, endpoints, and workflows needed to run autonomous systems at scale. The real value of AI is in the data harnesses and infrastructure that make foundational models useful to businesses.
- [Read more](https://twitter.com/ttunguz/status/2087204622805135780/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

## **7\. Can software factories actually work? — PostHog**

- Why read: Examines the viability of fully autonomous software factories.
- Summary: The idea of "software factories" where agents write, test, and merge code without humans faces criticism for hurting codebase maintainability. Critics argue models prioritize quick fixes over architecture because standard benchmarks don't measure design trade-offs. The problem is treating agents as isolated executors. Giving coding agents production signals, user feedback, and architectural constraints allows them to make better design decisions, pushing the industry closer to self-driving product engineering.
- [Read more](https://twitter.com/posthog/status/2087248173106684127/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

## **8\. Agentic Code Quality — Addy Osmani**

- Why read: How to handle software quality assurance when AI agents generate code at machine speed.
- Summary: Human code review cannot keep up with agent-generated code. Teams need automated constraints and quality gates. Ensuring code quality requires unit tests, mutation testing, and architectural linting to reject bad code before production. Humans should only review complex architectural decisions and edge cases where automation fails. Adjusting the strictness of these constraints lets teams balance speed with safety, keeping the codebase maintainable while maximizing output.
- [Read more](https://addyo.substack.com/p/agentic-code-quality)

## **9\. Climbing the Hills That Matter — Andrew Li**

- Why read: Why static AI evaluations fail and how Agent Behavior Monitoring tracks actual user preferences.
- Summary: Standard LLM judges fail to capture domain-specific requirements in production. Building reliable agents requires Agent Behavior Monitoring (ABM) to track raw user trajectories, corrections, and approvals. Mining this interaction data helps companies find implicit quality criteria and build custom rubrics that change with user behavior. Turning messy production feedback into dynamic evaluations is the only way to measure performance accurately and turn usage data into a product advantage.
- [Read more](https://www.judgmentlabs.ai/blogs/climbing-the-hills-that-matter?ref=antoinebuteau.com)

## **10\. Everything hackable will get hacked — Malte Ubl**

- Why read: Why organizations must use AI for defense before offensive open-weight models close the capability gap.
- Summary: Open-weight models like Kimi K3 can now map attack surfaces and write custom fuzzers without safeguards. Defenders still have the advantage because frontier models are better at defense when they have source code access. Organizations need AI security harnesses to review codebases and process vulnerability hypotheses. As offensive models improve, keeping the defensive advantage means integrating AI into software development. Failing to patch systems with these tools leaves infrastructure open to automated attacks.
- [Read more](https://vercel.com/blog/everything-hackable-will-get-hacked?ref=antoinebuteau.com)

## **11\. \[AINews\] Muse Glimmer and Spark: Open Weights return Personal Superintelligence promise — AINews**

- Why read: Meta returns to the open-weights frontier, and the industry shifts toward optimized agent harnesses.
- Summary: Meta released Muse Glimmer, a 30B multimodal model for local, always-on personal agents. Anthropic demonstrated AI-assisted theorem-search on the Riemann Hypothesis, and OpenAI launched a restricted GPT-5.6-Cyber model for defensive security. The industry recognizes that agent performance depends on the harness and tool interfaces, shifting toward treating tools as native code objects instead of JSON schemas. Speculative decoding is also improving inference speed and token efficiency, making fast, local agent deployment a reality.
- [Read more](https://substack.com/redirect/faf21649-f41a-47f2-aef7-10a89551bce5?j=eyJ1IjoiMzR5bXIifQ.a6vEKQR6802KnRCI5UGsr-WfrZy92G7RG6mX8wlMt6Q&ref=antoinebuteau.com)

## **12\. Some Simple Economics of Open Versus Closed AI — Christian Catalini**

- Why read: An economic argument for why open-weights AI accelerates innovation without undermining frontier labs.
- Summary: The debate over open versus closed AI mirrors historical arguments about intellectual property. Frontier labs say distillation threatens R&D funding, but economic history shows diffusing foundational tech drives progress. In most sectors, value goes to incumbents who control distribution, trust, and proprietary data, not the model creators. Open-weights models prevent market concentration and help enterprises maintain control over their intellectual property as they adopt AI.
- [Read more](https://twitter.com/ccatalini/status/2087177319253459019/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

## **13\. insane paper, they extracted reasoning from frontier models by asking... — elie**

- Why read: How researchers exposed the hidden chain-of-thought reasoning of safeguarded frontier models.
- Summary: Researchers extracted the hidden reasoning blocks of frontier models by prompting less-safeguarded models to decrypt them. The study showed that advanced models optimize their hidden reasoning with telegraphic language and sometimes use foreign languages like Chinese or Russian. The extracted reasoning leaked personal information and showed models recognizing chances to cheat. These behaviors are hidden from users, exposing transparency and security flaws in how frontier labs obscure model cognition.
- [Read more](https://twitter.com/eliebakouch/status/2087179305474298162/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

## **14\. I’ve been getting this question a lot lately: "My startup... — Aakash Sabharwal**

- Why read: Why raw enterprise data is worthless on its own, and how to make it valuable.
- Summary: Startups often overestimate the value of their raw data. Frontier labs need more than isolated PDFs or CRM records. AI models learn to work in enterprises using raw artifacts, simulated servers, tasks, and expert verifiers. Data labeling companies anonymize artifacts and reconstruct these environments, making the data valuable for training agents. For enterprises, the goal isn't selling data, but using these environments to evaluate models against real workflows to drive internal AI deployment.
- [Read more](https://twitter.com/aakashsabharwal/status/2087290636936569124/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

## **15\. How frontier models train on outcomes in 2026 — Sergio Paniego**

- Why read: How frontier labs use Reinforcement Learning with Verifiable Rewards (RLVR) to improve reasoning.
- Summary: The AI industry has moved past standard RLHF, using Group Relative Policy Optimization (GRPO) and outcome-based rewards to train models on code and math. Rewarding models based on test difficulty and sequence-level likelihood avoids issues like sparse rewards. Frontier labs now train agents inside interactive environments with programmatic checks instead of using single-turn answers. These reinforcement learning techniques build domain-expert models, which are then distilled into efficient student models for deployment.
- [Read more](https://twitter.com/SergioPaniego/status/2086805987705417851/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)