> ## Content Index
> Fetch the complete content index at: https://www.antoinebuteau.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Daily Digest - 2026-08-20
- URL: https://www.antoinebuteau.com/daily-digest-2026-08-20/
- Published: 2026-08-21T10:39:23.000Z
- Updated: 2026-08-21T10:39:23.000Z
- Description: Explains why adding more parameters isn't enough anymore; we need more compute and better post-training data for reasoning. Tracking parameter counts is getting outdated.
- Author: Antoine Buteau
- Tags: Digest

**1\. \[AINews\] Death of Params: Z.ai CEO Jie Tang on GLM 5.3 and the new Post-training Scaling Law — Substack**

- Why read: Explains why adding more parameters isn't enough anymore; we need more compute and better post-training data for reasoning.
- Summary: Tracking parameter counts is getting outdated. Inference compute and post-training data quality matter more now. The GLM-5.3 model improved significantly just by using reinforcement learning in complex, real-world engineering setups. It tackles 20+ step reasoning problems on its own, without humans breaking the tasks down. The entire system of environments, judging, and verification is becoming fully synthetic so the model can improve itself. Engineering teams need to focus on data mixtures, synthetic evaluators, and post-training infrastructure instead of just scaling up parameters.
- [Read more](https://substack.com/app-link/post?publication%5Fid=1084089&%3Bpost%5Fid=211952724&%3Butm%5Fsource=post-email-title&%3Butm%5Fcampaign=email-post-title&%3BisFreemail=true&%3Br=34ymr&%3Btoken=eyJ1c2VyX2lkIjo1MjcwMzU1LCJwb3N0X2lkIjoyMTE5NTI3MjQsImlhdCI6MTc4NzIwMzA2NiwiZXhwIjoxNzg5Nzk1MDY2LCJpc3MiOiJwdWItMTA4NDA4OSIsInN1YiI6InBvc3QtcmVhY3Rpb24ifQ.CPUmsUcEfd5S5SnVTeZt%5F5ZBN0fGgMXOtARH0JSmdmg&ref=antoinebuteau.com)

**2\. The Model Stops Learning At Deployment — X (formerly Twitter)**

- Why read: Explains why transformers struggle to learn continuously and adapt after deployment.
- Summary: Transformers learn during training but freeze after deployment, struggling to adapt to new information. Trying to make them learn continuously either wastes data or causes them to forget older knowledge. Labs stick to transformers because they are currently profitable, ignoring flaws like token-by-token reasoning. Finding a better architecture will mean automating research to test ideas faster. Keep an eye out for new designs and hardware that fix these continuous learning problems.
- [Read more](https://twitter.com/gokulr/status/2090460769234424240/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

**3\. The HALO Effect — kwokchain**

- Why read: Breaks down the controversial HALO deals that are changing how big tech acquires AI startups.
- Summary: Tech giants are using "Hire and License Out" (HALO) deals to bypass antitrust rules. They hire a startup's team and license their IP, leaving the startup as a shell. This shows that an AI company's real value is its talent, not its code or products. While legally complex, these deals quickly pay out founders, employees, and investors. For startups, proven AI talent is worth more than ever, so keep your team tight and move fast.
- [Read more](https://kwokchain.com/2025/07/15/the-halo-effect/?ref=antoinebuteau.com)

**4\. System Design for Agent Systems (Part 1) — X (formerly Twitter)**

- Why read: An engineering guide on building reliable agent architectures that handle the chaos of non-deterministic workloads.
- Summary: Agent systems aren't like web servers. They are unpredictable, wait on I/O, and can burn money quickly if stuck in a loop. A good agent framework needs typed tools, clear prompts, and separated short-term and long-term memory. Running agents in sequence usually works better than running them in parallel because it makes tracking state easier. Always set strict stop conditions like limits on steps, time, or tokens to prevent endless loops. Use an LLM as a judge for continuous evaluation, as standard CI tests won't catch reasoning regressions.
- [Read more](https://twitter.com/kmeanskaran/status/2090436724250026177/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

**5\. Data Is a Great Place to Start an AI Company, and a Dangerous Place to Stop — X (formerly Twitter)**

- Why read: Why the AI data supply chain is changing, and why owning a specific workflow is better than hoarding static data.
- Summary: Labs spend billions on training data, but their needs change with every new model. As models improve, they outgrow synthetic data and need real-world logs of actual work. Chinese startups are building tight loops of task design, generation, and evaluation to meet this need. Data vendors can't just sell datasets forever; they need to build their own models or embed themselves into enterprise software. The real prize is ongoing access to real user workflows and the mistakes people make.
- [Read more](https://twitter.com/hannahhaina/status/2090519081279705359/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

**6\. From Demo to Business: The Race in Physical AI — X (formerly Twitter)**

- Why read: Looks at the race to bring physical AI out of the lab and into the real world.
- Summary: Physical AI is a massive market, but robots don't have the internet's free data to learn from. Some teams are building general foundation models, while others build specialized robots for single tasks in logistics or data centers. Foundation models can learn quickly but still miss the 99.9% reliability needed for real work. The specialized, deployment-first companies are pulling ahead because they gather real-world data directly from enterprise clients. The winners will be those who can afford the hardware while constantly collecting specific, real-world data.
- [Read more](https://twitter.com/kshenster/status/2090482498275070358/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

**7\. To Invent Waymo, They Had to Reinvent PM — Substack**

- Why read: Lessons from Waymo on how to manage AI products when you can't run standard A/B tests.
- Summary: Standard SaaS product management doesn't work for AI because correct behavior is hard to define. Waymo treated their evaluation metrics as products themselves, hiring PMs just to make their tests harder and more realistic. Since AI projects take years, PMs have to set hard, short-term deadlines to keep the team moving. Leading these teams requires technical credibility, the guts to delay an unsafe release, and empathy when big ideas don't work out. Focus on defining exactly what success looks like, and update that definition constantly.
- [Read more](https://substack.com/app-link/post?publication%5Fid=92860&%3Bpost%5Fid=211073969&%3Butm%5Fsource=cross-post&%3Butm%5Fcampaign=10845&%3BisFreemail=true&%3Br=34ymr&%3Btoken=eyJ1c2VyX2lkIjo1MjcwMzU1LCJwb3N0X2lkIjoyMTEwNzM5NjksImlhdCI6MTc4NzI2OTkzMCwiZXhwIjoxNzg5ODYxOTMwLCJpc3MiOiJwdWItOTI4NjAiLCJzdWIiOiJwb3N0LXJlYWN0aW9uIn0.h2bOAxJNAkQsDr0Zt7E5aSXtYQC2YV6OlLqRHcRjD6k&ref=antoinebuteau.com)

**8\. 🧠 Nvidia is the Fannie Mae for Compute — beehiiv.com**

- Why read: Compares the complex financial structures funding AI compute today with the mortgages that caused the 2008 financial crisis.
- Summary: AI infrastructure is being funded by debt, special purpose vehicles, and contracts that assume demand will always grow. Even though GPUs are being used and generating revenue, power grid limits and delayed data center permits mean chips sit idle. If demand slows down, these highly leveraged deals could collapse. Compute is becoming a financial asset, packaged and sold to Wall Street. Tech leaders should review their ties to these providers and brace for impact if the funding dries up.
- [Read more](https://link.mail.beehiiv.com/v2/c/c51f830dce02661cdc6b7d91c89609ec8b113b42ec26b6effbf7aaf9662414adf8f16a26627038eb8942c1876bfb653ad5feb158c42f974124a5bc268d5ab27353cc0e59bee250c4b7311009eaba2afec05ab76458a814c421cfed00e8d55b9e1bcd94f2214840492be6d8b01d4b1ea888e8f1b367798bbd7dca1fd73055c9178878d837b629554a259bfc337259e3574f4384165455c74cfcb53ce6c070eab3/41e36e810d566593?ref=antoinebuteau.com)

**9\. RL Policy Churn — X (formerly Twitter)**

- Why read: John Carmack explains "policy churn," a quirk in reinforcement learning where neural networks randomly change their minds.
- Summary: In reinforcement learning, networks experience "policy churn." The system randomly changes its predicted best action during training, even if you tell it not to explore. This happens because the network generalizes too much; fixing one mistake alters predictions everywhere else. This internal chaos often overrides any deliberate exploration rules engineers write. If you try to fix this churn with sparser networks, you have to manually add exploration back in, or performance will tank.
- [Read more](https://twitter.com/ID%5FAA%5FCarmack/status/2090514515129520516/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

**10\. Scaling Real-Time Text-to-Speech Inference — X (formerly Twitter)**

- Why read: A technical look at restructuring text-to-speech pipelines to get the low latency needed for voice agents.
- Summary: To make voice agents respond instantly, you have to split the text-to-speech pipeline into separate stages: prefill, latent generation, and waveform decoding. This keeps the GPU busy. Batching work constantly and using paged KV caching stops the system from wasting time copying data. Using CUDA Graphs speeds things up further by skipping CPU overhead on repetitive tasks. Done right, this drops the time before the audio starts to under 30 milliseconds. In production, you have to balance memory usage with setup times to keep the voice smooth under heavy load.
- [Read more](https://twitter.com/DecagonAI/status/2090497367233773904/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

**11\. forget about the ARR — X (formerly Twitter)**

- Why read: A warning for investors: buying an AI startup based on standard SaaS metrics is a massive risk.
- Summary: Standard due diligence doesn't work for AI companies. Acquirers mistake API wrappers for real platforms. They miss big risks like dependence on third-party models, shaky data rights, and having only one engineer who understands the system. You have to read the code to see if the product has a real advantage or if someone could clone it in a weekend. Skipping this step leads to bad deals that fail 80% of the time. Buyers need to look hard at the tech before signing anything, or they might just buy an expensive configuration file.
- [Read more](https://twitter.com/mardehaym/status/2090478724290359520/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

**12\. Rewriting Infrastructure Primitives for Agents — X (formerly Twitter)**

- Why read: Why developer tools and databases need a redesign now that AI agents are the ones using them.
- Summary: AI agents are breaking existing software infrastructure. Tools built for humans can't handle the speed of machines. Standard version control struggles with the volume of code agents write. Databases are shifting from permanent fixtures to temporary sessions spun up by the thousands. Even search engines need an overhaul to deliver dense, cheap text instead of visual links. Infrastructure companies need to focus entirely on APIs and protocols like MCP, because visual UIs matter less every day.
- [Read more](https://twitter.com/ProbyShandilya/status/2090468621663277240/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

**13\. Context Is Data: AI Builders Should Start Treating It That Way — X (formerly Twitter)**

- Why read: Why we need to treat the text we feed into AI context windows with the same care as traditional database records.
- Summary: Most AI apps just dump text into context windows. They don't track where the data came from. Since context drives the model's output, losing track of sources makes the AI unpredictable and hard to verify. We need to track dependencies so that when source documents change, the AI updates its context. We also need to save the intermediate steps models generate instead of throwing them away. Managing AI context like traditional data will make systems more reliable.
- [Read more](https://twitter.com/JoshARosen/status/2090461097178341452/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

**14\. Toward measuring recursive self-improvement — X (formerly Twitter)**

- Why read: Looks at how researchers are testing AI agents that can write code to optimize themselves.
- Summary: Recursive self-improvement happens when an AI manages its own training loops without a human. In current tests, agents are writing and saving optimization code to use on themselves later. The main challenge is making sure the feedback they use to grade themselves is actually correct. If the AI handles the optimization, human engineers can focus entirely on checking that feedback layer. We need better ways to measure this progress as agents shift from chatbots to self-teaching systems.
- [Read more](https://twitter.com/anndvision/status/2090411367819809047/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

**15\. Why Real-World Data Is the Next Battleground — X (formerly Twitter)**

- Why read: Why logs of actual human work and economic activity are becoming the most valuable data for training AI.
- Summary: Having experts perform fake tasks for training data isn't working anymore. To train long-term agents, models need logs of real work: what the user wanted, what tools they clicked, their mistakes, and the final result. Real workflows end in clear outcomes like a merged PR or a closed deal, which provide a reliable grading system. Getting this data means navigating messy privacy and consent laws. Securing the legal rights to these internal workflow logs will soon be more valuable than the data itself.
- [Read more](https://twitter.com/hannahhaina/status/2090519081279705359/?rw%5Ftt%5Fthread=True&ref=antoinebuteau.com)

### Themes from yesterday

- **Data and Context Rigor:** The industry is starting to treat AI context and training data like traditional database records, tracking where they come from and how they change over time.
- **Overhauling AI Finance:** Investors and acquirers are waking up. Deal teams are reading code instead of just looking at revenue, and there are growing concerns about the debt funding the massive GPU buildout.
- **The Agent Harness Becomes the Moat:** Adding parameters isn't enough anymore. The real advantage comes from the engineering around the model: reinforcement learning setups, state tracking, and automated testing loops.
- **Infrastructure Built for Agents:** Developer tools, databases, and search engines are being rebuilt from the ground up because software is now being consumed by fast-moving AI agents, not humans.