In this digest
- Building Git infrastructure for agent-scale development
- The 4 levels of agentic software development
- How to Automate Inbound
- Agent Evals 101: Stop Guessing, Start Measuring
- How we do evals at Crosby
- Data is the application
- Inference Is the Most Important Market in Software
- How Personal Agents Get Paid
- On the Nature of the Swarm
- How to Turn Your Skills into Agents
- Designing an Agent-first API: How we re-architected supermemory
- The state of the tech industry in 2026
- Asset Light Software
- Everyone is building the same thing
- How to become an inference engineer in 6 months (builder's guide)
Themes from yesterday
- Core infrastructure updates for agent workloads: From GitHub separating Git compute from storage to Supermemory optimizing API responses for context limits, teams are rebuilding infrastructure to handle persistent agent fleets and high-frequency automated commits.
- Deterministic guardrails for reliable agent execution: Workflows at Vercel, Crosby, and enterprise skill deployments show that dependable automation requires encoding business rules and validation checks directly in code, reserving model reasoning strictly for unstructured inputs.
- Evaluation suites replacing public model leaderboards: Public benchmarks and parameter counts fail to reflect production reliability, prompting teams to build decoupled evaluation harnesses, blinded pairwise human reviews, and regression suites drawn from live failure logs.
- Software margins shifting from fixed SaaS to inference costs: Growing model usage is turning software vendors into inference resellers, squeezing traditional 72% gross margins and shifting long-term business defensibility toward proprietary vertical data and specialized workflows.
1. Building Git infrastructure for agent-scale development — GitHub (X)
- Why read: GitHub is redesigning its core Git architecture to support agent fleets that push millions of commits every day, decoupling compute from durable storage to increase write throughput by up to 35 times.
- Summary: Monthly Git activity on GitHub reached 473.3 billion events after autonomous coding agents drove a fivefold increase in commits and an almost fivefold jump in pushes. Because agents commit after nearly every step, writes created latency bottlenecks in GitHub's legacy Spokes architecture, where every replica had to participate in every write. To support this volume without downtime, GitHub is splitting compute from durable storage. Authoritative repository data now lives in Azure Blob Storage, while stateless, cache-backed workers handle read requests. The redesign also moves compaction and garbage collection off the serving path so reference updates and push acknowledgments stay fast. In internal benchmarks, this decoupled setup delivers up to 35 times higher write throughput while maintaining branch protection rules and audit controls. This prevents core Git infrastructure from bottlenecking parallel agent fleets as automated commits surpass human output.
- Read more
2. The 4 levels of agentic software development — Kaspar von Grünberg (platformengineering.org)
- Why read: A framework showing why platform infrastructure, rather than model capability, sets the ceiling on enterprise productivity gains across four stages of agent autonomy.
- Summary: Enterprise teams often hit a wall with AI productivity when they treat models as desktop assistants instead of adapting their development platforms. The report defines four stages of adoption: Level 1 (human in the loop), Level 2 (humans on the loop with parallel execution), Level 3 (humans as orchestrators of continuous background jobs), and Level 4 (fully autonomous systems responding to telemetry). Moving from Level 1 to Level 2 is the hardest hurdle because manual code reviews collapse under high volumes of concurrent pull requests. Progressing past Level 1 requires non-human machine identities, isolated sandboxes, and automated feedback loops that send test failures directly back to the agent for repair. At Levels 3 and 4, human oversight is reserved for exceptions, shifting engineering effort toward policy rules and token spend. Teams that build these platform foundations can run continuous background refactoring and maintenance that manual processes cannot sustain.
- Read more
3. How to Automate Inbound — Tomasz Tunguz (X)
- Why read: How Vercel automated its top-of-funnel inbound sales qualification for about $1,000 a year by separating deterministic business logic from model judgment.
- Summary: Vercel COO Jeanne DeWitt Grosser explains how a single go-to-market engineer automated inbound lead qualification with a hybrid workflow. The project began with a 1,000-line prompt drafted with top SDRs, but the team found that language models could not reliably follow rigid qualification rules. Engineers resolved this by decoupling deterministic logic from subjective assessment, encoding 14 routing and CRM rules directly in code while using the model solely for unstructured research on prospective accounts. A second escalation agent reviews production runs to catch rule violations and surface edge cases for human inspection. The setup handles company-wide inbound qualification at an operating cost of roughly $1,000 per year, freeing SDRs from routine email screening to focus on outbound conversations.
- Read more
4. Agent Evals 101: Stop Guessing, Start Measuring — Paul Iusztin (X)
- Why read: Practical patterns for building decoupled agent evaluation harnesses, and why domain-specific benchmarks matter much more than public leaderboard rankings.
- Summary: Public model rankings rarely reflect how agents perform in production, illustrated by an optimized 35B model scoring 95% on a coding benchmark while a 120B model scored only 53%. Evaluating agents reliably requires a decoupled test harness where the agent runs inside an isolated sandbox and an independent referee program scores the results. A production testing pipeline relies on three tiers: development benchmarks that evaluate final outputs, regression suites that check multi-step execution traces, and online monitors that sample live traffic. Teams should build regression suites out of real production failures through structured error triage instead of writing synthetic test cases. Measuring performance with binary pass/fail checks and deterministic tool-use metrics gives clearer signal than 1-to-5 rating scales, allowing engineers to refine prompts, tools, and token budgets without causing silent regressions.
- Read more
5. How we do evals at Crosby — Aveek Duttagupta (crosby.ai)
- Why read: How legal AI company Crosby sped up contract reviews fivefold using an evaluation system that combines deterministic structural checks with blinded pairwise human grading.
- Summary: Measuring quality in legal redlining is difficult because human experts often disagree; in one test, two experienced lawyers agreed on only 45% of edits across identical non-disclosure agreements. Crosby addressed this by running automated deterministic checks first, catching structural issues like misplaced comments and undefined terms before reviewing legal substance. For subjective legal quality, the team dropped numerical scoring rubrics in favor of blinded pairwise comparisons, asking attorneys which draft they would send to a client. Engineers also built a Microsoft Word add-in where lawyers review diffs and adjust context parameters without editing prompt code. This feedback loop between legal staff and automated checks reduced NDA review times by 74% and Master Services Agreement reviews by 80%, demonstrating the value of using deterministic filters for formatting while keeping domain experts in familiar interfaces.
- Read more
6. Data is the application — Mattias Geniar (ttias.be)
- Why read: Why code and interface changes made by AI agents are easily reversible, while database architecture and data migrations remain irreversible decisions that require careful human control.
- Summary: Coding agents let engineers build features, update user interfaces, and patch bugs quickly, turning most application code into easily reversible decisions. That flexibility stops at the database. Bad migrations, dropped columns, and corrupted tables cause permanent data loss that cannot be undone with a deployment rollback or reconstructed by a model. While architectures like event sourcing allow data replaying, they create their own rigid requirements around append-only logs and privacy deletion laws. Applying rapid agent iteration to schema changes and data retention policies introduces serious operational and legal risks. In an agent-driven development cycle, application code is disposable, but stored data remains the permanent asset.
- Read more
7. Inference Is the Most Important Market in Software — Tomasz Tunguz (X)
- Why read: Why AI inference spending is projected to reach $350 billion by 2027, surpassing the database market, and how it is squeezing standard SaaS gross margins.
- Summary: Spending on AI model inference grew from $25 billion in 2025 to roughly $130 billion in 2026, and is projected to hit $350 billion by 2027, roughly double the size of the database market. This expansion is turning software providers into inference resellers. As compute and token usage overtake fixed subscription fees, standard software gross margins of around 72% are falling. To stay profitable, vendors must either build custom software harnesses that cut token consumption or shift to Bring-Your-Own-Key models that sacrifice top-line revenue in exchange for higher software margins. Even as unit costs per token decline, newer frontier models keep inference compute as a major cost of goods sold. Software operators now have to manage token budgets with the same scrutiny logistics companies apply to fuel consumption.
- Read more
8. How Personal Agents Get Paid — Tanay Jaipuria (X)
- Why read: An analysis of how consumer AI agents can generate revenue through transaction cuts, affiliate referral fees, marketplace aggregation, and sponsored placement.
- Summary: As personal AI agents begin completing transactions directly for consumers, developers are evaluating how to monetize their services. Revenue is organizing into four models: subscriptions, payment transaction fees, merchant referral partnerships, and advertising. An agent that simply automates checkout earns thin payment processing cuts of 0.2% to 0.5%. Agents that curate recommendations and drive qualified purchases can claim standard affiliate commissions between 1% and 8% across retail and travel. Platforms that aggregate full transactions by handling inventory, booking, and customer service can command take rates of 12% to 15%, though physical delivery remains difficult to disintermediate. Adding sponsored merchant placements yields higher revenue, but creates a conflict between serving the user's best interest and prioritizing advertising partners.
- Read more
9. On the Nature of the Swarm — hallerite's blog
- Why read: An examination of context window limits, explaining why scaling test-time compute leads teams toward persistent multi-agent swarms organized in peer meshes.
- Summary: Scaling test-time compute improves model reasoning, but multi-step tasks still fill the one-million-token context windows of modern frontier models. Context compaction discards detail prematurely because the system cannot predict which points will be relevant later. Storing state on disk avoids data loss, but subsequent agents must burn fresh tokens parsing those files, while ephemeral sub-agents lose their working memory as soon as they finish a prompt. A more resilient pattern keeps persistent sub-agents active in memory, expanding available context across parallel windows. To stop top-level coordinators from running out of context, architectures are shifting from strict hierarchies to peer mesh networks that read and write shared state. These swarms still carry token overhead, requiring teams to balance task decomposition against inter-agent communication costs.
- Read more
10. How to Turn Your Skills into Agents — Eyad (X)
- Why read: How to prevent prompt skill drift in enterprise AI by managing skills like software code, with dedicated owners, automated evaluations, and weekly test reviews.
- Summary: More than half of enterprise executives report seeing no financial returns from AI adoption, while staff spend over a third of their saved hours fixing flawed model outputs. This issue usually stems from treating prompt skills like unmanaged text documents, leaving departments with dozens of conflicting versions and no shared testing. In tasks like invoice coding or financial variance reporting, this variation creates silent errors and prevents teams from rolling out model updates safely. Organizations should manage prompt skills like production software, giving each workflow a single owner and storing prompts in a central repository. Teams should maintain regression test suites of 20 to 50 historical failure cases and run automated evals before merging prompt changes. Reviewing flagged runs weekly to expand the test library routinely improves agent first-pass accuracy from 80% to 95%.
- Read more
11. Designing an Agent-first API: How we re-architected supermemory — Dhravya Shah (X)
- Why read: How Supermemory rebuilt its API for autonomous agents by optimizing for context-window token limits, explicit URL namespaces, and markdown responses.
- Summary: As autonomous coding agents replace humans as the primary consumers of APIs, interfaces need to be designed around agent limitations rather than human convenience. Supermemory rebuilt its API after discovering that traditional responses, such as 23-field database dumps, were consuming too much agent context. The new design replaces ambiguous tenant tags with clean URL paths, clarifying multi-tenant permissions for language models. The endpoints now support markdown content negotiation, delivering compact markdown tailored for model context instead of nested JSON objects that require extra parsing tokens. Supermemory also combined duplicate search endpoints, standardized query filters, and added an automated feedback endpoint to route runtime errors directly to engineers. Using compact schemas and standard REST structures saves tokens and makes agent integrations more dependable.
- Read more
12. The state of the tech industry in 2026 — The Pragmatic Engineer (Substack)
- Why read: A report on software engineering in late 2026, tracking the decline of manual coding, the evolution of the traditional IDE, and the rise of parallel agent fleets.
- Summary: Manual programming is declining across high-performing software teams. Monthly agent-authored pull requests on GitHub increased ninefold in eight months, overtaking human PR volume by August 2026. Engineers at companies like Linear and Cockroach Labs now work primarily as coordinators, managing five to ten parallel agent sessions in separate Git worktrees. This change is turning traditional code editors like Cursor and Antigravity into verification dashboards and agent control planes. At the same time, most mid-sized tech companies are building internal harnesses to manage agent sandboxes, permissions, and tool execution. Engineering teams now depend on automated validation pipelines to verify agent-generated code without overwhelming human reviewers.
- Read more
13. Asset Light Software — Will Manidis (X)
- Why read: How generative AI is reshaping SaaS economics by converting fixed development and go-to-market payroll into variable costs, creating asset-light software businesses.
- Summary: Traditional software companies typically spend about 90% of recurring revenue on fixed headcount across engineering, sales, and administration. Generative AI is changing this cost profile by automating junior knowledge work, turning software development and marketing expenses into variable, consumption-based costs. While the internet drove the marginal cost of distribution to zero, AI is reducing the marginal cost of code and content production. This flattens engineering teams into small groups of senior architects who direct agent workflows. As a result, software founders need less venture capital to reach scale, often choosing non-dilutive financing and early cash flow over equity dilution. Because standard software features are becoming easier to reproduce, business defensibility is moving toward proprietary vertical data integrations and direct customer relationships.
- Read more
14. Everyone is building the same thing — Angular Ventures
- Why read: An investor assessment of why common agent loops, memory graphs, and sandboxes are now commodity infrastructure, leaving startup defensibility to vertical domain expertise.
- Summary: Many AI startups are converging on identical architectures built around agent loops, memory layers, tool connectors, and sandboxed runtimes. Just as user authentication and databases turned into standard web building blocks after 2005, these agent plumbing components are becoming basic table stakes instead of competitive advantages. Pitch decks frequently feature the same infrastructure patterns without clear business differentiation. Sustainable advantages come from deep industry expertise and customer relationships rather than custom prompt harnesses, token tricks, or context graphs. The strongest opportunities sit with vertical software products embedded in complex enterprise workflows, where private operational data and strict accuracy standards protect against easy model replacement. Founders create more value by fixing specific operational bottlenecks than by rebuilding common agent scaffolding.
- Read more
15. How to become an inference engineer in 6 months (builder's guide) — Avid (X)
- Why read: A 26-week technical study plan for mastering production model inference, covering GPU kernels, KV cache optimization, continuous batching, and distributed serving.
- Summary: AI engineering jobs now represent nearly 7% of technical job postings, making inference infrastructure a central specialty. This 26-week, project-based curriculum outlines how to move from basic model execution to production-grade serving systems. Engineers start by deploying models in vLLM and SGLang to set performance baselines for time-to-first-token, inter-token latency, and total goodput. The guide then covers writing custom CUDA and Triton kernels, profiling memory with Nsight Compute, and setting up KV cache techniques like PagedAttention and prefix caching. Later sections cover multi-tenant serving, chunked prefill scheduling, Kubernetes autoscaling, gateway fault tolerance, and multi-GPU tensor parallelism. Understanding these low-level systems helps teams maximize GPU utilization and keep inference costs under control.
- Read more