In this digest
- Crazy few weeks for model releases!
- Automating eval design and hillclimbing with Claude
- OpenAI DevDay: Architectural Trends to Watch Tomorrow
- Why has Shopify dropped React Native?
- Aggregation Theory 2.0: who wins when agents do the buying
- Is Jev overhyped? We tested it on 4 real enterprise tasks.
- 🧠Agent Wars: The agentic bank run?
- Are Pre-AI Vendors Spiraling Out?
- Five frontier launches and one unanswered question
- OpenAI Understands Something Important and Rare
- Ghostwriter: When AI goes from tool to teammate
- To Seek a Newer World
- GTM Engineering, Explained: The Role, the Stack and the First 5 Systems to Build.
- You can't point a GPU at Cancer
- This is what I'm seeing right now from over 50...
Themes from yesterday
- The Economic Shift from Benchmark Leaderboards to Task-Level Costs: Model evaluation is moving away from headline benchmark scores toward the total cost of finishing a job. Databricks' internal testing and new model-routing harnesses show that choosing the right model tier for each task matters more than chasing top-line rankings.
- The Erosion of Cross-Platform Software Abstractions: AI coding assistants make maintaining separate codebases much cheaper. As Shopify's return to native Swift and Kotlin shows, the engineering savings of cross-platform frameworks no longer outweigh the performance and tooling benefits of native code.
- Agentic Disintermediation and Incumbent Defensiveness: Consumer agents are managing transactions directly in banking and retail. In response, incumbent platforms and financial institutions are alternating between blocking bot traffic and charging metering fees to protect their customer relationships.
- From Interactive Tools to Proactive Systems: AI software is evolving from reactive chat windows into background teammates, managed runtimes, and automated sales pipelines that run continuously to hit operational targets.
1. Crazy few weeks for model releases! — Patrick Wendell (X)
- Why read: Databricks tested frontier models across 2,400 internal engineers, showing how Opus 5.5 and GPT-6 Luna change the cost and performance baseline for daily software work.
- Summary: Databricks evaluated recent frontier models using offline benchmarks and live telemetry from 2,400 engineers. Opus 5.5 emerged as the team's primary coding model, outperforming earlier Opus versions and GPT-6 variants while reducing inference costs by 20% on identical tasks. At the lower end, GPT-6 Luna costs at least 20 times less than Opus 5.5 while matching older flagship models such as Opus 4.6. That shift reflects a 100x drop in cost over nine months, resetting the economics of high-frequency agent loops. For engineering teams, the data points to routing routine agent subtasks to budget models like Luna while keeping primary interactive coding on Opus 5.5 through a centralized gateway.
- Read more
2. Automating eval design and hillclimbing with Claude — Lance Martin (Anthropic)
- Why read: Anthropic outlines practical methods for building repository-level evaluations and automating application improvements without overfitting to test harnesses.
- Summary: Anthropic added guided workflows to its claude-api skill for creating test suites and running automated hillclimbing inside code repositories. The approach pulls test inputs from production traffic and human-verified edge cases rather than cataloging model-specific quirks. To prevent harness overfitting, where changes raise benchmark numbers without helping real users, the workflow enforces strict train-test splits and confirms that score gains exceed measurement noise before patches merge. Evaluation checks follow a hierarchy: deterministic programmatic tests run first, followed by calibrated model judges with verifiable rubrics for open-ended outputs. Teams running production agents can use these regression gates and diagnostics to tune prompts, tools, latency, and costs.
- Read more
3. OpenAI DevDay: Architectural Trends to Watch Tomorrow — Josh Rosen (X)
- Why read: A breakdown of ten architectural patterns heading into DevDay, focusing on the transition from standalone model calls to managed agent infrastructure.
- Summary: Ahead of DevDay, OpenAI's product direction centers on managed harnesses, execution environments, and asynchronous tool calling rather than small benchmark gains. Key patterns include front-model and back-model decoupling, which separates responsive user interfaces from heavy background reasoning, and programmatic tool calling, where models write executable code to run loops and transform data. Dynamic features like Tool Search pull in tools on demand to avoid context bloat, while cloud sandboxes now provide full virtual machines with browsers and filesystems. Multi-agent systems are also moving away from static graphs, treating subagents as dynamic compute allocations within an elastic inference budget. As the Responses API turns into low-level plumbing, engineering teams will increasingly build on top of managed agent runtimes.
- Read more
4. Why has Shopify dropped React Native? — The Pragmatic Engineer (Substack)
- Why read: Shopify is leaving React Native after five years because AI coding tools have made maintaining separate Swift and Kotlin codebases much cheaper.
- Summary: Shopify is moving its mobile apps back to fully native development after standardizing six apps on React Native in 2020. The primary catalyst is the progress of AI coding agents, which now assist with writing, porting, testing, and reviewing code across Swift and Kotlin. To validate the approach, Shopify's mobile team rewrote its flagship Shop app into native code in twelve weeks. Because AI reduces the staffing overhead of supporting two codebases, the convenience of a shared abstraction no longer offsets native performance, debugging tools, and platform integrations. Engineering leaders may want to re-examine their mobile tech stacks, since AI assistance makes separate native apps practical again.
- Read more
5. Aggregation Theory 2.0: who wins when agents do the buying — Igor Chalhub (X)
- Why read: An analysis of how buying agents alter Ben Thompson's Aggregation Theory by adding real inference costs to browsing and prompting major retailers to block automated access.
- Summary: Classic internet aggregation depended on near-zero distribution costs and modular suppliers who had to participate where consumer demand aggregated. Purchasing agents capture user intent end-to-end, but they break the zero-marginal-cost model because browsing and research runs consume inference tokens even when users never buy. Simultaneously, dominant retailers like Amazon block third-party agents to protect their customer relationships and merchant-of-record positions. As commerce splits into six layers (intent, selection, payment, merchant of record, servicing, and data), value moves toward proprietary inference hardware and non-routable inventory. Teams building consumer shopping agents must watch unit economics closely to ensure commission revenues cover the compute costs of multi-step sessions.
- Read more
6. Is Jev overhyped? We tested it on 4 real enterprise tasks. — Tony Gentilcore (X)
- Why read: Glean tested Jev's typed zero-shot decision model across four enterprise tasks, benchmarking it against standard LLMs, fine-tuned classifiers, and search rerankers.
- Summary: Glean evaluated Jev across query classification, expert model routing, search reranking, and citation-support checks. Fine-tuned small models still won on known classification tasks, and dedicated rerankers performed better on precision search. However, Jev proved superior for model routing, delivering an 8.1x median speedup and better accuracy than prompt-based LLM baselines. For citation evaluation, Jev produced more consistent paragraph-level verification than high-reasoning GPT-5.6 configurations while using far less latency and compute. Jev functions well as a zero-shot, fast decision layer for fixed categorical choices, but generative models remain necessary when workflows require dynamic arguments or freeform text. Teams can use typed classifiers to route traffic and prune agent branches, saving generative calls for complex tasks.
- Read more
7. 🧠Agent Wars: The agentic bank run? — Simon Taylor (beehiiv.com)
- Why read: Personal consumer agents like Meta's Muse could disintermediate retail banks by automating rate shopping and eroding the profits built on customer inertia.
- Summary: Meta's launch of Muse raised concerns among banks that autonomous agents might trigger deposit outflows by moving idle balances to higher-yielding accounts. The more immediate threat is disintermediation. Consumers will likely keep their direct-deposit checking accounts while letting software choose their lending, insurance, and savings providers. Retail banking profits rely on user inertia; McKinsey estimates deposit profit pools could drop by 20% or more if automated agents shop every renewal. While some banks attempt to block automated traffic, durable adaptations will require secure APIs and Model Context Protocol (MCP) endpoints that accommodate external agents. Institutions will need to compete on product terms as headless infrastructure rather than relying on proprietary mobile app engagement.
- Read more
8. Are Pre-AI Vendors Spiraling Out? — SaaStr
- Why read: Legacy SaaS vendors are adding surcharges and metering for agent access, risking customer defection as enterprise software and token budgets grow.
- Summary: Established software vendors are introducing new fees for autonomous agent access, including HubSpot metering custom agent credits and Salesforce charging Flex Credits on third-party API and MCP calls. Adding tollgates to records risks driving customers toward AI-native platforms that do not penalize automated tooling. At the same time, visibility into AI spending remains weak; median engineering teams now spend $213 per developer weekly on coding model tokens without clear metrics on business output. With consumption costs rising, enterprise buyers increasingly require secure microVM isolation and verified data boundaries from their vendors. IT leaders should review their SaaS agreements for agent-specific fees to prevent renewal surprises and favor vendors that offer straightforward API access.
- Read more
9. Five frontier launches and one unanswered question — Emergence Capital (Substack)
- Why read: Public AI benchmarks mask extreme compute costs, making cost per completed task the essential operational metric for enterprise deployments.
- Summary: Ten days brought five frontier model releases, but headline benchmark scores obscure steep inference expenses. On ARC-AGI-3, for instance, a 37-point score difference separated an optimized $26,000 run from a brute-force multi-pass test that cost over $220,000. Because production environments operate under fixed budgets, public leaderboards offer little insight into whether a workflow is commercially practical. The relevant operational metric is cost per completed task, which accounts for model efficiency, inference throughput, and hardware pricing. Maintaining viable margins requires optimizing across the stack, including sparse open-weight models, fast inference runtimes, and routing harnesses that send simpler subtasks to cheaper models. Teams should run internal evaluations that track end-to-end execution costs, including tool calls and retry loops.
- Read more
10. OpenAI Understands Something Important and Rare — David George (X)
- Why read: Why broad user distribution and programmable developer primitives give OpenAI a durable advantage as raw model performance commoditizes across labs.
- Summary: With baseline model capabilities converging across leading labs, competitive durability depends on distribution and shaping user behavior. OpenAI's advantage stems from its direct exposure to consumer and enterprise workflows, which surfaces useful product primitives faster than benchmark testing can. Instead of building narrow vertical applications or relying on exclusive partnerships, OpenAI is developing a programmable runtime where users write code and coordinate agents to handle bespoke tasks. This feedback loop ensures platform features mirror real-world usage patterns. For founders and product teams, expanding aggregate demand through flexible platform infrastructure matters more than protecting proprietary model weights against rival providers.
- Read more
11. Ghostwriter: When AI goes from tool to teammate — Bret Taylor (X)
- Why read: How Sierra shifted its Ghostwriter agent from a prompt-based assistant to a proactive coworker embedded in enterprise Slack and Teams channels.
- Summary: Sierra updated Ghostwriter from a setup utility into an active participant in team Slack and Microsoft Teams channels. The system reviews millions of customer support interactions to identify recurring pain points, propose system fixes, and suggest A/B tests that improve resolution rates. In one enterprise rollout, the agent analyzed call transcripts, identified escalation drivers, and launched statistically sound experiments without waiting for an employee prompt. To prevent channel clutter, Sierra designed the agent to recognize conversation flow, respect workplace tempo, and speak up only when it holds confident, actionable recommendations. This marks a shift from reactive chat prompts to autonomous tools that work alongside teams on operational goals.
- Read more
12. To Seek a Newer World — Fei-Fei Li (X)
- Why read: Dr. Fei-Fei Li details AMD's acquisition of World Labs to pair spatial foundation models with dedicated chip design.
- Summary: AMD has agreed to acquire World Labs, with founder Dr. Fei-Fei Li joining as Executive Vice President and Chief Scientist. World Labs started from the premise that language models alone cannot model the physical world, developing spatial systems like Atlas to generate 3D scenes from sparse 2D photographs. Training and running models for robotics simulation, multi-view reconstruction, and physical reasoning demands specialized compute architectures. The combined group plans to merge World Labs' 3D generative architectures with AMD's silicon design and inference optimizations to support physical-world AI. The move indicates that spatial foundation models and real-world simulation are now strategic priorities for major chipmakers.
- Read more
13. GTM Engineering, Explained: The Role, the Stack and the First 5 Systems to Build. — Adam Rahman (X)
- Why read: A practical guide to building a go-to-market engineering stack that combines multi-directory prospecting, dual email verification, and automated reply routing.
- Summary: Go-to-market engineering combines software development, sales operations, and data pipelines to automate routine SDR work. A modern GTM stack divides tasks cleanly: Jev handles scoring and classification, while coding agents like Claude Code run queries and aggregate cross-database records. The process starts by mapping target accounts across multiple directories, followed by dual-engine email verification before outbound sends. Incoming responses are categorized by intent, sending warm prospects to account executives within minutes alongside automatically generated call notes while filtering out opt-outs. Sales leaders can implement these modules sequentially, maintaining human review on messaging copy and advertising budgets while letting code handle data aggregation and routing.
- Read more
14. You can't point a GPU at Cancer — dylan ツ (X)
- Why read: Why AI compute concentrates on commercial problems like digital advertising rather than complex medical research, and what is required to change that.
- Summary: Compute clusters follow clear scoreboards, fast feedback loops, and paying customers. Digital advertising offers instant click metrics and direct billing, whereas biology lacks a single loss function and requires multi-year wet-lab and clinical validation. DeepMind's AlphaFold succeeded because structural biologists spent fifty years assembling the Protein Data Bank and running CASP benchmark competitions, providing clean training data before GPUs were applied. Bringing AI to hard scientific problems requires foundational legwork: automated wet labs, standardized open datasets, and faster physical assays. Funders and research institutions need to build clear physical-world benchmarks and advance market commitments so complex biological problems become computationally tractable.
- Read more
15. This is what I'm seeing right now from over 50... — Zain Manji (X)
- Why read: Practical lessons from more than 50 enterprise AI implementations across fintech, healthcare, retail, and SaaS.
- Summary: Deployments across dozens of companies show that enterprise AI adoption depends on data preparation, often requiring two to four months of basic cleanup before software layers can run. At the same time, enterprise rollouts face a shortage of forward-deployed engineers who can bridge the gap between initial demos and reliable production deployments. Security teams and regulators are mandating data sovereignty and centralized governance hubs, partly to rein in employees using unsanctioned personal model accounts. Customer support agents are seeing heavy investment from private equity firms looking for direct cost reductions and 24/7 revenue capture. Technology executives should organize their data pipelines and security controls before rolling out agent workflows to keep projects from stalling.
- Read more