Research & Deep Dives

Research Explainers

Research papers, reports, scenarios, and technical findings translated into practical implications for builders and operators.

Open Research & Deep Dives

AI Adoption Fails When Firms Cannot Map It to Work

A field experiment with 515 startups suggests AI creates firm-level gains when teams learn where to reorganize work around it, rather than merely gaining access to tools. This explainer connects adoption outcomes to workflow mapping, experimentation, and organizational change.

Continuous Scoring Turns Verification Into a New AI Scaling Axis

By replacing discrete judgments with continuous probabilistic scoring, researchers show how verification can improve agent performance without additional training. This explainer examines why better scoring may become a distinct scaling axis alongside model size, inference compute, and data.

Agentic AI Turns Work Into Delegation

OpenAI's Codex usage data suggests the agentic shift is less about better chat answers and more about delegating longer, reusable, parallel workflows. It explains why duration, concurrency, and reuse matter more than conversational polish.

Prompt Tuning Stops Working When the Workflow Is the Problem

FAPO uses Claude Code to optimize multi-step LLM pipelines by inspecting intermediate failures, trying prompt fixes first, and changing the workflow only when evidence points to a structural bottleneck. The result is a practical diagnostic loop for builders.

How Much Future Has the Market Already Bought?

Michael Mauboussin and Dan Callahan use PVGO to separate current business value from future growth expectations, giving investors a cleaner view of what the market has already priced in. It works as an expectations lens, not a timing signal.

The Rise of the AI-Native Firm: How Startups Scale Without Large Teams

AI-native startups are smaller, flatter, more technical, and more valuable per employee because AI is embedded in the product, not merely internal workflows. The analysis connects organizational design, technical founder density, and valuation economics.

AI 2027 Is a Scenario, Not a Prediction

A detailed scenario about agents, automated AI research, model misalignment, geopolitical competition, and the choice between racing and slowing down. Its value lies in testing assumptions about timelines, coordination, and control rather than treating the scenario as destiny.

Agent Coding Costs Hide in Review, Not Generation

A study of ChatDev traces finds that agentic coding systems spend most tokens on review and repeated context passing, not initial code generation. The finding shifts optimization toward handoffs, critique loops, and context efficiency.

Choose your reading rhythm.

Start with a weekly briefing, add daily notes, or hear only when a durable essay or research update is ready.

You've successfully subscribed to Antoine Buteau
You've successfully subscribed to Antoine Buteau
Welcome back! You've successfully signed in.