Research & Deep Dives

Research Explainers

Research papers, reports, scenarios, and technical findings translated into practical implications for builders and operators.

Open Research & Deep Dives

Prompt Tuning Stops Working When the Workflow Is the Problem

FAPO uses Claude Code to optimize multi-step LLM pipelines by inspecting intermediate failures, trying prompt fixes first, and changing the workflow only when evidence points to a structural bottleneck. The result is a practical diagnostic loop for builders.

How Much Future Has the Market Already Bought?

Michael Mauboussin and Dan Callahan use PVGO to separate current business value from future growth expectations, giving investors a cleaner view of what the market has already priced in. It works as an expectations lens, not a timing signal.

The Rise of the AI-Native Firm: How Startups Scale Without Large Teams

AI-native startups are smaller, flatter, more technical, and more valuable per employee because AI is embedded in the product, not merely internal workflows. The analysis connects organizational design, technical founder density, and valuation economics.

AI 2027 Is a Scenario, Not a Prediction

A detailed scenario about agents, automated AI research, model misalignment, geopolitical competition, and the choice between racing and slowing down. Its value lies in testing assumptions about timelines, coordination, and control rather than treating the scenario as destiny.

Agent Coding Costs Hide in Review, Not Generation

A study of ChatDev traces finds that agentic coding systems spend most tokens on review and repeated context passing, not initial code generation. The finding shifts optimization toward handoffs, critique loops, and context efficiency.

Agents Change Work by Lowering the Cost of Execution

A Perplexity field study compares conversational search with autonomous agent execution and finds large time savings, lower dissatisfaction, and broader task scope. The comparison shows why autonomy changes both the economics and ambition of knowledge work.

Synthetic Consumers Work Better When They Talk First

A consumer research paper finds that LLMs can match human purchase-intent surveys when they answer in free text before being mapped back to Likert ratings. The result favors better survey design over treating synthetic respondents as wholesale human replacements.

AGI May Be a Phase, Not the Finish Line

A DeepMind report argues that human-level AGI may be followed by several pathways toward superintelligence, each constrained by data, compute, embodiment, and coordination bottlenecks. It reframes AGI as an intermediate capability threshold rather than a stable endpoint.

Choose your reading rhythm.

Start with a weekly briefing, add daily notes, or hear only when a durable essay or library update is ready.

You've successfully subscribed to Antoine Buteau
Great! Next, complete checkout to get full access to all premium content.
Welcome back! You've successfully signed in.
Unable to sign you in. Please try again.
Success! Your account is fully activated, you now have access to all content.
Error! Stripe checkout failed.
Success! Your billing info is updated.
Error! Billing info update failed.