Research & Deep Dives

Research Explainers

Research papers, reports, scenarios, and technical findings translated into practical implications for builders and operators.

Open Research & Deep Dives

Agents Change Work by Lowering the Cost of Execution

A Perplexity field study compares conversational search with autonomous agent execution and finds large time savings, lower dissatisfaction, and broader task scope. The comparison shows why autonomy changes both the economics and ambition of knowledge work.

Synthetic Consumers Work Better When They Talk First

A consumer research paper finds that LLMs can match human purchase-intent surveys when they answer in free text before being mapped back to Likert ratings. The result favors better survey design over treating synthetic respondents as wholesale human replacements.

AGI May Be a Phase, Not the Finish Line

A DeepMind report argues that human-level AGI may be followed by several pathways toward superintelligence, each constrained by data, compute, embodiment, and coordination bottlenecks. It reframes AGI as an intermediate capability threshold rather than a stable endpoint.

Post-Training Is Where Models Learn Bad Habits

A post-training paper shows how interpretability tools can audit preference data, expose unwanted learning signals, and reshape rewards before models absorb bad habits. The method turns interpretability into a practical intervention in data curation and reward design.

Agent Harnesses Can Learn From Their Own Failures

Self-Harness shows that agents can use execution traces, targeted edits, and regression tests to improve the scaffolding around their own model. The important constraint is preserving task performance while the runtime changes itself.

New Technologies Take Time to Reach the Productivity Stats

Paul David's dynamo analogy explains why general-purpose technologies can spread through business before appearing clearly in productivity statistics. The lesson is to distinguish implementation lag from evidence that the technology lacks value.

The Agent That Updates the Harness Is Not the Bottleneck

A paper separates the model that improves an agent's harness from the model that must use it. Cheap models may propose useful scaffold changes, but task performance still depends on an executor capable enough to load and follow them.

Bigger Models Remember the Rare Stuff Long Enough to Learn It

A scaling paper argues that larger models win partly because frequent tasks stop overwriting rare-task features. The proposed mechanism links parameter scale to reduced interference and better retention of infrequent capabilities over time.

Choose your reading rhythm.

Start with a weekly briefing, add daily notes, or hear only when a durable essay or research update is ready.

You've successfully subscribed to Antoine Buteau
You've successfully subscribed to Antoine Buteau
Welcome back! You've successfully signed in.