Research & Deep Dives

Research Explainers

Research papers, reports, scenarios, and technical findings translated into practical implications for builders and operators.

Open Research & Deep Dives

Long Context Needs Recursion, Not Bigger Windows

Recursive Language Models argue that very long prompts should become an external environment the model can inspect, decompose, and recursively call itself over. Recursion converts context management from passive storage into an active reasoning process.

Agent Logs Should Be the System of Record

An agent architecture argues that the append-only event log should be the runtime's source of truth, not an observability layer bolted on afterward. That choice makes recovery, replay, observability, and policy enforcement part of one architecture.

Synthetic Data Needs Recipes, Not Bigger Generators

FinePhrase shows that high-quality synthetic pretraining data comes from prompt recipes, mix-in strategy, output diversity, and data infrastructure, not simply larger generator models. The result is a repeatable engineering discipline for controlling quality, coverage, and cost.

The Harness Is the Reliability Layer

A survey argues that agent reliability depends as much on the execution harness as on the model and gives builders a vocabulary for the infrastructure surrounding agents. It organizes execution, context, state, tools, recovery, and evaluation into one reliability model.

Agents Need Reliability Tests After Day One

AgingBench argues that deployed agents can degrade as their memory state changes, so reliability needs lifespan testing rather than only day-one benchmarks. Its evaluation tests how accumulated state and environmental change affect performance over time.

Long-Context Models Need Time to Consolidate

Language Models Need Sleep argues that long-context systems may require offline consolidation time, not merely larger memory, to reason over context they can no longer attend to. The analogy motivates architectures that compress experience between active reasoning cycles.

AI Authorship Has Been Here Before

A legal history argues that AI authorship is not a new copyright problem but the latest form of mediated authorship. Earlier disputes over photography and mechanical creation reveal recurring legal tests for control and originality.

Train the Agent Where It Actually Runs

Polar shows how to train coding agents through the harnesses they already use instead of rebuilding those environments for reinforcement learning. The approach preserves realistic tools and feedback while making training operationally tractable.

Choose your reading rhythm.

Start with a weekly briefing, add daily notes, or hear only when a durable essay or research update is ready.

You've successfully subscribed to Antoine Buteau
You've successfully subscribed to Antoine Buteau
Welcome back! You've successfully signed in.