Research & Deep Dives

Research Explainers

Research papers, reports, scenarios, and technical findings translated into practical implications for builders and operators.

Open Research & Deep Dives

The Harness Is the Reliability Layer

A survey argues that agent reliability depends as much on the execution harness as on the model and gives builders a vocabulary for the infrastructure surrounding agents. It organizes execution, context, state, tools, recovery, and evaluation into one reliability model.

Agents Need Reliability Tests After Day One

AgingBench argues that deployed agents can degrade as their memory state changes, so reliability needs lifespan testing rather than only day-one benchmarks. Its evaluation tests how accumulated state and environmental change affect performance over time.

Long-Context Models Need Time to Consolidate

Language Models Need Sleep argues that long-context systems may require offline consolidation time, not merely larger memory, to reason over context they can no longer attend to. The analogy motivates architectures that compress experience between active reasoning cycles.

AI Authorship Has Been Here Before

A legal history argues that AI authorship is not a new copyright problem but the latest form of mediated authorship. Earlier disputes over photography and mechanical creation reveal recurring legal tests for control and originality.

Train the Agent Where It Actually Runs

Polar shows how to train coding agents through the harnesses they already use instead of rebuilding those environments for reinforcement learning. The approach preserves realistic tools and feedback while making training operationally tractable.

Food AI Needs Knobs Beyond Recommendations

Epicure shows how ingredient embeddings can become navigable tools for cooking, menu design, and food AI. The system turns flavor relationships into controllable dimensions for substitution, exploration, menu planning, and creative composition rather than returning another opaque ranking.

Agent Skills Are Becoming Trainable Assets

SkillOpt treats reusable agent skills as operating assets that can be trained, validated, and reused. It shows how procedural instructions can be optimized against tasks, checked for regressions, and distributed with evidence.

Why AI Coding Agents Fail When Software Gets Real

Constraint decay explains why AI coding agents struggle when backend work carries real architectural and database rules. The paper traces how requirements fade across long trajectories and proposes stronger constraints, memory, and verification.

Choose your reading rhythm.

Start with a weekly briefing, add daily notes, or hear only when a durable essay or library update is ready.

You've successfully subscribed to Antoine Buteau
Great! Next, complete checkout to get full access to all premium content.
Welcome back! You've successfully signed in.
Unable to sign you in. Please try again.
Success! Your account is fully activated, you now have access to all content.
Error! Stripe checkout failed.
Success! Your billing info is updated.
Error! Billing info update failed.