Research & Deep Dives

Research Explainers

Research papers, reports, scenarios, and technical findings translated into practical implications for builders and operators.

Open Research & Deep Dives

Agent Skills Need a Trust Layer

A survey argues that agent skills are becoming the packaging layer for procedural AI work, making governance, permissions, and verification unavoidable. It maps the controls needed before reusable instructions can be trusted across teams and organizations.

The Agent Is the Whole System

A survey argues that agent quality is a system property spanning models, memory, tools, planners, verifiers, permissions, traces, and evaluation. It offers a practical architecture for diagnosing failures without blaming the base model by default.

Multi-Agent Systems Fail Like Organizations

A NeurIPS dataset paper finds that multi-agent LLM systems fail through role confusion, broken handoffs, and weak verification, not only weak models. Its taxonomy turns coordination breakdowns into observable failure modes teams can test and repair.

Skill Bloat Is the New Context Tax

A paper argues that agent skills need a build-time optimization pass because many reusable instruction files waste context and make agents worse. Its proposed compiler trims redundancy while preserving the instructions that actually improve task performance.

AI Fiction Has a Plot Fingerprint

StoryScope suggests that AI fiction can be detected from narrative decisions, not only surface style, including tidy plots, explicit themes, and reduced structural variety. The findings suggest provenance detection should examine story structure alongside lexical fingerprints.

Synthetic Societies Need Belief Models, Not Personas

A NeurIPS position paper argues that LLM social simulations need traceable belief models, not merely fluent personas and plausible survey answers. Explicit belief state, update rules, and validation make simulated populations more scientifically useful.

Agent Skills Are the New Supply Chain

A systematization paper argues that reusable agent skills are becoming the procedural layer of agent systems, with meaningful performance gains and supply-chain risk. The survey maps benefits alongside provenance, permissions, poisoning, and dependency management.

Agent Reliability Is a Runtime Problem

A preprint argues that agent reliability is governed by the runtime harness around the model: execution, tools, context, state, lifecycle policy, and evaluation. The framework gives teams a way to reason about reliability beyond benchmark scores.

Choose your reading rhythm.

Start with a weekly briefing, add daily notes, or hear only when a durable essay or library update is ready.

You've successfully subscribed to Antoine Buteau
Great! Next, complete checkout to get full access to all premium content.
Welcome back! You've successfully signed in.
Unable to sign you in. Please try again.
Success! Your account is fully activated, you now have access to all content.
Error! Stripe checkout failed.
Success! Your billing info is updated.
Error! Billing info update failed.