Topic

AI

Essays, profiles, and reading notes on AI products, AI systems, adoption, agents, and company strategy.

Start Here

A few useful paths through AI

Related Hubs

Move sideways by problem area.

Post-Training Is Where Models Learn Bad Habits

A post-training paper shows how interpretability tools can audit preference data, expose unwanted learning signals, and reshape rewards before models absorb bad habits. The method turns interpretability into a practical intervention in data curation and reward design.

Choose your reading rhythm.

Start with a weekly briefing, add daily notes, or hear only when a durable essay or library update is ready.

You've successfully subscribed to Antoine Buteau
Great! Next, complete checkout to get full access to all premium content.
Welcome back! You've successfully signed in.
Unable to sign you in. Please try again.
Success! Your account is fully activated, you now have access to all content.
Error! Stripe checkout failed.
Success! Your billing info is updated.
Error! Billing info update failed.