Antoine Buteau

Page 153 of 506 · Back to latest writing

Lessons from Neel Nanda

Neel Nanda leads a mechanistic interpretability team at Google DeepMind after working at Anthropic. Through grokking, induction heads, and Othello-GPT, he reverse-engineers neural networks to improve how researchers read internals, test safety, and conduct experiments.

Lessons from Robert Miles

Robert Miles is an AI safety communicator who makes dense alignment theory accessible through thought experiments. Explaining instrumental convergence and specification gaming, he shows why building safe, goal-directed AI remains a difficult technical challenge.

Lessons from Carl Shulman

Carl Shulman models humanity’s future through economics, evolutionary biology, and decision theory. His compute-centric forecasts examine how AGI could automate research and reshape the economy, linking AI development with cognitive enhancement and moral philosophy.

Lessons from Victoria Krakovna

Victoria Krakovna is a Google DeepMind research scientist focused on AI alignment. Her work on specification gaming and penalties for unintended side effects asks how systems can follow human intent without optimizing toward hidden goals of their own.

Lessons from Jan Leike

Jan Leike is an AI safety researcher known for work on RLHF and scalable oversight across DeepMind, OpenAI, and Anthropic. He argues for using AI to automate alignment research while prioritizing safety culture as models approach superintelligence.

Lessons from Paul Christiano

Paul Christiano co-developed Reinforcement Learning from Human Feedback and founded the Alignment Research Center. His work asks how advanced AI can follow human intent instead of optimizing the wrong goals, using scalable oversight and rigorous model evaluation.

Lessons from Stuart Russell

Stuart Russell co-authored the artificial intelligence textbook after decades studying machine decisions. He argues that fixed-objective maximization invites failure, proposing that machines remain uncertain about human preferences so they can coexist with us safely.

Lessons from Nick Bostrom

Nick Bostrom is a philosopher studying existential risk and humanity’s future, known for the simulation argument and vulnerable world hypothesis. His frameworks ask how civilization can survive technological development and the shift to advanced artificial intelligence.

Choose your reading rhythm.

Start with a weekly briefing, add daily notes, or hear only when a durable essay or research update is ready.

You've successfully subscribed to Antoine Buteau
You've successfully subscribed to Antoine Buteau
Welcome back! You've successfully signed in.