Lessons from Neel Nanda
Neel Nanda leads a mechanistic interpretability team at Google DeepMind after working at Anthropic. Through grokking, induction heads, and Othello-GPT, he reverse-engineers neural networks to improve how researchers read internals, test safety, and conduct experiments.