Lessons from Chris Olah
Chris Olah is an AI researcher and Anthropic co-founder who popularized mechanistic interpretability. By isolating neural circuits instead of accepting black boxes, he maps how models form concepts and advances the effort to decode large language models.