This post opens a series on mechanistic interpretability: the attempt to reverse-engineer a trained network into human-understandable circuits. Everything you need to start is free.
The core: Transformer Circuits
Anthropic’s interpretability team publishes its work as the Transformer Circuits thread, free and online.
- Thread home: transformer-circuits.pub
- The foundational piece: A Mathematical Framework for Transformer Circuits (2021), followed by Toy Models of Superposition and Towards Monosemanticity.
Getting hands-on
- Neel Nanda’s getting-started guide, his TransformerLens library, and his “200 Concrete Open Problems in Mechanistic Interpretability.”
- ARENA, a free course with a full mechanistic-interpretability chapter and exercises.
The plan
From here the series works through the thread, paper by paper, with small experiments where they help.