Following Anthropic's Transformer Circuits thread.
Mechanistic Interpretability: Free Materials to Start
Reverse-engineering neural networks, circuit by circuit. Anthropic’s Transformer Circuits thread and a few companion resources are all free. Where to start, to open this series.
Transformer Circuits 1: Summary of Results
Reading the Framework paper, part 1. Treating tiny attention-only models as model organisms for fully reverse engineering transformers, and a map of the key results.
Transformer Circuits 2: The Transformer, Redrawn
Reading the Framework paper, part 2. Redrawing the residual stream as a communication channel and attention heads as independent information movers, through virtual weights and the QK/OV split.
Transformer Circuits 3: Zero-Layer Transformers
Reading the Framework paper, part 3. What does a transformer with no attention at all learn? Bigram statistics, and that answer turns out to be the key to the direct path in large models.
Transformer Circuits 4: One-Layer Models and Skip-Trigrams
Reading the Framework paper, part 4. Fully expanding a one-attention-layer model into paths: skip-trigrams, copying heads and positive eigenvalues, and the characteristic bugs the factored structure produces.
Transformer Circuits 5: Two-Layer Models and Induction Heads
Reading the Framework paper, part 5. With two layers, heads compose with heads: Q-, K-, and V-composition, and the paper’s biggest discovery, the mechanism of the induction head.
Transformer Circuits 6: Where Does This Leave Us?
Reading the Framework paper, part 6. What the framework secured and what it did not: the wall called MLP, and the follow-up research this paper announces.
Transformer Circuits 7: The Related-Work Landscape
Reading the Framework paper, final part. Whose shoulders this paper stands on, and how its approach differs from the lines of work before it.