Transformer Circuits 7: The Related-Work Landscape

Reading the Framework paper, final part. Whose shoulders this paper stands on, and how its approach differs from the lines of work before it.

August 9, 2025 · 3 min · rick

Transformer Circuits 6: Where Does This Leave Us?

Reading the Framework paper, part 6. What the framework secured and what it did not: the wall called MLP, and the follow-up research this paper announces.

July 26, 2025 · 3 min · rick

Transformer Circuits 5: Two-Layer Models and Induction Heads

Reading the Framework paper, part 5. With two layers, heads compose with heads: Q-, K-, and V-composition, and the paper’s biggest discovery, the mechanism of the induction head.

July 12, 2025 · 7 min · rick

Transformer Circuits 4: One-Layer Models and Skip-Trigrams

Reading the Framework paper, part 4. Fully expanding a one-attention-layer model into paths: skip-trigrams, copying heads and positive eigenvalues, and the characteristic bugs the factored structure produces.

June 28, 2025 · 7 min · rick

Transformer Circuits 3: Zero-Layer Transformers

Reading the Framework paper, part 3. What does a transformer with no attention at all learn? Bigram statistics, and that answer turns out to be the key to the direct path in large models.

June 14, 2025 · 3 min · rick

Transformer Circuits 2: The Transformer, Redrawn

Reading the Framework paper, part 2. Redrawing the residual stream as a communication channel and attention heads as independent information movers, through virtual weights and the QK/OV split.

May 31, 2025 · 9 min · rick

Transformer Circuits 1: Summary of Results

Reading the Framework paper, part 1. Treating tiny attention-only models as model organisms for fully reverse engineering transformers, and a map of the key results.

May 17, 2025 · 3 min · rick

Mechanistic Interpretability: Free Materials to Start

Reverse-engineering neural networks, circuit by circuit. Anthropic’s Transformer Circuits thread and a few companion resources are all free. Where to start, to open this series.

May 10, 2025 · 1 min · rick

Steering GPT-2's Emotions with Sparse Autoencoders

Finding emotion-related features in GPT-2 using OpenAI’s pretrained SAE, then training one from scratch. Feature patching turns ‘good person’ into ‘shit’.

February 16, 2025 · 6 min · rick