This post opens a series on mechanistic interpretability: the attempt to reverse-engineer a trained network into human-understandable circuits. Everything you need to start is free.

The core: Transformer Circuits

Anthropic’s interpretability team publishes its work as the Transformer Circuits thread, free and online.

Getting hands-on

  • Neel Nanda’s getting-started guide, his TransformerLens library, and his “200 Concrete Open Problems in Mechanistic Interpretability.”
  • ARENA, a free course with a full mechanistic-interpretability chapter and exercises.

The plan

From here the series works through the thread, paper by paper, with small experiments where they help.