<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Transformer Circuits on 3rd layer</title><link>https://3rdlayer.uk/tags/transformer-circuits/</link><description>Recent content in Transformer Circuits on 3rd layer</description><generator>Hugo -- 0.157.0</generator><language>en-US</language><lastBuildDate>Sat, 09 Aug 2025 00:00:00 +0000</lastBuildDate><atom:link href="https://3rdlayer.uk/tags/transformer-circuits/index.xml" rel="self" type="application/rss+xml"/><item><title>Transformer Circuits 7: The Related-Work Landscape</title><link>https://3rdlayer.uk/posts/framework-07-related-work/</link><pubDate>Sat, 09 Aug 2025 00:00:00 +0000</pubDate><guid>https://3rdlayer.uk/posts/framework-07-related-work/</guid><description>Reading the Framework paper, final part. Whose shoulders this paper stands on, and how its approach differs from the lines of work before it.</description></item><item><title>Transformer Circuits 6: Where Does This Leave Us?</title><link>https://3rdlayer.uk/posts/framework-06-where-does-this-leave-us/</link><pubDate>Sat, 26 Jul 2025 00:00:00 +0000</pubDate><guid>https://3rdlayer.uk/posts/framework-06-where-does-this-leave-us/</guid><description>Reading the Framework paper, part 6. What the framework secured and what it did not: the wall called MLP, and the follow-up research this paper announces.</description></item><item><title>Transformer Circuits 5: Two-Layer Models and Induction Heads</title><link>https://3rdlayer.uk/posts/framework-05-two-layer/</link><pubDate>Sat, 12 Jul 2025 00:00:00 +0000</pubDate><guid>https://3rdlayer.uk/posts/framework-05-two-layer/</guid><description>Reading the Framework paper, part 5. With two layers, heads compose with heads: Q-, K-, and V-composition, and the paper&amp;rsquo;s biggest discovery, the mechanism of the induction head.</description></item><item><title>Transformer Circuits 4: One-Layer Models and Skip-Trigrams</title><link>https://3rdlayer.uk/posts/framework-04-one-layer/</link><pubDate>Sat, 28 Jun 2025 00:00:00 +0000</pubDate><guid>https://3rdlayer.uk/posts/framework-04-one-layer/</guid><description>Reading the Framework paper, part 4. Fully expanding a one-attention-layer model into paths: skip-trigrams, copying heads and positive eigenvalues, and the characteristic bugs the factored structure produces.</description></item><item><title>Transformer Circuits 3: Zero-Layer Transformers</title><link>https://3rdlayer.uk/posts/framework-03-zero-layer/</link><pubDate>Sat, 14 Jun 2025 00:00:00 +0000</pubDate><guid>https://3rdlayer.uk/posts/framework-03-zero-layer/</guid><description>Reading the Framework paper, part 3. What does a transformer with no attention at all learn? Bigram statistics, and that answer turns out to be the key to the direct path in large models.</description></item><item><title>Transformer Circuits 2: The Transformer, Redrawn</title><link>https://3rdlayer.uk/posts/framework-02-transformer-overview/</link><pubDate>Sat, 31 May 2025 00:00:00 +0000</pubDate><guid>https://3rdlayer.uk/posts/framework-02-transformer-overview/</guid><description>Reading the Framework paper, part 2. Redrawing the residual stream as a communication channel and attention heads as independent information movers, through virtual weights and the QK/OV split.</description></item><item><title>Transformer Circuits 1: Summary of Results</title><link>https://3rdlayer.uk/posts/framework-01-summary/</link><pubDate>Sat, 17 May 2025 00:00:00 +0000</pubDate><guid>https://3rdlayer.uk/posts/framework-01-summary/</guid><description>Reading the Framework paper, part 1. Treating tiny attention-only models as model organisms for fully reverse engineering transformers, and a map of the key results.</description></item><item><title>Mechanistic Interpretability: Free Materials to Start</title><link>https://3rdlayer.uk/posts/mech-interp-resources/</link><pubDate>Sat, 10 May 2025 00:00:00 +0000</pubDate><guid>https://3rdlayer.uk/posts/mech-interp-resources/</guid><description>Reverse-engineering neural networks, circuit by circuit. Anthropic&amp;rsquo;s Transformer Circuits thread and a few companion resources are all free. Where to start, to open this series.</description></item></channel></rss>