The letter

One letter a week, in your inbox.

The signal of the week, what shipped, what to try, and the editor's note. No tracking, no ads, nothing else.

We keep your address, your language and the date you joined, nothing else. Every letter has a one-click unsubscribe link that deletes the record.

Transformer Circuits

25 stories

Transformer Circuits

Transformer CircuitsResearchCharacterizing interference weights in a tiny language model We identify interference weights in a 1-layer transformer by measuring their effect on model outputs and loss.

August 21
Transformer Circuits

Transformer CircuitsResearchVerbalizable Representations Form a Global Workspace in Language Models We find that Claude maintains a small, privileged set of representations it can report on, control, and reason with, atop a much larger volume of automatic processing.

July 6
Transformer Circuits

Transformer CircuitsResearchCircuits Updates — June 2026 A short update on turn-averaged sparse autoencoders.

June 30
Transformer Circuits

Transformer CircuitsResearchCircuits Updates — May 2026 A short update on understanding features through downstream connections.

June 1
Transformer Circuits

Transformer CircuitsResearchNatural Language Autoencoders Produce Unsupervised Explanations of LLM Activations We train Claude to translate its internal state into natural language.

May 7
Transformer Circuits

Transformer CircuitsResearchHeadVis We develop an interactive visualization tool to help us understand the behaviors of attention heads in language models.

May 4
Transformer Circuits

Transformer CircuitsResearchEmotion Concepts and their Function in a Large Language Model We find representations of emotion concepts in Claude Sonnet 4.5 and show that they causally influence its outputs.

April 2
Transformer Circuits

Transformer CircuitsResearchCircuits Cross-Post — Activation Oracles We train language models to answer questions about their own activations in natural language.

December 19
Transformer Circuits

Transformer CircuitsResearchCircuits Updates — November 2025 A short update on harm pressure.

November 26
Transformer Circuits

Transformer CircuitsResearchEmergent Introspective Awareness in Large Language Models We find evidence that language models can introspect on their internal states.

October 29
Transformer Circuits

Transformer CircuitsResearchCircuits Updates — October 2025 Small updates on visual features and dictionary initialization.

October 24
Transformer Circuits

Transformer CircuitsResearchWhen Models Manipulate Manifolds: The Geometry of a Counting Task We find geometric structure underlying the mechanisms of a fundamental language model behavior.

October 21
Transformer Circuits

Transformer CircuitsResearchCircuits Updates — September 2025 A small update on features and in-context learning.

September 29
Transformer Circuits

Transformer CircuitsResearchCircuits Updates — August 2025 A small update: How does a persona modify the assistant’s response?

August 28
Transformer Circuits

Transformer CircuitsResearchA Toy Model of Mechanistic (Un)Faithfulness When transcoders go awry.

August 7
Transformer Circuits

Transformer CircuitsResearchTracing Attention Computation Through Feature Interactions We describe and apply a method to explain attention patterns in terms of feature interactions, and integrate this information into attribution graphs.

July 31
Transformer Circuits

Transformer CircuitsResearchA Toy Model of Interference Weights Unpacking "interference weights" in some more depth.

July 29
Transformer Circuits

Transformer CircuitsResearchSparse mixtures of linear transforms We investigate sparse mixture of linear transforms (MOLT), a new approach to transcoders.

July 25
Transformer Circuits

Transformer CircuitsResearchCircuits Updates — July 2025 A collection of small updates: revisiting A Mathematical Framework and applications of interpretability to biology.

July 25
Transformer Circuits

Transformer CircuitsResearchAutomated Auditing A note on using agents to perform automated alignment audits, including using interpretability tools.

July 24
Transformer Circuits

Transformer CircuitsResearchCircuits Updates — April 2025 A collection of small updates: jailbreaks, dense features, and spinning up on interpretability.

April 29
Transformer Circuits

Transformer CircuitsResearchProgress on Attention An update on our progress studying attention.

April 28
Transformer Circuits

Transformer CircuitsResearchOn the Biology of a Large Language Model We investigate the internal mechanisms used by Claude 3.5 Haiku — Anthropic's lightweight production model — in a variety of contexts.

March 27
Transformer Circuits

Transformer CircuitsResearchCircuit Tracing: Revealing Computational Graphs in Language Models We describe an approach to tracing the "step-by-step" computation involved when a model responds to a single prompt.

March 27