Transformer Circuits
25 stories
Transformer Circuits
Transformer CircuitsResearchCharacterizing interference weights in a tiny language model We identify interference weights in a 1-layer transformer by measuring their effect on model outputs and loss.
Transformer Circuits
Transformer CircuitsResearchVerbalizable Representations Form a Global Workspace in Language Models We find that Claude maintains a small, privileged set of representations it can report on, control, and reason with, atop a much larger volume of automatic processing.
Transformer Circuits
Transformer CircuitsResearchCircuits Updates — June 2026 A short update on turn-averaged sparse autoencoders.
Transformer Circuits
Transformer CircuitsResearchCircuits Updates — May 2026 A short update on understanding features through downstream connections.
Transformer Circuits
Transformer CircuitsResearchNatural Language Autoencoders Produce Unsupervised Explanations of LLM Activations We train Claude to translate its internal state into natural language.
Transformer Circuits
Transformer CircuitsResearchHeadVis We develop an interactive visualization tool to help us understand the behaviors of attention heads in language models.
Transformer Circuits
Transformer CircuitsResearchEmotion Concepts and their Function in a Large Language Model We find representations of emotion concepts in Claude Sonnet 4.5 and show that they causally influence its outputs.
Transformer Circuits
Transformer CircuitsResearchCircuits Cross-Post — Activation Oracles We train language models to answer questions about their own activations in natural language.
Transformer Circuits
Transformer CircuitsResearchCircuits Updates — November 2025 A short update on harm pressure.
Transformer Circuits
Transformer CircuitsResearchEmergent Introspective Awareness in Large Language Models We find evidence that language models can introspect on their internal states.
Transformer Circuits
Transformer CircuitsResearchCircuits Updates — October 2025 Small updates on visual features and dictionary initialization.
Transformer Circuits
Transformer CircuitsResearchWhen Models Manipulate Manifolds: The Geometry of a Counting Task We find geometric structure underlying the mechanisms of a fundamental language model behavior.
Transformer Circuits
Transformer CircuitsResearchCircuits Updates — September 2025 A small update on features and in-context learning.
Transformer Circuits
Transformer CircuitsResearchCircuits Updates — August 2025 A small update: How does a persona modify the assistant’s response?
Transformer Circuits
Transformer CircuitsResearchA Toy Model of Mechanistic (Un)Faithfulness When transcoders go awry.
Transformer Circuits
Transformer CircuitsResearchTracing Attention Computation Through Feature Interactions We describe and apply a method to explain attention patterns in terms of feature interactions, and integrate this information into attribution graphs.
Transformer Circuits
Transformer CircuitsResearchA Toy Model of Interference Weights Unpacking "interference weights" in some more depth.
Transformer Circuits
Transformer CircuitsResearchSparse mixtures of linear transforms We investigate sparse mixture of linear transforms (MOLT), a new approach to transcoders.
Transformer Circuits
Transformer CircuitsResearchCircuits Updates — July 2025 A collection of small updates: revisiting A Mathematical Framework and applications of interpretability to biology.
Transformer Circuits
Transformer CircuitsResearchAutomated Auditing A note on using agents to perform automated alignment audits, including using interpretability tools.
Transformer Circuits
Transformer CircuitsResearchCircuits Updates — April 2025 A collection of small updates: jailbreaks, dense features, and spinning up on interpretability.
Transformer Circuits
Transformer CircuitsResearchProgress on Attention An update on our progress studying attention.
Transformer Circuits
Transformer CircuitsResearchOn the Biology of a Large Language Model We investigate the internal mechanisms used by Claude 3.5 Haiku — Anthropic's lightweight production model — in a variety of contexts.
Transformer Circuits
Transformer CircuitsResearchCircuit Tracing: Revealing Computational Graphs in Language Models We describe an approach to tracing the "step-by-step" computation involved when a model responds to a single prompt.
Nothing matches this filter yet.