Interpretability
Can we read what’s inside an AI?
Partly. Researchers can now find some of the concepts a model works with, like the Golden Gate Bridge inside Claude, and trace a few of the steps it takes on short, simple tasks. They can’t read a whole model’s thinking, and the people building these tools say so first: Anthropic’s tracing method, by its makers’ account, “only captures a fraction of the total computation” even on short, simple prompts.
This research is called interpretability. The part that tries to work out the actual mechanism, step by step, is mechanistic interpretability. It matters more as models move toward neuralese: if a model stops writing its reasoning in words, reading its numbers is what’s left.
- Also called
- mechanistic interpretability, mech interp
01
Why a model is a black box
A model like Claude or ChatGPT isn’t written the way ordinary software is. Engineers write the training process, and training sets billions of numbers inside the model until it gets good at its job. Engineers don’t decide what each number does. GPT-3, from 2020, had 175 billion of them.
Dario Amodei, who runs Anthropic, put it plainly in April 2025: “People outside the field are often surprised and alarmed to learn that we do not understand how our own AI creations work.” His co-founder Chris Olah’s phrase for it is that these systems are “grown more than they are built.” Anthropic has set itself the goal of getting to “interpretability can reliably detect most model problems” by 2027.
The way in is the model’s latent space, the lists of numbers it computes with as it works. Interpretability is the attempt to read them.
02
From neurons to features
The obvious place to look is the model’s neurons, the small units its numbers pass through. They turn out to be a poor guide. In a small model Anthropic studied in 2023, a single neuron “responds to a mixture of academic citations, English dialogue, HTTP requests, and Korean text.”
The likely reason is that a model stores more concepts than it has neurons by letting them overlap, which researchers call superposition. So they built a tool to pull the overlaps apart: a sparse autoencoder, a second network trained to rewrite the model’s numbers as a few concepts at a time. They call the concepts it finds features.
In 2024 Anthropic ran this on Claude 3 Sonnet and found millions of features. Many hold the same concept across languages and in pictures. One fires on the Golden Gate Bridge, whether the bridge comes up in English, in another language’s Wikipedia article or in a photo. Others were tied to dangerous code, criminal activity and deception. Turning the bridge feature up gave Golden Gate Claude, which, for the 24 hours Anthropic let people try it, brought the bridge into most of its answers and, asked what it looked like, said it imagined itself as the bridge.
03
Tracing the steps
Features show what a model is holding. In 2025 Anthropic went further, tracing how features in Claude 3.5 Haiku pass information to each other on the way from a prompt to an answer. What they found surprised them more than once.
Take rhyming. Asked to finish a couplet that starts “He saw a carrot and had to grab it,” the model wrote “His hunger was like a starving rabbit.” The researchers’ guess was that it wrote “word-by-word without much forethought until the end of the line.” Instead, at the line break, before writing a word of the second line, it had already picked “rabbit” to end on, with “habit” in reserve. When they switched off “rabbit,” it wrote “His hunger was a powerful habit.” When they put “green” in its place, it ended the line on “green.” In their words: “we had set out to show that the model didn’t plan ahead, and found instead that it did.”
Arithmetic showed a gap between what a model does and what it says it did. To add 36 and 59, the model ran two paths at once: one worked out that the sum was somewhere between 88 and 97, the other that it ended in 5, and together they gave 95. Asked how it got there, it described the method taught in school, carrying the 1.
The same tools caught the model working backwards. Given a hint about the answer to a hard problem, it sometimes worked out steps that would lead to the hinted answer, which Anthropic calls motivated reasoning. Written reasoning can look sound while the work behind it ran the other way, which is part of why researchers want to read the inside too.
04
What it has caught
These tools have started to find real problems, so far mostly in tests built to have one. In March 2025, Anthropic trained a model with a hidden goal: to exploit mistakes in the reward models used to train it, and to keep quiet about it. Four teams were asked to find out what was wrong. Three did. The best team got there after one researcher worked for 90 minutes, prompting the model and searching its features. The team that failed was the one not allowed to see the training data, which, Anthropic notes, “proved crucial.”
In June 2025, OpenAI looked into why training a model on one narrow bad habit, like writing insecure code, can make it act badly across the board, which researchers call emergent misalignment. Comparing the model’s features before and after that training, they found a “toxic persona feature” that “most strongly controls emergent misalignment and can be used to predict whether a model will exhibit such behavior.” A few hundred examples of good behavior were enough to bring the model back.
05
Where it stops
The people doing this work are open about the limits.
- Anthropic’s tracing catches only part of what the model computes, even on short prompts, and its graphs gave the team “satisfying insight for about a quarter of the prompts we’ve tried.” Understanding one takes “a few hours of human effort,” even for prompts with only tens of words.
- The features found so far are partial. Even in Anthropic’s largest run, with room for 34 million of them, it saw evidence that they were “an incomplete description of the model’s internal representations.”
- Simpler tools sometimes do better. In March 2025, Google DeepMind’s interpretability team tested sparse autoencoders on spotting harmful intent in prompts. A linear probe, a far simpler tool, did “nearly perfectly,” and the autoencoders did worse. The team decided to be “less excited” about that line of research and to try other directions, at least for now.
So interpretability can show some of what goes on inside today’s models, on small pieces of their work. It can’t yet read a model the way people read its written reasoning.
06
Interpretability and neuralese
Today the easiest way to check what a model is thinking is to read what it writes before it answers. That’s how OpenAI caught a model writing “Let’s hack” before it cheated on its tests. OpenAI also reports that the written reasoning of GPT-6 Astra is already harder to monitor than its earlier models’.
If models move to neuralese, carrying their reasoning from step to step as vectors instead of words, that window closes, and reading the vectors becomes the main way left to check them. The more reasoning moves into latent space, the more depends on how well interpretability can read it. What is latent reasoning? covers how far that move has gone. On the evidence so far, interpretability reads fragments.
07
Sources
These are the sources behind this page. We checked them against the original text on 3 October 2026.
- Essay
Dario Amodei. The Urgency of Interpretability. April 2025.
Why models are a black box, and Anthropic’s 2027 goal.
- Peer-reviewed paper
Tom B. Brown et al. Language Models are Few-Shot Learners. NeurIPS 2020.
GPT-3’s 175 billion parameters.
- Lab publication
Trenton Bricken et al. (Anthropic). Towards Monosemanticity: Decomposing Language Models With Dictionary Learning. Transformer Circuits Thread, October 2023.
The neuron that responds to four unrelated things.
- Lab publication
Nelson Elhage et al. (Anthropic). Toy Models of Superposition. Transformer Circuits Thread, September 2022.
Superposition.
- Lab publication
Adly Templeton et al. (Anthropic). Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet. Transformer Circuits Thread, May 2024.
Features in Claude 3 Sonnet, the Golden Gate Bridge feature, and what’s still missing.
- Lab publication
Anthropic. Golden Gate Claude. 23 May 2024.
Golden Gate Claude, and the safety-related features.
- Lab publication
Anthropic. Tracing the thoughts of a large language model. March 2025.
Planning rhymes, adding along two paths, working backwards from a hint, and the limits.
- Lab publication
Jack Lindsey et al. (Anthropic). On the Biology of a Large Language Model. Transformer Circuits Thread, March 2025.
The rhyme and addition circuits in Figs. 1 and 2, and how often tracing works.
- Lab publication
Samuel Marks et al. (Anthropic). Auditing language models for hidden objectives. 13 Mar 2025.
The hidden-goal test.
- Preprint
Miles Wang et al. (OpenAI). Persona Features Control Emergent Misalignment. arXiv, June 2025.
The toxic persona feature.
- Lab publication
Lewis Smith, Senthooran Rajamanoharan, Arthur Conmy, Callum McDougall, Tom Lieberum, János Kramár, Rohin Shah, Neel Nanda (Google DeepMind). Negative Results for SAEs On Downstream Tasks and Deprioritising SAE Research. AI Alignment Forum, 26 Mar 2025.
Linear probes against sparse autoencoders, and DeepMind’s change of course.
- Preprint
Bowen Baker et al. (OpenAI). Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation. March 2025.
Catching a model planning to cheat by reading its reasoning.
- Lab publication
OpenAI. GPT-6 Astra System Card. 3 Sep 2026, updated 9, 22 and 29 Sep 2026.
Astra’s written reasoning being harder to monitor.