Latent reasoning
What is latent reasoning?
Latent reasoning is a language model working through a problem in its latent space, the internal numbers it computes with, instead of in written words.
The phrase covers two things. One is the reasoning a model does inside a single step, before it writes a word. The other is a set of newer methods that train a model to carry its whole chain of reasoning in numbers, which is what this site calls neuralese.
- Also called
- latent chain of thought, latent space reasoning, continuous thought, implicit reasoning
01
Reasoning inside a single step
Before a large language model writes a word, its input runs through dozens of layers, and some reasoning happens in there. Take “The mother of the singer of ‘Superstition’ is”. To finish that sentence, a model has to work out that the singer is Stevie Wonder, then recall his mother, without writing either step down.
In 2024, researchers at Google DeepMind, UCL and Tel Aviv University looked inside LLaMA-2 models to see whether that happens. The first hop, finding Stevie Wonder, often did, and more so in bigger models. The second hop was patchier and didn’t improve with size. For some kinds of question the models took the full two-step path in more than 80% of cases. Averaged across the kinds they tested, the evidence was moderate.
A later study from Apollo Research, the UK AI Security Institute and UC Berkeley used invented facts, so the models couldn’t have memorized the answers. When a model learned two invented facts separately, it couldn’t combine them: 0% correct. When one fact was invented and the other was something it already knew, it often could. Taught that Kevin’s favorite programming language is Python, it could say who created Kevin’s favorite language. The authors’ conclusion is that models can do this kind of silent reasoning, and that experiments can easily make them look better or worse at it than they are.
Written-out reasoning matters because of limits like these. A transformer can only pass information from its later layers back to its earlier ones by writing a token, so a long enough chain of reasoning has to go through words. As 41 researchers put it in 2025, “for sufficiently difficult tasks, Transformers must use chain of thought as a form of working memory.” Methods that reason in numbers remove that constraint, and the rest of this page is about them. The window is narrowing without them, too: OpenAI says GPT-6 Astra can do far more without writing any reasoning than earlier models could.
02
How researchers move reasoning into numbers
The simplest approach changes only the training. Yuntian Deng and colleagues start with a model that writes its reasoning out, then delete the written steps a few at a time as training goes on, until the model does them silently. A small GPT-2 model trained this way multiplied two nine-digit numbers with up to 99% accuracy without writing a step, where standard training couldn’t get past four digits. Mistral 7B, trained the same way, got over half of a standard set of grade-school math problems right. An earlier paper by the same lead author called this reasoning “vertically”, through the model’s layers, instead of “horizontally”, through words.
The second approach replaces the written steps themselves. COCONUT, short for Chain of Continuous Thought, comes from Meta and UC San Diego. The model takes its internal state at the end of a step and feeds it straight back in as the next input, without picking a word. These steps are called continuous thoughts. CODI teaches a model the same trick by having it learn from its own written reasoning. Sakana AI’s Continuous Thought Machine, from 2025, shares the name but is a different idea: a network whose neurons each keep their own timing, tested on mazes and image recognition, not as a chatbot.
The third keeps words for most of the reasoning and swaps the rest for new tokens. Token Assorted, from 2025, compresses the first part of a written reasoning trace into learned tokens that stand for chunks of reasoning but aren’t words, and trains the model on a mix of the two.
Recurrent depth changes something else: how much work goes into each step. A recurrent-depth model runs a block of its layers in a loop, so it can think longer before it writes the next word. Huginn-0125, a 3.5-billion-parameter model from Jonas Geiping and colleagues, works this way, and its weights are public. It’s also the technique The Information reported for GPT-6 Astra. LOTUS, from June 2026, combines the loop with continuous thoughts.
Xinghao Chen and colleagues’ survey maps the rest of the papers, sorted by method.
03
Does latent reasoning work?
On small models and narrow tasks, often. Beyond that, there’s little evidence yet.
COCONUT, tested on GPT-2, a small model from 2019, beat written reasoning on logic puzzles that need a lot of searching and backtracking, and used fewer steps to do it. On grade-school math it did worse: 34% correct against 43%. CODI closed that gap at the same scale, matching written reasoning on the GSM8K math set with reasoning 3.1 times shorter. Token Assorted’s mixed traces came out shorter than fully written ones and scored higher on its benchmarks. LOTUS matched written reasoning on its test problems at 3 billion parameters.
The methods are fragile. Adding more continuous steps can make training collapse, because the steps turn into near-copies of each other. SIM-CoT fixed this by supervising each step with a helper decoder that’s dropped after training. Silent reasoning also often leans on shortcuts. In a 2025 study, small models trained from scratch learned to skip the written steps only when problems followed a fixed pattern. Given varied problems, they overfit, and the authors saw the same limit in large current models.
The results also come from models of 8 billion parameters or fewer, far below the frontier. We found no frontier lab saying it uses continuous thoughts in a released model. GPT-6 Astra was reported to use recurrent depth, a different technique, and OpenAI hasn’t confirmed even that. Which AI models think in neuralese? goes model by model.
04
Can anyone read the numbers?
A continuous thought is a long list of numbers with no dictionary. Researchers have found ways to read some of it, and good reasons to doubt the readings.
In theory, one continuous thought can hold several lines of reasoning at once. Hanlin Zhu and colleagues proved in 2025 that on a graph-search problem, continuous thoughts can follow many paths in parallel, where written reasoning has to pick one path at a time. Their models found this strategy in training without being taught it. COCONUT’s authors saw something similar inside their model and compared it to breadth-first search.
Some steps can be turned back into words. SIM-CoT’s helper decoder gives a rough reading of each continuous step, and LOTUS’s hidden states could be read back as reasoning steps.
A 2026 retest urges caution. Researchers at Université Grenoble Alpes, Université Paris-Saclay and NAVER LABS Europe compared COCONUT and CODI with control models built without the feedback loop or the staged training. The patterns that looked like reasoning, including the search frontier and arithmetic that could be decoded, showed up in the controls too, and in some cases had no effect on what the models did. Their conclusion is that latent thoughts are “hidden computation, not hidden explanation”. A step you can decode isn’t proof of what the model used it for.
That matters for safety. A person or a weaker model can read written reasoning, even when it leaves things out. Reading continuous thoughts takes interpretability tools, and so far those give rough readings of small models.
05
Sources
These are the sources behind this page. We checked them against the original text where we could read it.
- Peer-reviewed paper
Sohee Yang et al. Do Large Language Models Latently Perform Multi-Hop Reasoning? ACL 2024.
The two hops inside LLaMA-2.
- Preprint
Mikita Balesni, Tomek Korbak, Owain Evans. Lessons from Studying Two-Hop Latent Reasoning. November 2024, revised November 2025.
Two-hop reasoning over invented facts.
- Preprint
Tomek Korbak et al. Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety. July 2025.
Why long reasoning has to go through words.
- Lab publication
OpenAI. GPT-6 Astra System Card. 3 Sep 2026, updated 9, 22 and 29 Sep 2026.
Astra doing more without written reasoning.
- Preprint
Yuntian Deng, Yejin Choi, Stuart Shieber. From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step. May 2024.
Removing written steps during training.
- Preprint
Yuntian Deng et al. Implicit Chain of Thought Reasoning via Knowledge Distillation. November 2023.
Reasoning “vertically” through the layers.
- Peer-reviewed paper
Shibo Hao et al. Training Large Language Models to Reason in a Continuous Latent Space. COLM 2025.
Continuous thoughts, and how they did.
- Peer-reviewed paper
Zhenyi Shen et al. CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation. EMNLP 2025.
Learning continuous thoughts from written reasoning.
- Peer-reviewed paper
Luke Darlow et al. (Sakana AI). Continuous Thought Machines. NeurIPS 2025.
The Continuous Thought Machine.
- Peer-reviewed paper
DiJia Su et al. Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning. ICML 2025.
Learned tokens that aren’t words.
- Peer-reviewed paper
Jonas Geiping et al. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach. NeurIPS 2025.
Huginn-0125 and recurrent depth.
- Journalism
Amir Efrati, Stephanie Palazzolo, Rocket Drew. OpenAI Technique in ‘Astra’ Model Sparks Security Concerns. The Information, 1 Sep 2026.
The report that Astra uses recurrent depth. Paywalled. We haven’t read the full text, and the details here match how other outlets summarized it.
- Preprint
Ying Fan, Anej Svete, Kangwook Lee. Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers. June 2026.
Continuous thoughts in a looped model.
- Peer-reviewed paper
Xinghao Chen et al. Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning. Findings of EMNLP 2026.
A map of the papers, sorted by method.
- Peer-reviewed paper
Xilin Wei et al. SIM-CoT: Supervised Implicit Chain-of-Thought. ICLR 2026.
Training collapse, and reading each step.
- Peer-reviewed paper
Tianhe Lin, Jian Xie, Siyu Yuan, Deqing Yang. Implicit Reasoning in Transformers is Reasoning through Shortcuts. Findings of ACL 2025.
Shortcuts in silent reasoning.
- Peer-reviewed paper
Hanlin Zhu et al. Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought. NeurIPS 2025.
Several paths in one continuous thought.
- Preprint
Darpan Aswal, Thomas Palmeira Ferraz, Yongxin Zhou, Maxime Peyrard. Observable Patterns Are Not Explanations: A Causal-Geometric Analysis of Latent Reasoning Models. June 2026.
The 2026 retest of COCONUT and CODI.