Recurrent depth
What is recurrent depth?
Recurrent depth lets a language model think longer by running part of itself in a loop. Before it writes each word, a block of its layers can go over the same numbers several times, putting more work into that one step.
The loop runs in numbers, not words, so the extra thinking doesn’t show up in a written chain of thought. That’s why it comes up in talk about neuralese, and why a report that GPT-6 Astra uses it drew so much attention.
- Also called
- looped transformers, looped language models, recursive reasoning
01
How recurrent depth works
An ordinary language model sends each word through its stack of layers once, then writes the next word. To think harder, reasoning models write more words first. A recurrent-depth model has another option: it sends the same numbers through some of its layers again.
The clearest example is Huginn-0125, a 3.5-billion-parameter model that Jonas Geiping and colleagues released in February 2025, weights included. It has eight real layers: two at the start that turn words into numbers, four in the middle that loop, and two at the end that turn the numbers back into a word. The authors call them the prelude, the recurrent block and the coda, and say looping lets them “put an indefinite amount of verses in our song.” The name comes from Huginn, a raven in Norse mythology whose name means “thought.”
In training, the number of loops changed from one step to the next, 32 on average, so the model learned to work with however many it gets. At 32 loops, its eight layers do the work of 132. Turning the loops up at test time pays off: on grade-school science questions, the easy half of the AI2 Reasoning Challenge, the same model scored 49% with 4 loops and 70% with 32.
It can also decide for itself. The model stops looping on a word once another pass would barely change its prediction, so some questions get fewer loops than others. In the authors’ tests it settled sooner on high-school math than on questions about logical fallacies or moral scenarios.
The loops leave traces you can plot. On many words the model’s state settles quickly. When a prompt needs numerical reasoning, the state often falls into an orbit instead, circling through the same region. The authors saw orbits, and states that drift steadily in one direction, used for things like arithmetic.
02
Where the idea came from
The idea is older than chatbots. In 2016, Alex Graves let a recurrent network decide how many internal steps to take on each input, so harder inputs got more computation. Universal Transformers, in 2018, applied the same transformer block over and over and let each position stop when it was done.
In 2023, researchers showed how far a loop can go in principle. With its weights set by hand, a transformer of 13 layers, run in a loop, worked as a small programmable computer: it ran a calculator, a linear algebra library and a learning algorithm.
Huginn took the idea to the scale of a language model in 2025. It was trained on 800 billion tokens, the word pieces a model reads, using 4,096 GPUs on the Frontier supercomputer at Oak Ridge National Laboratory.
That June, Sapient, a lab in Singapore, published the Hierarchical Reasoning Model, with two loops running at different speeds: a slow one for planning and a fast one for detail, loosely modeled on the brain. With 27 million parameters and about 1,000 training examples, it solved hard Sudoku puzzles and large mazes almost perfectly, and beat much bigger models on ARC-AGI, a test of abstract visual puzzles.
In October, Alexia Jolicoeur-Martineau at Samsung’s AI lab in Montréal cut it down to one network of two layers and 7 million parameters. Her Tiny Recursive Model reported 45% on ARC-AGI-1 and 8% on the harder ARC-AGI-2, higher than DeepSeek R1, o3-mini and Gemini 2.5 Pro with less than 0.01% of their parameters.
The same month, ByteDance Seed and university partners released Ouro, named after the ouroboros, the snake that eats its own tail. Its looped language models of 1.4 and 2.6 billion parameters, trained on 7.7 trillion tokens, match models of up to 12 billion, the authors report. LOTUS, in June 2026, runs COCONUT-style continuous thoughts through a looped transformer, joining recurrent depth with the other big idea in latent reasoning.
03
Does it work?
On puzzles and at small scale, often, and sometimes strikingly. Beyond that, there’s little evidence yet.
Huginn’s extra loops paid off most where there was something to work out. Easy question sets settled after a few loops, while grade-school math kept improving as loops were added. Its scores kept rising up to the computation a 50-billion-parameter model would spend.
The puzzle results need a closer look. When the ARC Prize team retested the Hierarchical Reasoning Model on hidden tasks in August 2025, it scored 32% on ARC-AGI-1 and 2% on ARC-AGI-2. In their experiments, the two-speed, brain-inspired design mattered little: an ordinary transformer of the same size came within about 5 points. What drove the score was an outer loop that let the model take more passes at its own answer, and, in their words, “the large majority of the performance is driven by training on the tasks seen at evaluation time.”
At the frontier, the evidence comes down to one report. The Information wrote on 1 September 2026 that GPT-6 Astra loops some of its layers, and OpenAI hasn’t confirmed it. We found no frontier lab saying it ships a looped model.
- Confirmed
Huginn-0125 thinks longer by running its layers in a loop.
- Confirmed
LOTUS runs continuous thoughts through a looped model.
- Reported
GPT-6 Astra loops some of its layers, a technique called recurrent depth.
Which AI models think in neuralese? goes model by model.
04
Loops and reading AI’s reasoning
The loop’s work doesn’t turn into words. A model that can think longer inside each step has less need to write its reasoning out, and reasoning it keeps in the loop can’t be read the way a chain of thought can.
That worries the researchers who study chain-of-thought monitoring. In a 2025 position paper, 41 of them, from labs including OpenAI, Anthropic and Google DeepMind, point to architectures like Huginn’s and warn that such models “might not need to verbalize any of their thoughts and would thus lose the safety advantages that CoT confers.” For GPT-6 Astra, OpenAI says it is “quite confident” that the change in how well Astra controls its reasoning isn’t down to architecture.
Loops may be easier to read than they sound. Ouro’s authors report that their looped models’ reasoning traces line up with the final answer more closely than written chain of thought does, and LOTUS’s hidden states could be read back as reasoning steps. Whether that holds at the scale of a frontier model, researchers can’t say yet.
In our view, recurrent depth is the form of neuralese to watch: it needs no special training data, it has worked at a few billion parameters, and it’s the technique a frontier model has been reported to use.
05
Sources
These are the sources behind this page. We checked each claim against the original.
- Peer-reviewed paper
Jonas Geiping et al. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach. NeurIPS 2025.
Huginn-0125: how it loops and what the loops bought.
- Preprint
Alex Graves. Adaptive Computation Time for Recurrent Neural Networks. March 2016.
Adaptive computation, 2016.
- Peer-reviewed paper
Mostafa Dehghani et al. Universal Transformers. ICLR 2019.
Universal Transformers.
- Peer-reviewed paper
Angeliki Giannou et al. Looped Transformers as Programmable Computers. ICML 2023.
A loop as a programmable computer.
- Preprint
Guan Wang et al. (Sapient). Hierarchical Reasoning Model. June 2025.
The Hierarchical Reasoning Model.
- Lab publication
ARC Prize Team. The Hidden Drivers of HRM’s Performance on ARC-AGI. 15 Aug 2025.
The ARC Prize retest of HRM.
- Preprint
Alexia Jolicoeur-Martineau (Samsung SAIL Montréal). Less is More: Recursive Reasoning with Tiny Networks. October 2025.
The Tiny Recursive Model.
- Preprint
Rui-Jie Zhu et al. (ByteDance Seed and others). Scaling Latent Reasoning via Looped Language Models. October 2025, revised July 2026.
Ouro’s looped language models.
- Preprint
Ying Fan, Anej Svete, Kangwook Lee. Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers. June 2026.
Continuous thoughts in a loop.
- Preprint
Tomek Korbak et al. Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety. July 2025.
The warning about architectures that reason in numbers.
- Journalism
Amir Efrati, Stephanie Palazzolo, Rocket Drew. OpenAI Technique in ‘Astra’ Model Sparks Security Concerns. The Information, 1 Sep 2026.
The report that Astra uses recurrent depth. Paywalled. We haven’t read the full text, and the details here match how other outlets summarized it.
- Lab publication
OpenAI. GPT-6 Astra System Card. 3 Sep 2026, updated 9, 22 and 29 Sep 2026.
OpenAI on Astra’s architecture and reasoning.