Neuralese Wiki

Last updated

Recurrent depth

What is recurrent depth?

Recurrent depth lets a language model think longer by running part of itself in a loop. Before it writes each word, a block of its layers can go over the same numbers several times, putting more work into that one step.

The loop runs in numbers, not words, so the extra thinking doesn’t show up in a written chain of thought. That’s why it comes up in talk about neuralese, and why a report that GPT-6 Astra uses it drew so much attention.

Also called
looped transformers, looped language models, recursive reasoning
Last updated

01

How recurrent depth works

An ordinary language model sends each word through its stack of layers once, then writes the next word. To think harder, reasoning models write more words first. A recurrent-depth model has another option: it sends the same numbers through some of its layers again.

One step through Huginn-0125The words so far pass through a prelude of two layers, then a block of four layers that runs in a loop, 32 times on average in training, then a coda of two layers that produces the next word. That is 2 + 4 × 32 + 2 = 132 layers of work for each word.the words so farnext wordPRELUDE2 LAYERSLOOP4 LAYERSCODA2 LAYERS×322 + 4 × 32 + 2 = 132 LAYERS FOR EACH WORD
Fig. 1 One step in Huginn-0125. Two layers turn the words so far into numbers, four layers loop over them, and two more turn the result into the next word. The model was trained to loop 32 times on average, and at test time it can run more loops or fewer.

The clearest example is Huginn-0125, a 3.5-billion-parameter model that Jonas Geiping and colleagues released in February 2025, weights included. It has eight real layers: two at the start that turn words into numbers, four in the middle that loop, and two at the end that turn the numbers back into a word. The authors call them the prelude, the recurrent block and the coda, and say looping lets them “put an indefinite amount of verses in our song.” The name comes from Huginn, a raven in Norse mythology whose name means “thought.”

In training, the number of loops changed from one step to the next, 32 on average, so the model learned to work with however many it gets. At 32 loops, its eight layers do the work of 132. Turning the loops up at test time pays off: on grade-school science questions, the easy half of the AI2 Reasoning Challenge, the same model scored 49% with 4 loops and 70% with 32.

It can also decide for itself. The model stops looping on a word once another pass would barely change its prediction, so some questions get fewer loops than others. In the authors’ tests it settled sooner on high-school math than on questions about logical fallacies or moral scenarios.

The loops leave traces you can plot. On many words the model’s state settles quickly. When a prompt needs numerical reasoning, the state often falls into an orbit instead, circling through the same region. The authors saw orbits, and states that drift steadily in one direction, used for things like arithmetic.

02

Where the idea came from

The idea is older than chatbots. In 2016, Alex Graves let a recurrent network decide how many internal steps to take on each input, so harder inputs got more computation. Universal Transformers, in 2018, applied the same transformer block over and over and let each position stop when it was done.

In 2023, researchers showed how far a loop can go in principle. With its weights set by hand, a transformer of 13 layers, run in a loop, worked as a small programmable computer: it ran a calculator, a linear algebra library and a learning algorithm.

Huginn took the idea to the scale of a language model in 2025. It was trained on 800 billion tokens, the word pieces a model reads, using 4,096 GPUs on the Frontier supercomputer at Oak Ridge National Laboratory.

That June, Sapient, a lab in Singapore, published the Hierarchical Reasoning Model, with two loops running at different speeds: a slow one for planning and a fast one for detail, loosely modeled on the brain. With 27 million parameters and about 1,000 training examples, it solved hard Sudoku puzzles and large mazes almost perfectly, and beat much bigger models on ARC-AGI, a test of abstract visual puzzles.

In October, Alexia Jolicoeur-Martineau at Samsung’s AI lab in Montréal cut it down to one network of two layers and 7 million parameters. Her Tiny Recursive Model reported 45% on ARC-AGI-1 and 8% on the harder ARC-AGI-2, higher than DeepSeek R1, o3-mini and Gemini 2.5 Pro with less than 0.01% of their parameters.

The same month, ByteDance Seed and university partners released Ouro, named after the ouroboros, the snake that eats its own tail. Its looped language models of 1.4 and 2.6 billion parameters, trained on 7.7 trillion tokens, match models of up to 12 billion, the authors report. LOTUS, in June 2026, runs COCONUT-style continuous thoughts through a looped transformer, joining recurrent depth with the other big idea in latent reasoning.

03

Does it work?

On puzzles and at small scale, often, and sometimes strikingly. Beyond that, there’s little evidence yet.

Huginn’s extra loops paid off most where there was something to work out. Easy question sets settled after a few loops, while grade-school math kept improving as loops were added. Its scores kept rising up to the computation a 50-billion-parameter model would spend.

The puzzle results need a closer look. When the ARC Prize team retested the Hierarchical Reasoning Model on hidden tasks in August 2025, it scored 32% on ARC-AGI-1 and 2% on ARC-AGI-2. In their experiments, the two-speed, brain-inspired design mattered little: an ordinary transformer of the same size came within about 5 points. What drove the score was an outer loop that let the model take more passes at its own answer, and, in their words, “the large majority of the performance is driven by training on the tasks seen at evaluation time.”

At the frontier, the evidence comes down to one report. The Information wrote on 1 September 2026 that GPT-6 Astra loops some of its layers, and OpenAI hasn’t confirmed it. We found no frontier lab saying it ships a looped model.

  1. Confirmed

    Huginn-0125 thinks longer by running its layers in a loop.

    Geiping et al. 2025

  2. Confirmed

    LOTUS runs continuous thoughts through a looped model.

    LOTUS paper

  3. Reported

    GPT-6 Astra loops some of its layers, a technique called recurrent depth.

    The Information
    Astra system card

Which AI models think in neuralese? goes model by model.

04

Loops and reading AI’s reasoning

The loop’s work doesn’t turn into words. A model that can think longer inside each step has less need to write its reasoning out, and reasoning it keeps in the loop can’t be read the way a chain of thought can.

That worries the researchers who study chain-of-thought monitoring. In a 2025 position paper, 41 of them, from labs including OpenAI, Anthropic and Google DeepMind, point to architectures like Huginn’s and warn that such models “might not need to verbalize any of their thoughts and would thus lose the safety advantages that CoT confers.” For GPT-6 Astra, OpenAI says it is “quite confident” that the change in how well Astra controls its reasoning isn’t down to architecture.

Loops may be easier to read than they sound. Ouro’s authors report that their looped models’ reasoning traces line up with the final answer more closely than written chain of thought does, and LOTUS’s hidden states could be read back as reasoning steps. Whether that holds at the scale of a frontier model, researchers can’t say yet.

In our view, recurrent depth is the form of neuralese to watch: it needs no special training data, it has worked at a few billion parameters, and it’s the technique a frontier model has been reported to use.

05

Sources

These are the sources behind this page. We checked each claim against the original.

  1. Peer-reviewed paper

    Huginn-0125: how it loops and what the loops bought.

  2. Preprint

    Adaptive computation, 2016.

  3. Peer-reviewed paper

    Mostafa Dehghani et al. Universal Transformers. ICLR 2019.

    Universal Transformers.

  4. Peer-reviewed paper

    Angeliki Giannou et al. Looped Transformers as Programmable Computers. ICML 2023.

    A loop as a programmable computer.

  5. Preprint

    Guan Wang et al. (Sapient). Hierarchical Reasoning Model. June 2025.

    The Hierarchical Reasoning Model.

  6. Lab publication

    The ARC Prize retest of HRM.

  7. Preprint

    Alexia Jolicoeur-Martineau (Samsung SAIL Montréal). Less is More: Recursive Reasoning with Tiny Networks. October 2025.

    The Tiny Recursive Model.

  8. Preprint

    Rui-Jie Zhu et al. (ByteDance Seed and others). Scaling Latent Reasoning via Looped Language Models. October 2025, revised July 2026.

    Ouro’s looped language models.

  9. Preprint

    Continuous thoughts in a loop.

  10. Preprint

    The warning about architectures that reason in numbers.

  11. Journalism

    Amir Efrati, Stephanie Palazzolo, Rocket Drew. OpenAI Technique in ‘Astra’ Model Sparks Security Concerns. The Information, 1 Sep 2026.

    The report that Astra uses recurrent depth. Paywalled. We haven’t read the full text, and the details here match how other outlets summarized it.

  12. Lab publication

    OpenAI. GPT-6 Astra System Card. 3 Sep 2026, updated 9, 22 and 29 Sep 2026.

    OpenAI on Astra’s architecture and reasoning.