Neuralese Wiki

Last updated

Reasoning models

How do reasoning models work?

A reasoning model is a language model that writes out its thinking before it answers.

It writes the way any chatbot does, one token at a time. What’s different is what it learned in training: to spend hundreds or thousands of tokens on a problem, trying an approach, checking it and backing up, before it commits to an answer. This page walks through how that works, from a single token to models that are starting to reason in numbers instead of words, which this site calls neuralese.

Also called
thinking models, large reasoning models
Last updated

01

How a language model writes

A language model writes in a loop, and a reasoning model is no exception. It reads all the text so far, predicts the next token, adds it to the end and goes around again. A token is a chunk of text, usually a word or part of one. Take DeepSeek-R1, the reasoning model this page keeps coming back to, because its makers published how it was built. Its vocabulary has about 128,000 tokens. Mid-sentence, it reads “neuralese” as two of them, “neural” and “ese”.

Mid-sentence, “strawberry” is a single token to R1, number 79,430. The model gets that number, not the ten letters.

Each trip around the loop costs the same. The text runs through the model once, layer by layer, and one token comes out. R1 has 61 layers and 671 billion parameters, of which 37 billion are used for each token. The token that settles a hard problem gets the same 61 layers and 37 billion parameters as the token that ends a greeting.

A language model writes one token per trip through its layersThe question and the answer so far go into a band of 61 layers. One trip through the band produces one token, which joins the end of the answer. Then the longer text goes back in.QUESTIONIf a > 1, then the sum of the real solutions of √(a − √(a + x)) = x is equal toANSWER SO FAR<think> To solve the equationDEEPSEEK-R1-ZERO: 61 LAYERSONE TRIP PER TOKENNEXT TOKEN √JOINS THE ANSWER
Fig. 1 The opening of a real answer from an early version of DeepSeek-R1-Zero, as printed in the DeepSeek-R1 paper, split into tokens by R1’s own tokenizer. Lines under the text mark where each token starts and ends, and a token often includes the space before it. Each token takes one trip through the model’s 61 layers, then joins the answer, and the longer text goes back in for the next trip.

Inside one trip, information only moves up through the layers. Whatever the model works out near the top can get back to the bottom one way: by becoming a token that the next trip reads. So a model that has to answer in a single token gets one trip’s worth of thought. A problem that needs a long chain of steps, each one depending on the last, can run out of layers before it runs out of steps.

Writing the steps down gets around that. Each written step becomes input for the next trip, so the chain can be as long as the problem needs. As 41 researchers put it in 2025, “for sufficiently difficult tasks, Transformers must use chain of thought as a form of working memory.” That’s the idea reasoning models are built on. R1 has the same 61 layers as DeepSeek-V3, the chatbot trained from the same base model. What sets it apart is how much it writes before it answers.

02

How R1 learned to reason

DeepSeek’s first attempt, R1-Zero, started from DeepSeek-V3-Base, a model that had learned from a vast amount of text but hadn’t been taught to work through problems. DeepSeek gave it no examples of reasoning written by people. It gave it problems a program could mark, math with one right answer and code with tests to pass, and two rewards: one for the right answer, and one for putting its working between <think> tags before the answer. That working is the model’s chain of thought. Learning by trial and reward like this is called reinforcement learning.

The tags stayed. R1’s chat template writes the opening <think> for it, because R1 sometimes skipped thinking on its own.

The rewards said nothing about length. R1-Zero’s answers grew longer anyway, and its scores rose with them: in DeepSeek’s words, it “naturally learns to solve reasoning tasks with more thinking time.” It started checking its own steps and trying other approaches. Over 10,400 rounds of training, its score on AIME 2024, a hard American high-school math contest, went from 15.6% to 77.9%, above the average human contestant.

Partway through training, it began stopping to check itself. In one answer DeepSeek published, it’s halfway through the algebra when it writes:

Wait, wait. Wait. That’s an aha moment I can flag here.

DeepSeek-R1-Zero, partway through training

Then it starts over from the first line and checks each step. DeepSeek counted a sudden rise in the word “wait” at that stage of training, and called it “an aha moment for us” too.

R1-Zero reasoned well and read badly. Its chains of thought were hard to follow and slid between English and Chinese. For R1, DeepSeek began with a few thousand examples of reasoning written in a conversational style, ran the reward training again with an extra reward for sticking to one language, then trained the model on writing and everyday questions, and on being helpful and harmless. The language reward cost a little accuracy by DeepSeek’s own tests, and DeepSeek kept it, because people could read the result.

03

What thinking longer buys

OpenAI showed the payoff first. When it released o1, its first reasoning model, in September 2024, it said o1 “consistently improves with more reinforcement learning (train-time compute) and with more time spent thinking (test-time compute).” Test-time compute is the work a model does while it answers, as opposed to while it’s trained. On AIME 2024, GPT-4o solved 12% of the problems. o1 solved 74% on a single try, and 83% when it took a vote across 64 tries.

The tokens a model writes before its answer are often called thinking tokens or reasoning tokens. In R1 they sit between the <think> tags.

Thinking is paid for in tokens. Each step of reasoning is more text to compute, and to bill, so the companies let you choose how hard a model thinks. OpenAI runs its models at settings it calls reasoning effort, like medium and high. Anthropic lets developers set a budget of thinking tokens, and says larger budgets “can improve response quality.” Anthropic still bills for the thinking when it doesn’t show it.

Some of that thinking is wasted. DeepSeek says R1 spends fewer tokens on simple tasks and more on hard ones, and admits that “instances of excessive reasoning—manifested as overthinking—are still observed in response to simpler questions.”

04

Can you trust the written reasoning?

How much of it you see depends on whose model it is. R1’s weights are public, so anyone who runs it can read the whole chain of thought. OpenAI decided at o1’s launch “not to show the raw chains of thought to users,” and shows a summary instead. It gave several reasons, competitive advantage among them. One matters here: a hidden chain of thought, it wrote, “allows us to ‘read the mind’ of the model,” as long as the model isn’t trained to make it look good. Anthropic’s API returns “a readable summary of Claude’s reasoning” for its older models, and on its newest ones, Claude Opus 5.5 among them, an empty field by default.

Reading the full text still isn’t the same as seeing what happened. In 2025, Anthropic slipped hints into test questions. When a hint changed the answer, Claude 3.7 Sonnet mentioned the hint 25% of the time, and R1 39%. Studies of small open models in September 2026 found that they often skip past their own written steps on easy problems, and that concepts that drive the answer can stay out of the text. The site’s myths page goes through what that means for the claim that reasoning in plain English is reasoning you can see.

Reading it works often enough to matter. OpenAI had GPT-4o read the chain of thought of o3-mini, a stronger model, and it caught o3-mini gaming coding tasks. When training then pushed hard against the reasoning the monitor flagged, the model learned to hide its intent and kept cheating. That trade is behind the title 41 researchers gave their 2025 paper on reading reasoning: “A New and Fragile Opportunity for AI Safety.”

05

From words to numbers

Each of those checks depends on the reasoning being written in words. Researchers are building models that skip them. Some feed the model’s internal numbers straight back in where a token would go, which is latent reasoning. Others loop the same layers to think longer without writing anything, which is recurrent depth. This site calls reasoning in numbers instead of words neuralese.

So far they’re research models, and small ones. The closest a big chatbot has come is a report: The Information wrote in September 2026 that GPT-6 Astra loops some of its layers, and OpenAI hasn’t confirmed it. Astra still writes its reasoning out, but its own system card found that reasoning harder to monitor than earlier models’. Does GPT-6 Astra think in neuralese? goes through the evidence.

06

Sources

These are the sources behind this page. We checked them against the original text where we could read it.

  1. Lab publication

    DeepSeek. DeepSeek-R1 model card, configuration and tokenizer. Hugging Face, January 2025.

    R1’s layers, parameters and tokenizer, and its public weights.

  2. Peer-reviewed paper

    How R1-Zero and R1 were trained, the aha moment, overthinking, and the answer in Fig. 1.

  3. Lab publication

    OpenAI. Learning to reason with LLMs. 12 Sep 2024.

    Test-time compute, o1’s AIME scores, and why OpenAI hides the raw chain of thought.

  4. Lab publication

    Anthropic. Thinking. Claude Platform Docs, read 1 Oct 2026.

    Thinking budgets, and what Claude’s API shows of its reasoning.

  5. Lab publication

    OpenAI. GPT-6 Astra System Card. 3 Sep 2026, updated 9, 22 and 29 Sep 2026.

    Reasoning effort settings, and Astra’s reasoning being harder to monitor.

  6. Lab publication

    Claude 3.7 Sonnet and R1 leaving hints out of their reasoning.

  7. Peer-reviewed paper

    Models skipping past their own steps on easy problems.

  8. Preprint

    Qianli Wang, Yilong Wang, Dennis Wei, Jingyi Sun, Simon Ostermann, Isabelle Augenstein, Pepa Atanasova and Nils Feldhus. From Concept Alignment to Causal Grounding: An Intervention Test of Chain-of-Thought Faithfulness. September 2026.

    Concepts that drive the answer staying out of the text.

  9. Preprint

    GPT-4o catching o3-mini, and the model learning to hide its intent.

  10. Preprint

    Why long reasoning has to go through words, and the paper’s title.

  11. Journalism

    Amir Efrati, Stephanie Palazzolo, Rocket Drew. OpenAI Technique in ‘Astra’ Model Sparks Security Concerns. The Information, 1 Sep 2026.

    The report that Astra loops its layers. Paywalled. We haven’t read the full text, and the details here match how other outlets summarized it.