neuralese.wiki

Last checked

Claim ledger

Does GPT-6 Astra think in neuralese?

Not in the way the phrase suggests. Astra still writes out its reasoning before it answers, and nothing public shows that it has swapped words for vectors.

What OpenAI has confirmed is different, and it’s the part OpenAI itself is worried about. Astra’s written reasoning is harder to monitor than earlier models’, and Astra can do far more without writing any reasoning at all.

The story about Astra’s design is still unconfirmed. The Information reported that Astra loops some of its layers, a technique called recurrent depth. OpenAI hasn’t confirmed that, and says it is “quite confident” the change in how well Astra controls its reasoning isn’t down to architecture.

Model
GPT-6 Astra, OpenAI
Released
3 September 2026
Claims tracked
10
Last checked
22 September 2026

01

Claim by claim

Each row is something people have said about Astra. The mark on the left says where the evidence stands, and the text under each claim says why.

Confirmed
The primary source says it on the record.
Reported
Credible journalism, not confirmed by the source.
Not established
Claimed or implied, but the evidence doesn’t support it yet.
Wrong
Contradicted by the evidence.
Open question
Nobody knows yet, or researchers disagree.
  1. Confirmed

    #Astra’s chain of thought is harder to monitor than earlier OpenAI models’.

    The system card reports “a substantial decrease in chain-of-thought monitorability compared to previous models.” Most of the evidence comes from adversarial tests, where the model is told to evade monitors.

    Astra system card

  2. Confirmed

    #Astra can do much more without writing any reasoning at all.

    With its reasoning switched off, Astra could complete a set of tasks “greatly expanded compared to prior models.” The card cites UK AI Security Institute evaluations suggesting the length of task it can handle this way “may have increased by about an order of magnitude.” Reasoning that never gets written down can’t be read by a monitor.

    Astra system card

  3. Confirmed

    #Astra is better at controlling what shows up in its chain of thought.

    The card says Astra is “significantly more able to control its own chain-of-thought.” In one test, telling it that a monitor was watching made it cut its reasoning short, which let it occasionally slip past monitors that read only the chain of thought.

    Astra system card

  4. Reported

    #Astra uses recurrent depth, a technique that runs some of its layers in a loop.

    The Information reported this on 1 September, citing a source familiar with the matter. Summaries of the paywalled article say OpenAI limited the technique so Astra’s written reasoning stays readable. OpenAI hasn’t confirmed it, and the system card doesn’t describe Astra’s architecture. Jakub Pachocki, OpenAI’s chief scientist, didn’t deny it in his public reply. He said the depth of the computation graph for OpenAI’s current frontier models, Astra included, “is within a factor of two of GPT-4.” That puts a ceiling on how deep Astra’s computation goes, looped or not.

    The Information
    Transformer
    Astra system card

  5. Not established

    #Recurrent depth caused the drop in monitorability.

    OpenAI says it is “quite confident” that the change in controllability is not “differentially due to any architectural changes,” and that the change tracks Astra’s gains at working without reasoning. In his 6 September essay, Pachocki blamed “a combination of factors”: more complex, tool-heavy tasks, models getting better at manipulating their own reasoning, and models becoming “much smarter even without using verbalized reasoning at all.” Architecture isn’t on his list.

    Astra system card
    An Alien Mind

  6. Not established

    #Astra reasons in COCONUT-style continuous thoughts.

    Nothing public says so. The reporting is about recurrent depth, which is a different technique. How the two differ.

    COCONUT paper
    Geiping et al. 2025

  7. Not established

    #Astra hides its reasoning inside ordinary-looking text.

    OpenAI’s current evaluations “indicate Astra is not able to perform steganographic CoT reasoning,” which means tucking hidden reasoning into unrelated text. OpenAI takes this to mean that evading monitors mostly works for misbehavior that doesn’t need written reasoning in the first place.

    Astra system card

  8. Not established

    #“Neuralese” is OpenAI’s term for how Astra works.

    The word appears nowhere in the system card or in Pachocki’s essay. It does appear in Microsoft’s draft code of conduct, which isn’t about Astra.

    Astra system card
    An Alien Mind
    Microsoft draft code

  9. Wrong

    #Astra has stopped thinking in words.

    Astra still writes a chain of thought, and most of OpenAI’s monitorability testing is about reading it. What changed is how much that chain of thought gives away.

    Astra system card

  10. Open question

    #A large share of Astra’s reasoning happens outside its written chain of thought.

    Nobody outside OpenAI can measure this yet. The results with reasoning switched off suggest more happens silently than in earlier models. OpenAI says it is working on alternatives to reading the chain of thought, such as monitoring the model’s activations directly.

    Astra system card

02

What happened, in order

  1. 1 Sep 2026

    The Information reports that Astra uses recurrent depth, citing a source familiar with the matter.

    The Information

  2. 1–2 Sep 2026

    Jakub Pachocki, OpenAI’s chief scientist, replies on X. He says the depth of computation in OpenAI’s current frontier models, Astra included, is within a factor of two of GPT-4. He doesn’t deny the report.

    Transformer

  3. 3 Sep 2026

    OpenAI releases GPT-6 Astra. Its system card reports lower chain-of-thought monitorability, more control over the chain of thought, and far more ability to work without one.

    Astra system card

  4. 6 Sep 2026

    Pachocki publishes “An Alien Mind.” He writes that OpenAI’s ability to rely on chain-of-thought monitoring “is progressively diminishing,” and that no lab has solved alignment and monitoring well enough “to continue responsibly scaling at maximum speed for much longer.”

    An Alien Mind

  5. 14 Sep 2026

    Microsoft AI publishes a draft code of conduct saying its models “do not communicate in neuralese or any form beyond simple human understanding.” Public comment runs for six weeks.

    Microsoft draft code

03

What would change these statuses

We move a claim when the evidence moves, and note every change below. These are the likeliest triggers.

04

Changes to this page

  1. 22 Sep 2026

    First version.

05

Sources

Every claim on this page traces to one of these. All were checked on 22 September 2026, against the original text wherever we could read it.

  1. Lab publication

    OpenAI. GPT-6 Astra System Card. 3 Sep 2026, updated 9 and 22 Sep 2026.

    The monitorability, controllability and no-reasoning findings.

  2. Lab publication

    Jakub Pachocki (OpenAI). An Alien Mind. 6 Sep 2026.

    Pachocki on why monitorability is falling.

  3. Journalism

    Amir Efrati, Stephanie Palazzolo, Rocket Drew. OpenAI Technique in ‘Astra’ Model Sparks Security Concerns. The Information, 1 Sep 2026.

    The recurrent-depth report. Paywalled. We haven’t read the full text, and the details here match how other outlets summarized it.

  4. Journalism

    Shakeel Hashim. What’s neuralese and why is everyone so concerned about it? Transformer, 3 Sep 2026.

    Pachocki’s reply on X.

  5. Draft policy

    Microsoft AI. Humanist AI Code of Conduct for MAI Models. Draft for public consultation, 14 Sep 2026.

    Neuralese in Microsoft’s draft code.

  6. Peer-reviewed paper

    What recurrent depth is.

  7. Peer-reviewed paper

    What continuous thoughts are.