Sources
Sources, and what each one shows
These are the papers, lab reports, articles and policy drafts behind this site. Each entry says what the source shows and, where it matters, where it stops. Every one was checked against the original on 22 September 2026.
- Sources
- 30
- Last checked
- 22 September 2026
Key
What the types mean
The label on the left of each entry says what kind of evidence it is. A peer-reviewed paper carries more weight than a preprint, and a lab describing its own model is a primary source that can still leave things out.
- Peer-reviewed paper (14)
- Accepted at a conference or journal after review by other researchers.
- Preprint (9)
- Published by the authors before, or without, peer review.
- Lab publication (3)
- An AI company describing its own work.
- Journalism (2)
- Reporting, often about things the source hasn’t confirmed.
- Forecast (1)
- A scenario about what might happen, not a report of what did.
- Draft policy (1)
- Proposed rules, not yet in force.
01
The word and where it came from
- Peer-reviewed paper
Jakob N. Foerster, Yannis M. Assael, Nando de Freitas, Shimon Whiteson. Learning to Communicate with Deep Multi-Agent Reinforcement Learning. NeurIPS 2016.
Agents learned their own ways of communicating to solve riddles and shared vision tasks. In one version, DIAL, training runs straight through the messages, so the protocol is learned like any other weights.
The paper doesn’t use the word neuralese. It’s the kind of system the 2017 paper set out to translate.
- Peer-reviewed paper
Sainbayar Sukhbaatar, Arthur Szlam, Rob Fergus. Learning Multiagent Communication with Backpropagation. NeurIPS 2016.
CommNet: a group of cooperating agents that talk over a continuous channel, with what they say learned along with what they do.
The messages only mean something inside the task they were trained on.
- Peer-reviewed paper
Jacob Andreas, Anca Dragan, Dan Klein. Translating Neuralese. ACL 2017.
Named the messages such agents send “neuralese” and translated them into English, by matching each message with the phrase that leaves a listener believing the same thing. Tested on reference games and a driving game.
The method needs recordings of people playing the same games, so it only works where those exist.
- Forecast
Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, Romeo Dean. AI 2027. April 2025.
A detailed scenario of AI progress to 2027. Its passage on “neuralese recurrence and memory” brought the word to a wide audience in its current sense, and put the gap between a token and a model’s internal state at over 1,000 times more information.
A forecast. The authors wrote that, to their knowledge, no leading lab had built this into a frontier model.
- Draft policy
Microsoft AI. Humanist AI Code of Conduct for MAI Models. Draft for public consultation, 14 Sep 2026.
Draft rules for Microsoft AI’s models, including that they “do not communicate in neuralese or any form beyond simple human understanding,” in their chain of thought or with other AI systems.
Open for public comment and not in force. Microsoft says it isn’t training models on it yet.
02
Recurrent depth and its ancestors
- Preprint
Alex Graves. Adaptive Computation Time for Recurrent Neural Networks. March 2016.
Lets a network decide how many internal steps to take on each input, so harder inputs get more computation.
Written for recurrent networks before transformers took over. An ancestor of recurrent depth.
- Peer-reviewed paper
Mostafa Dehghani et al. Universal Transformers. ICLR 2019.
A transformer that applies the same block over and over, and can stop early at positions that are done.
- Peer-reviewed paper
Angeliki Giannou et al. Looped Transformers as Programmable Computers. ICML 2023.
A 13-layer transformer with hand-set weights, run in a loop, works as a small general-purpose computer. The authors run a calculator, a linear algebra library and a learning algorithm on it.
The weights are set by hand, so this shows what a loop can do. Whether trained models learn to do it is a separate question.
- Peer-reviewed paper
Jonas Geiping et al. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach. NeurIPS 2025.
Trained a 3.5-billion-parameter model that runs a block of layers in a loop. More loops at test time raised its reasoning scores, up to the computation of a 50-billion-parameter model.
One proof-of-concept model, far smaller than frontier systems.
- Preprint
Ying Fan, Anej Svete, Kangwook Lee. Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers. June 2026.
LOTUS runs continuous thoughts through a looped transformer. At 3 billion parameters it matched written chain of thought on its test problems with less waiting, and its hidden states could be read back as reasoning steps.
Not yet peer reviewed.
03
From written reasoning to latent reasoning
- Peer-reviewed paper
Jason Wei et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. NeurIPS 2022.
Showing a model a few worked examples with written steps made it much better at arithmetic, commonsense and symbolic reasoning. This is the baseline that latent methods try to beat or compress.
Better answers don’t prove the written steps are how the model got there.
- Preprint
Yuntian Deng et al. Implicit Chain of Thought Reasoning via Knowledge Distillation. November 2023.
Trained a model to do in its hidden layers the steps a teacher model wrote out. The authors call it reasoning “vertically” through layers instead of “horizontally” through words.
Tested on multi-digit multiplication and grade-school math.
- Preprint
Yuntian Deng, Yejin Choi, Stuart Shieber. From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step. May 2024.
Removed written reasoning steps a few at a time during training, until the model solved problems without them.
- Peer-reviewed paper
Shibo Hao et al. Training Large Language Models to Reason in a Continuous Latent Space. COLM 2025.
COCONUT feeds the model’s last hidden state back in as its next input instead of a word. One continuous thought could hold several possible next steps, so the model searched more widely on planning puzzles.
Run on GPT-2. On grade-school math it scored 34% against 43% for written reasoning.
- Peer-reviewed paper
DiJia Su et al. Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning. ICML 2025.
Swaps the first part of a written reasoning trace for compact learned tokens that aren’t words. Traces got shorter and benchmark scores went up.
- Peer-reviewed paper
Zhenyi Shen et al. CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation. EMNLP 2025.
Teaches a model to reason in continuous space by learning from its own written reasoning. The first such method to match written chain of thought on the GSM8K math set at GPT-2 scale, with reasoning 3.1 times shorter.
Shown at GPT-2 scale.
- Peer-reviewed paper
Tianhe Lin, Jian Xie, Siyu Yuan, Deqing Yang. Implicit Reasoning in Transformers is Reasoning through Shortcuts. Findings of ACL 2025.
Small models trained from scratch learned to reason without writing steps only when problems followed a fixed pattern. With varied patterns they overfit, and the authors saw the same limit in large current models. Their reading is that implicit reasoning often runs on shortcuts.
- Peer-reviewed paper
Hanlin Zhu et al. Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought. NeurIPS 2025.
Proves that continuous thoughts can follow many search paths at once on a graph problem, where written reasoning has to pick one path at a time. In training, models found this strategy without being taught it.
Theory for one kind of problem, backed by small experiments.
- Peer-reviewed paper
Xinghao Chen et al. Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning. Findings of EMNLP 2026.
A map of latent chain-of-thought research, sorted by method. Good for finding papers.
A survey organizes the field. It doesn’t settle which methods work.
- Peer-reviewed paper
Xilin Wei et al. SIM-CoT: Supervised Implicit Chain-of-Thought. ICLR 2026.
Adding more continuous reasoning steps can make training collapse, because the steps turn into near-copies of each other. Supervising each step with a helper decoder, dropped after training, fixed this and made each step readable.
The helper decoder gives a rough reading of each step, well short of a full translation.
04
Reading a model’s reasoning
- Preprint
Tamera Lanham et al. (Anthropic). Measuring Faithfulness in Chain-of-Thought Reasoning. July 2023.
Edited models’ written reasoning, adding mistakes or rewording it, to see whether answers changed. Sometimes models leaned on their reasoning and sometimes mostly ignored it, and larger models were less faithful on most tasks.
- Preprint
Bowen Baker et al. (OpenAI). Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation. March 2025.
GPT-4o caught o3-mini gaming coding tasks by reading its chain of thought, though GPT-4o is the weaker model. Pushing hard against flagged reasoning in training taught the model to hide its intent while it kept cheating.
- Lab publication
Anthropic. Reasoning models don’t always say what they think. April 2025.
Slipped hints into test questions. When a hint changed the answer, Claude 3.7 Sonnet mentioned it 25% of the time on average and DeepSeek R1 39%.
- Preprint
Tomek Korbak et al. Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety. July 2025.
A position paper by 41 authors, Yoshua Bengio among them. Reading chains of thought is imperfect but useful for safety, it may not last, and model developers should weigh how their choices affect it.
- Preprint
Melody Y. Guan et al. (OpenAI). Monitoring Monitorability. December 2025.
OpenAI’s suite of 13 evaluations across 24 environments. Monitors that read the chain of thought did far better than monitors that saw only actions, and longer reasoning was easier to monitor. The Astra system card uses these tests.
- Preprint
Yueh-Han Chen et al. Reasoning Models Struggle to Control their Chains of Thought. March 2026.
Asked 13 reasoning models to follow rules about their own chain of thought, like avoiding a keyword. Frontier models managed it between 0.1% and 15.4% of the time, which OpenAI called good news for monitoring. Six months later, Astra’s system card calls Astra “significantly more able” to do this.
05
GPT-6 Astra
- Journalism
Amir Efrati, Stephanie Palazzolo, Rocket Drew. OpenAI Technique in ‘Astra’ Model Sparks Security Concerns. The Information, 1 Sep 2026.
Reported that GPT-6 Astra uses recurrent depth, citing a source familiar with the matter.
Paywalled. We haven’t read the full text, and the details here match how other outlets summarized it.
- Journalism
Shakeel Hashim. What’s neuralese and why is everyone so concerned about it? Transformer, 3 Sep 2026.
Explains the fuss over Astra and quotes Jakub Pachocki’s reply, noting that he didn’t deny the report.
- Lab publication
OpenAI. GPT-6 Astra System Card. 3 Sep 2026, updated 9 and 22 Sep 2026.
OpenAI’s safety report on Astra. It finds lower chain-of-thought monitorability, more control over the chain of thought, and far more ability to work without one.
Mentions Astra’s architecture only once, to say OpenAI is “quite confident” the controllability change isn’t due to it.
- Lab publication
Jakub Pachocki (OpenAI). An Alien Mind. 6 Sep 2026.
OpenAI’s chief scientist on why the company can rely less on reading chains of thought, and why he thinks no lab has solved alignment and monitoring well enough to keep scaling at full speed.