Neuralese Wiki

Last updated

Latent space

What is latent space?

Latent space is the space of numbers an AI model works in. When a model reads a word or looks at a picture, it turns it into a long list of numbers, and that list marks a point in the space. Things the model treats as alike end up close together, so the space works like the model’s map of meaning.

It’s called latent, meaning hidden, because you see only what the model makes of the numbers, not the numbers themselves. When people say a model might reason in neuralese, they mean it would do its thinking in this space, without putting it into words.

Also called
embedding space, representation space
Last updated

01

Vectors, in plain words

A vector is a list of numbers. Two numbers can pin down a place on a map: its latitude and longitude. Three can pin down a spot in a room. Each number is a direction you can move in, and researchers call each direction a dimension.

AI models use much longer lists. GPT-3, which OpenAI described in 2020, turns each piece of a word into a list of 12,288 numbers. People can’t picture a space with 12,288 directions, but it works like the map: lists with similar numbers sit close together, and you can measure how far apart any two are.

Researchers don’t choose what each number means. Training sets them, nudging the numbers again and again toward whatever helps the model do its job, such as guessing the next word. What comes out is an arrangement where words used in similar ways land near each other, though the directions don’t come with labels.

A word becomes a list of numbers, and the list becomes a pointThe word “kitten” turns into a list of numbers that starts 1.2, 0.8. Those first two numbers place it on a map, close to cat, dog and puppy, and far from fruit and vehicles.A WORDkittenITS LIST OF NUMBERS[1.20.8−0.32.1−1.60.4… ]1ST2NDA REAL MODEL’S LIST HAS THOUSANDS1ST NUMBER2ND NUMBERapplepearbananacartruckbuscatdogpuppy1.20.8kittenA word becomes a list of numbers, and the list becomes a pointThe word “kitten” turns into a list of numbers that starts 1.2, 0.8. Those first two numbers place it on a map, close to cat, dog and puppy, and far from fruit and vehicles.A WORDkittenITS LIST OF NUMBERS[1.20.8−0.32.1−1.60.4… ]1ST2NDA REAL MODEL’S LIST HAS THOUSANDS1ST NUMBER2ND NUMBERapplepearbananacartruckbuscatdogpuppy1.20.8kitten
Fig. 1 A word becomes a list of numbers, and the list becomes a point. Here the first two numbers place it on a map, and it lands near the words used like it. A real model’s list has thousands of numbers, so its map has thousands of directions. The numbers and positions here are made up to show the idea.

02

Directions that mean something

In 2013, Tomas Mikolov and two colleagues at Microsoft Research found that some relationships between words show up as directions in this space. The step from man to woman pointed roughly the same way as the steps from uncle to aunt and from king to queen. So if you start at king, take away man and add woman, you land, in the paper’s words, “very close to ‘Queen.’”

The trick is famous, and it has a catch. The standard way of running it doesn’t let the answer be one of the words you started with. In a 2019 paper, Malvina Nissim and two colleagues at the University of Groningen let those words back in. Accuracy on the standard test fell from 74% to 21%, mostly because the answer came back as one of the starting words: man is to king as woman is to king. The directions are there, but they’re rougher than the famous example suggests.

King minus man plus womanArrows from man to woman and from uncle to aunt point the same way. The man-to-woman arrow, moved over to start at king, lands close to queen. The nearest word to where it lands is king; leaving out king, man and woman, it’s queen.manwomanuncleauntkingqueenking − man + womanTHE NEAREST WORDANY WORDkingLEAVING OUT KING, MAN AND WOMANqueenKing minus man plus womanArrows from man to woman and from uncle to aunt point the same way. The man-to-woman arrow, moved over to start at king, lands close to queen. The nearest word to where it lands is king; leaving out king, man and woman, it’s queen.manwomanuncleauntkingqueenking − man + womanTHE NEAREST WORDANY WORDkingLEAVING OUT KING, MAN AND WOMANqueen
Fig. 2 The step from man to woman, moved over to king, lands by queen. That’s in two directions. Google’s word vectors have hundreds, and searching them for the word nearest to king − man + woman finds king itself. The standard code leaves out the three starting words, which leaves queen. The drawing follows Mikolov and colleagues’ 2013 figure, and the answers are the ones Nissim and colleagues report.

03

Where latent space shows up

Word vectors were only the start. Image generators use latent space to save work. Stable Diffusion doesn’t make its pictures pixel by pixel. An encoder first squeezes an image into a smaller block of numbers: a picture 512 pixels square, which takes 786,432 numbers, becomes 16,384. The model learns to make pictures in that smaller space, and a decoder turns the result back into pixels. Its makers called the method latent diffusion, after the space where the work happens.

Some models put different kinds of things in one space. OpenAI’s CLIP, from 2021, learned from 400 million pictures and their captions to place each picture near its own caption, so a photo of a dog and the words “a photo of a dog” end up close together. Stable Diffusion uses CLIP’s text half to turn a prompt into numbers it can follow.

In a chatbot, latent space is where the work gets done. Each piece of the text becomes a vector, and each layer of the model reads those vectors and adds what it has worked out. Researchers at Anthropic call this running total the residual stream, “a communication channel” that “all layers communicate through.” At the top, the last vector becomes a guess at the next word.

04

What researchers have found inside

If concepts sit along directions, you could find the direction for a concept, read it off, or turn it up. Anthropic’s interpretability team, which studies what goes on inside models, starts from that idea: that a network stores meaningful concepts as directions in its space.

It’s harder than it sounds, because a model can store more concepts than it has directions by letting them overlap a little. Anthropic’s researchers showed how in small test networks in 2022 and called it superposition. How far those toy results carry over to real models was, in their words, “very unclear.”

In 2024 they tried a real one. They trained a second network, a sparse autoencoder, to pull the overlapping concepts in Claude 3 Sonnet apart, and found millions of what they call features. One fired on the Golden Gate Bridge, in text in several languages and in pictures. Turned up, it made the model bring the bridge into most of its replies. For 24 hours anyone could try this Golden Gate Claude: asked how to spend $10, it suggested driving across the bridge and paying the toll. Anthropic says the same method can turn safety-related features up or down, like those for dangerous code or deception.

In 2025 Anthropic looked inside Claude 3.5 Haiku and found the same features standing for the same concepts in English, French and Chinese, which it says suggests “a kind of universal ‘language of thought.’”

These tools see part of the space. Even with 34 million features, the researchers found evidence that what they had uncovered was “an incomplete description of the model’s internal representations.”

05

Latent space and neuralese

A chatbot that reasons in words leaves latent space at each step. Its last vector, thousands of numbers, comes out as a single word or part of one, and the next step starts from there. AI 2027 described the alternative: passing those vectors back to the model’s early layers, “potentially transmitting over 1,000 times more information.”

That’s what this site calls neuralese: reasoning that stays in latent space. In COCONUT, a method published in 2024, a model feeds its last vector straight back in as its next input instead of writing a word. Recurrent-depth models loop through the same layers, working on the same vectors for longer before they write anything. What is latent reasoning? goes through these methods and how well they work, and Which AI models think in neuralese? tracks which models use them.

It’s also why neuralese worries the people who check AI’s reasoning. Written reasoning can be read. Reasoning in vectors has to be decoded with tools like the ones above, which see only part of the space. As AI 2027 put it, “Unlike English words, these high-dimensional vectors are likely quite difficult for humans to interpret.”

06

Sources

These are the sources behind this page. We checked them against the original text on 1 October 2026.

  1. Peer-reviewed paper

    Tom B. Brown et al. Language Models are Few-Shot Learners. NeurIPS 2020.

    GPT-3’s 12,288 numbers for each piece of a word.

  2. Peer-reviewed paper

    Tomas Mikolov, Wen-tau Yih, Geoffrey Zweig. Linguistic Regularities in Continuous Space Word Representations. NAACL-HLT 2013.

    King − man + woman, and the drawing Fig. 2 follows.

  3. Peer-reviewed paper

    Malvina Nissim, Rik van Noord, Rob van der Goot. Fair Is Better than Sensational: Man Is to Doctor as Woman Is to Doctor. Computational Linguistics, 2020.

    The catch, and the two answers in Fig. 2.

  4. Peer-reviewed paper

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Björn Ommer. High-Resolution Image Synthesis with Latent Diffusion Models. CVPR 2022.

    Latent diffusion.

  5. Lab publication

    Robin Rombach, Patrick Esser. Stable Diffusion v1-4 model card. Hugging Face, August 2022.

    How much Stable Diffusion squeezes a picture, and its use of CLIP.

  6. Peer-reviewed paper

    CLIP’s shared space for pictures and captions.

  7. Lab publication

    Nelson Elhage et al. (Anthropic). A Mathematical Framework for Transformer Circuits. Transformer Circuits Thread, December 2021.

    The residual stream.

  8. Lab publication

    Nelson Elhage et al. (Anthropic). Toy Models of Superposition. Transformer Circuits Thread, September 2022.

    Superposition, and how far it carries over.

  9. Lab publication

    Adly Templeton et al. (Anthropic). Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet. Transformer Circuits Thread, May 2024.

    Features in Claude 3 Sonnet, the Golden Gate Bridge feature, and what’s still missing.

  10. Lab publication

    Anthropic. Golden Gate Claude. 23 May 2024.

    Golden Gate Claude, and turning features up or down.

  11. Lab publication

    The same features across English, French and Chinese.

  12. Forecast

    Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, Romeo Dean. AI 2027. April 2025.

    Passing vectors back instead of words, and why they’re hard to read.

  13. Peer-reviewed paper

    Feeding the last vector back in.

  14. Peer-reviewed paper

    Recurrent depth.