Latent

How large language models work

A mind made of numbers

You are about to watch a machine learn to write. There is no magic in it — only numbers, and a patient way of changing them. Scroll, and we will build one, starting from a single number.

01 One knob

Can a machine learn a number?

A shop sells apples: 1 kg costs $3, 2 kg $6, 4 kg $12. Our machine has one knob, w, and guesses price = w × kg. It starts at w = 0.5. The loss — the average squared error — says how wrong it is. The slope of the loss says which way to turn the knob; the learning rate, how far. Guess, measure, nudge.

w0.50
loss46.88
step0
learning rate

Drag on the valley, or press Train. Fitting by least squares: Legendre, 1805. Gradient descent: Cauchy, 1847.

02 Two knobs

A line, and a bowl

Add a second knob, b, and the guess becomes w × x + b — any straight line. Here it learns Celsius → Fahrenheit. Over two knobs the loss is a bowl, and learning is a ball rolling to the bottom. Real models have billions of knobs. Nobody can picture that bowl, but the rule is the same: feel the slope, step downhill.

step0
w · b0.20 · −60.0
loss21,456

It lands on w = 1.8, b = 32 — and yes, −40 °C is −40 °F.

03 The neuron

Add, then bend

A neuron takes several numbers in, multiplies each by a weight, adds them up with a bias — then bends the result. The simplest bend, ReLU, is max(0, z): below zero becomes zero. A hinge. Why bend? Lines added to lines are still a line; without the bend, a thousand layers could only ever draw a line.

max(0, z)the whole bend

Loosely inspired by brain cells — McCulloch & Pitts, 1943; Rosenblatt’s perceptron, 1958 — but really just arithmetic.

04 The network

Many hinges, any shape

Each neuron adds one hinge to the curve; with enough hinges you can trace any smooth shape. Put neurons in layers and you have a network. When it is wrong, backpropagation sends the blame backwards through every connection, and each weight gets its own nudge — the same nudge as the first knob, for every knob at once.

knobs25
training step0
loss—

A real 25-knob network, training live in your browser. Backpropagation: Rumelhart, Hinton & Williams, 1986. Universal approximation, 1989.

05 Tokens

Words become numbers

A network only eats numbers, so text is cut into tokens: common words stay whole, rare ones break into pieces, and each piece is an integer ID. The pieces are learned by byte-pair encoding: start from single characters, keep merging the most frequent neighbouring pair. This page’s tokenizer just learned its pieces, in your browser, from the text you are reading.

— tokens · — characters · vocabulary of — on this page

50,257tokens in GPT-2’s vocabulary; modern ones ~100k–200k+

≈ ¾of an English word per token (~4 characters)

Byte-pair tokens: Sennrich et al., 2016, from a 1994 compression trick. Splits vary by model.

06 Embeddings illustrative map

Meaning becomes geometry

Every token ID looks up a row in a giant table: a long list of numbers, its embedding. They start random and are learned like every other knob. Tokens used in similar ways end up close together — cat near dog, Paris near Rome — and directions carry meaning: king − man + woman ≈ queen.

12,288numbers per token in GPT-3 — we can only draw three

word2vec, Mikolov et al., 2013. Positions here are hand-placed to show the idea.

07 The only game illustrative probabilities

Guess the next token

The whole model is trained on one task: given the text so far, predict the next token. It gives a probability to every token in its vocabulary. The loss is surprise: −log of the probability it gave the token that really came next. To write, it samples one, appends it, and runs again — one token at a time.

picked␣mat
its probability31%
surprise1.17

Low temperature: safe and predictable. High: adventurous. Surprise in nats (natural log).

08 Attention illustrative weights

Words look at each other

“Bank” means one thing by a river and another at an ATM: a token needs its context. With attention, each token makes a query (what am I looking for?), a key (what do I contain?) and a value (what I pass on). Queries meet earlier keys; the matches decide how much of each value to blend in. At “it” the model hedges between animal and street; at “tired”, it settles.

96 × 96heads per layer × layers in GPT-3

09 The Transformer

Stack it

“Attention Is All You Need” (Vaswani et al., 2017). One block is attention — tokens share information — plus a feed-forward network, where each token thinks on its own. Each block adds its result to a running vector per token: the residual stream, a river rising through the stack. At the top, the last token’s vector scores every token in the vocabulary.

96blocks stacked in GPT-3

in parallelevery token of a text is processed at once during training — perfect for GPUs. That’s why it scaled.

10 Scale

Then make it enormous

The recipe barely changed. The size did. Scaling laws showed loss falling along a smooth, predictable power law as parameters, data and compute grow — predictable enough to plan a run before it starts. Along the way, abilities nobody programmed appeared: translation, arithmetic, code.

175Bparameters in GPT-3 (2020), trained on ~300 billion tokens

15Ttokens read by Llama 3.1 (2024)

3.14 × 1023arithmetic operations to train GPT-3

~20tokens per parameter (Chinchilla)

Kaplan et al., 2020; Hoffmann et al., 2022. Today’s frontier models: sizes not disclosed. Curve shape illustrative.

11 From predictor to partner

Teaching it to help

A freshly pretrained model is a brilliant mimic, not an assistant. Post-training shapes it: fine-tuning on example conversations; RLHF, where people compare two answers and the model is nudged toward the preferred one (InstructGPT, 2022); and Constitutional AI (Anthropic, 2022), where it critiques and revises its answers against written principles. Underneath: gradients nudging numbers.

prompt

pretrained illustrative

after post-training illustrative

12 Thinking and doing

Reason, use tools, act

Given room to reason step by step, models solve harder problems (chain of thought, 2022); modern ones are trained to think before they answer. With tools they write a request — search, run code — and read the result back as tokens. Agents loop think → act → observe: write code, run the tests, fix the bug, repeat, for hours.

illustrative agent loop

    2,048 → 1,000,000tokens of context: GPT-3 (2020) → Claude Fable (≈ 750,000 words)

    13 Looking inside other features illustrative

    Reading the mind we grew

    Nobody wrote this program; it grew. Interpretability reads the numbers back. In 2024 Anthropic found millions of readable features inside Claude 3 Sonnet — one fires for the Golden Gate Bridge; turn it up, and the model steers every topic back to the bridge. In 2025, circuit tracing caught a model choosing the rhyme “rabbit” before writing its line: it plans ahead.

    The other feature names floating here are illustrative examples of the kinds of features found.

    14 The frontier

    Claude Fable

    At its core, the same recipe — tokens → embeddings → layers of attention and feed-forward networks → next-token probabilities — trained with gradients, and shaped by post-training into an assistant that thinks before it answers.

    tokens read at once — a large codebase, or a shelf of books

    tokens written in one reply

      Parameter count, training data and compute: not disclosed.

      14 All the way down

      One layer. One neuron. One weight.

      Zoom in on any of it. A layer is a crowd of neurons; a neuron is a few weights and a bend; a weight is a single number — a knob, like the very first one.

      Inside, every one of its numbers is a knob — like the first one — nudged a little at a time, until the numbers learned to talk.