How large language models work
A mind made of numbers
You are about to watch a machine learn to write. There is no magic in it — only numbers, and a patient way of changing them. Scroll, and we will build one, starting from a single number.
01 One knob
Can a machine learn a number?
A shop sells apples: 1 kg costs $3, 2 kg $6, 4 kg $12. Our machine has one knob, w, and guesses price = w × kg. It starts at w = 0.5. The loss — the average squared error — says how wrong it is. The slope of the loss says which way to turn the knob; the learning rate, how far. Guess, measure, nudge.
Drag on the valley, or press Train. Fitting by least squares: Legendre, 1805. Gradient descent: Cauchy, 1847.
02 Two knobs
A line, and a bowl
Add a second knob, b, and the guess becomes w × x + b — any straight line. Here it learns Celsius → Fahrenheit. Over two knobs the loss is a bowl, and learning is a ball rolling to the bottom. Real models have billions of knobs. Nobody can picture that bowl, but the rule is the same: feel the slope, step downhill.
It lands on w = 1.8, b = 32 — and yes, −40 °C is −40 °F.
03 The neuron
Add, then bend
A neuron takes several numbers in, multiplies each by a weight, adds them up with a bias — then bends the result. The simplest bend, ReLU, is max(0, z): below zero becomes zero. A hinge. Why bend? Lines added to lines are still a line; without the bend, a thousand layers could only ever draw a line.
max(0, z)the whole bend
Loosely inspired by brain cells — McCulloch & Pitts, 1943; Rosenblatt’s perceptron, 1958 — but really just arithmetic.
04 The network
Many hinges, any shape
Each neuron adds one hinge to the curve; with enough hinges you can trace any smooth shape. Put neurons in layers and you have a network. When it is wrong, backpropagation sends the blame backwards through every connection, and each weight gets its own nudge — the same nudge as the first knob, for every knob at once.
A real 25-knob network, training live in your browser. Backpropagation: Rumelhart, Hinton & Williams, 1986. Universal approximation, 1989.
05 Tokens
Words become numbers
A network only eats numbers, so text is cut into tokens: common words stay whole, rare ones break into pieces, and each piece is an integer ID. The pieces are learned by byte-pair encoding: start from single characters, keep merging the most frequent neighbouring pair. This page’s tokenizer just learned its pieces, in your browser, from the text you are reading.
tokens · characters · vocabulary of on this page
50,257tokens in GPT-2’s vocabulary; modern ones ~100k–200k+
≈ ¾of an English word per token (~4 characters)
Byte-pair tokens: Sennrich et al., 2016, from a 1994 compression trick. Splits vary by model.
06 Embeddings illustrative map
Meaning becomes geometry
Every token ID looks up a row in a giant table: a long list of numbers, its embedding. They start random and are learned like every other knob. Tokens used in similar ways end up close together — cat near dog, Paris near Rome — and directions carry meaning: king − man + woman ≈ queen.
12,288numbers per token in GPT-3 — we can only draw three
word2vec, Mikolov et al., 2013. Positions here are hand-placed to show the idea.
07 The only game illustrative probabilities
Guess the next token
The whole model is trained on one task: given the text so far, predict the next token. It gives a probability to every token in its vocabulary. The loss is surprise: −log of the probability it gave the token that really came next. To write, it samples one, appends it, and runs again — one token at a time.
Low temperature: safe and predictable. High: adventurous. Surprise in nats (natural log).
08 Attention illustrative weights
Words look at each other
“Bank” means one thing by a river and another at an ATM: a token needs its context. With attention, each token makes a query (what am I looking for?), a key (what do I contain?) and a value (what I pass on). Queries meet earlier keys; the matches decide how much of each value to blend in. At “it” the model hedges between animal and street; at “tired”, it settles.
96 × 96heads per layer × layers in GPT-3
09 The Transformer
Stack it
“Attention Is All You Need” (Vaswani et al., 2017). One block is attention — tokens share information — plus a feed-forward network, where each token thinks on its own. Each block adds its result to a running vector per token: the residual stream, a river rising through the stack. At the top, the last token’s vector scores every token in the vocabulary.
96blocks stacked in GPT-3
in parallelevery token of a text is processed at once during training — perfect for GPUs. That’s why it scaled.
10 Scale
Then make it enormous
The recipe barely changed. The size did. Scaling laws showed loss falling along a smooth, predictable power law as parameters, data and compute grow — predictable enough to plan a run before it starts. Along the way, abilities nobody programmed appeared: translation, arithmetic, code.
175Bparameters in GPT-3 (2020), trained on ~300 billion tokens
15Ttokens read by Llama 3.1 (2024)
3.14 × 1023arithmetic operations to train GPT-3
~20tokens per parameter (Chinchilla)
Kaplan et al., 2020; Hoffmann et al., 2022. Today’s frontier models: sizes not disclosed. Curve shape illustrative.
11 From predictor to partner
Teaching it to help
A freshly pretrained model is a brilliant mimic, not an assistant. Post-training shapes it: fine-tuning on example conversations; RLHF, where people compare two answers and the model is nudged toward the preferred one (InstructGPT, 2022); and Constitutional AI (Anthropic, 2022), where it critiques and revises its answers against written principles. Underneath: gradients nudging numbers.
prompt
pretrained illustrative
after post-training illustrative
12 Thinking and doing
Reason, use tools, act
Given room to reason step by step, models solve harder problems (chain of thought, 2022); modern ones are trained to think before they answer. With tools they write a request — search, run code — and read the result back as tokens. Agents loop think → act → observe: write code, run the tests, fix the bug, repeat, for hours.
2,048 → 1,000,000tokens of context: GPT-3 (2020) → Claude Fable (≈ 750,000 words)
13 Looking inside other features illustrative
Reading the mind we grew
Nobody wrote this program; it grew. Interpretability reads the numbers back. In 2024 Anthropic found millions of readable features inside Claude 3 Sonnet — one fires for the Golden Gate Bridge; turn it up, and the model steers every topic back to the bridge. In 2025, circuit tracing caught a model choosing the rhyme “rabbit” before writing its line: it plans ahead.
The other feature names floating here are illustrative examples of the kinds of features found.
14 The frontier
Claude Fable
At its core, the same recipe — tokens → embeddings → layers of attention and feed-forward networks → next-token probabilities — trained with gradients, and shaped by post-training into an assistant that thinks before it answers.
tokens read at once — a large codebase, or a shelf of books
tokens written in one reply
Parameter count, training data and compute: not disclosed.
14 All the way down
One layer. One neuron. One weight.
Zoom in on any of it. A layer is a crowd of neurons; a neuron is a few weights and a bend; a weight is a single number — a knob, like the very first one.