Inside a 3,184-parameter language model

A one-layer, two-head transformer trained on a 28-word toy corpus. Every number on this page is a live forward pass through those weights, run in your browser. Trained offline with hand-written backprop (gradient check: 7.9e-7 relative error) to a loss of 0.496 nats, against 3.332 for a uniform guess. The JavaScript forward pass was checked against a separate numpy implementation and agrees to 5e-7.

The model sees the last 8 words. Click a word to inspect what the model computes at that position; the bar beneath it is how surprised the model was by the word that followed (taller means more bits).

1

Embed

Each word indexes a row of a learned table of 16-number vectors. A position vector is added so the model can tell order. After this step the words are gone; the model holds only vectors.

x0t = E[idt] + Pt E: 28×16, P: 8×16
2

Attend

Every word is layer-normalised, then projected to a query, a key and a value, 8 numbers each per head. A query's dot product with an earlier word's key is a relevance score. Softmax turns the scores into weights, and the weights mix the values. Click any cell for the arithmetic.

Ah = softmax( β · mask( QhKhT / √8 ) )    x1 = x0 + Σh AhVhWOh β = 1 in the real model
3

Think

Each position passes alone through a small network: layer-norm, expand to 32 units, GELU, project back, add to the stream. Attention moves information between words; this step transforms it within one.

x2 = x1 + W2 · GELU( W1 · LN(x1) + c1 ) + c2
4

Read out

The final vector is multiplied by the unembedding matrix to give one score per word. Softmax makes probabilities. Temperature, top-k and top-p reshape them before a word is drawn.

p = softmax( x2U / τ )    U: 16×28

Logit lens: the prediction forming

The same unembedding applied to the stream after each stage, at the selected position. It shows what the model would say if it stopped there.

Sampling at the selected position

5

Poke the math

Everything above is recomputed with these knobs, and generation uses them too. The comparison shows the untouched model next to your settings at the selected position, at τ = 1.

6

Weight spectra

Each head's attention pattern depends on WQK = WQ,hWK,hT and what it writes on WOV = WV,hWO,h. Both are 16×16 but have rank at most 8. The bars are their singular values, computed by SVD of the trained weights.

Things to try

  • Switch off head 1 with the prompt "the cat sat on the". The prediction collapses to "mouse" at 92%, and the same default appears for most prompts (KL from 2 to 25 nats across seven test prompts). Switch off head 0 instead and the damage depends on the prompt: 0.35 nats for "the cat and the dog sat on the", 12.5 for "the black cat saw the". Click through head 1's attention map: in the windows checked, its later positions often point back at the verb ("chased", "sat").
  • Drag β to 0. Every score becomes 0, so each word averages equally over everything it can see. For "the big dog chased the" the answer moves from mouse, small and cat (26% to 28% each) to "dog" at 85%. Then drag β up to 3: for most prompts almost nothing changes (KL at most 0.02; the exception in these tests was "the black cat saw the" at 0.40), because the trained attention rows are already close to one-hot and sharpening a sharp softmax does little. (At position 0 β has no effect at all; the word can only attend to itself.)
  • Read the logit lens at the last word of "the cat sat on the". After embedding alone the guess is flat and says nothing about the prompt (mouse 5%, cheese 4%, bone 4%). After attention most of the right set is present (mat 35%, white 29%, rug 21%). The MLP then only refines it (0.6 nats). Here attention does nearly all of the work.
  • Switch off position embeddings on "the black cat saw the" (KL 6.0 nats; "white" 58% becomes "roof" 77%) and then on "a dog sat on a" (KL 0.00). Order matters to some predictions and not to others.