A one-layer, two-head transformer trained on a 28-word toy corpus. Every number on this page is a live forward pass through those weights, run in your browser. Trained offline with hand-written backprop (gradient check: 7.9e-7 relative error) to a loss of 0.496 nats, against 3.332 for a uniform guess. The JavaScript forward pass was checked against a separate numpy implementation and agrees to 5e-7.
Each word indexes a row of a learned table of 16-number vectors. A position vector is added so the model can tell order. After this step the words are gone; the model holds only vectors.
Every word is layer-normalised, then projected to a query, a key and a value, 8 numbers each per head. A query's dot product with an earlier word's key is a relevance score. Softmax turns the scores into weights, and the weights mix the values. Click any cell for the arithmetic.
Each position passes alone through a small network: layer-norm, expand to 32 units, GELU, project back, add to the stream. Attention moves information between words; this step transforms it within one.
The final vector is multiplied by the unembedding matrix to give one score per word. Softmax makes probabilities. Temperature, top-k and top-p reshape them before a word is drawn.
The same unembedding applied to the stream after each stage, at the selected position. It shows what the model would say if it stopped there.
Everything above is recomputed with these knobs, and generation uses them too. The comparison shows the untouched model next to your settings at the selected position, at τ = 1.
Each head's attention pattern depends on WQK = WQ,hWK,hT and what it writes on WOV = WV,hWO,h. Both are 16×16 but have rank at most 8. The bars are their singular values, computed by SVD of the trained weights.