GPT-2 Logit Lens

Logit lens · GPT-2 small · 12 layersWhat the model knows, and when

Residual stream — the running prediction MLP write — this layer's own contribution
Layer Residual streamMLP write

Layer 11 detail

p(emitted token)—
Residual entropy—
MLP entropy—

Belief in the emitted token

How much the residual stream backs the token actually emitted, layer by layer.

Residual entropy (nats)

Lower means more committed. A late rise means the model hedged after having decided.

Each layer's vector is projected through the model's own final LayerNorm and unembedding, turning any point mid-stack into a distribution over the vocabulary. Layer 11's residual equals the model's real output — which is the check that the projection is being done correctly.

Play your own prompt

GPT-2 doesn't run in this page — the traces are precomputed. Generate your own with the exporter, then paste the JSON here.

python export_lens_json.py --out lens.json --prompts "your prompt here"