Language enters as tokens, moves through a high-dimensional system, becomes a probability distribution, and is reshaped by training. Twenty interactive scenes make that invisible structure visible.
Exact scenes compute the real mathematics. Conceptual scenes illustrate documented effects. Each scene pairs the picture with a compact reading, its limits, and the connection to what follows.
drag to orbit · arrow keys · about & provenance
The complete atlas
All twenty scenes
Choose any idea. The chapters form a useful path, but every scene stands on its own.
LLM Geometry — about & provenance
How to read these pictures
Language models live in spaces with thousands of dimensions; these scenes are three-dimensional stories about that geometry. Each one wears a badge. Exact means the picture is computed from the real mathematics: points sampled from the true radial law of a high-dimensional ball, the fitted Chinchilla loss surface with its published constants, a measured packing of 42 feature directions, temperature and top-p applied to a real distribution. Conceptual means the surface is drawn to illustrate a documented effect, and any dynamics on it — gradient descent, kernel fits, gating, basin-hopping — are computed on the drawn surface.
Method receipts
The categorical palette (sky · amber · rose · violet on near-black) was validated programmatically for colorblind-safe separation (worst all-pairs CVD ΔE 13.8, normal-vision 18.7, contrast ≥ 3:1 on the surface); every colored element also carries a direct label. The 42-line packing converges to θmin ≈ 21.6° from every random seed. Shell fractions, median radii, and the compute-optimal frontier are computed live in this page, not drawn by hand.
Sources
Hoffmann et al. 2022, Training Compute-Optimal Large Language Models — the fitted loss surface (E = 1.69, A = 406.4, B = 410.7, α = 0.34, β = 0.28); Kaplan et al. 2020 for the earlier scaling picture; Schaeffer et al. 2023, Are Emergent Abilities a Mirage?
Elhage et al. 2022, Toy Models of Superposition; Bricken et al. 2023; Templeton et al. 2024, Scaling Monosemanticity — superposition and sparse-autoencoder features. Johnson & Lindenstrauss 1984.
Liu et al. 2023, Lost in the Middle, and 2024–25 long-context degradation reports — the context-volume effects.
Lambert et al. 2024 (Tülu 3) for RLVR; DeepSeek-R1 (2025) for its scale-up. Holtzman et al. 2019 and Meister et al. 2022 for typicality in decoding.
Dell'Acqua et al. 2023, Navigating the Jagged Technological Frontier; Mollick, Co-Intelligence (2024). Li et al. 2018; Aghajanyan et al. 2020; Hu et al. 2021 (LoRA); Garipov et al. 2018; Frankle et al. 2020; Wortsman et al. 2022 (soups).
Sennrich et al. 2016 (BPE tokenization). Yao et al. 2023 (Tree of Thoughts); Lightman et al. 2023 (process verification); Snell et al. 2024 (test-time compute); Turpin et al. 2023 (chain-of-thought faithfulness). Vaswani et al. 2017; Beltagy et al. 2020 (sliding-window attention).
Brown et al. 2020; Xie et al. 2022; von Oswald et al. 2023; Akyürek et al. 2023 (in-context learning); Sun et al. 2024 (test-time training). Carlini et al. 2023 (memorization); Kandpal et al. 2023 (long-tail facts); Lewis et al. 2020 (retrieval-augmented generation).
Power et al. 2022 (grokking); Nanda et al. 2023 (Fourier circuits in modular arithmetic). Elhage et al. 2021 (residual stream); Turner et al. 2023 (activation steering); Belrose et al. 2023 (tuned lens). Shazeer et al. 2017; Fedus et al. 2021; Jiang et al. 2024 (mixture-of-experts). Dettmers et al. 2022; Frantar et al. 2022 (quantization). Bengio et al. 2013 (manifold hypothesis).
Kirk et al. 2024 (RLHF and output diversity); Doshi & Hauser 2024 and Padmakumar & He 2024 (homogenization); diversity collapse under RLVR: arXiv 2509.07430 (2025), arXiv 2606.15455 (2026). Kobak et al. 2025, Science Advances (excess vocabulary); Juzek & Ward 2025, COLING (Why Does ChatGPT "Delve" So Much?); Mitchell et al. 2023 (DetectGPT); Shumailov et al. 2024, Nature (model collapse); Korbak et al. 2022 (KL-regularized RL as distribution tilting).
Colophon
A single self-contained HTML file: three.js r128, Newsreader & Inter subset and inlined, no trackers, no build step. The success-vs-process framing of RLVR follows Nate Jones. August 2026.