An interactive animation visualizes how tokens propagate through a large language model during inference, illustrating the internal mechanics of dense, mixture‑of‑experts, and linear architectures. The visualization, hosted at sambatista.com, includes an FAQ and allows users to switch between model types to observe differences in token flow.
Read original
reddit/r/LocalLLM