DeepSeek’s recent KV‑cache optimizations have slashed the memory footprint of long‑context inference by up to 437× relative to its V1 baseline, achieving a global cache size of roughly 890 bytes per token in the V4.1‑Flash release. The progression began with the MLA architecture, which compressed the cache about 15×, followed by Compressed Sparse Attention and Heavily Compressed Attention techniques. V4.1‑Flash adds CSA2, cross‑layer cache reuse, a causal encoder‑decoder layout, and FP4‑precision caching, together driving the dramatic reduction. These advances directly lower the VRAM required to hold the cache during serving, cutting inference costs. Adoption by Western labs is evident in pricing changes: Claude Opus 5.5 reduces cache‑read costs by 60 % versus Opus 5, while GPT‑6.1 Sol achieves an 80 % reduction versus the late‑July GPT‑5.6 Sol rate. The resulting cost per million tokens, shown in the accompanying graph, reflects a shift from expensive, GPU‑memory‑bound workloads to far cheaper, long‑session inference, enabling models such as Claude Opus 5.5 and GPT‑6.1 Sol to deliver flagship‑level quality with markedly lower operational expense.

Read original