medium

Your KV Cache Is Bigger Than Your Model

Satsawat Natakarnkitkul (Net) 2026-08-16

The KV cache used during inference can consume significantly more GPU memory than the model parameters themselves, with costs reaching 4.5 GiB at 128K context on gpt-oss-120b and up to 40 GiB in dense configurations. Thi…

→ View original source