Your KV Cache Is Bigger Than Your Model
The KV cache used during inference can consume significantly more GPU memory than the model parameters themselves, with costs reaching 4.5 GiB at 128K context on gpt-oss-120b and up to 40 GiB in dense configurations. Thi…
→ View original source