The author criticizes current model cards for omitting KV cache memory costs, pointing out that some small models require ~64 KB per token, which severely restricts context length despite fitting GPU memory for weights. They highlight Ling-3.0‑Tiny as an example where low KV cache usage enables >600 k tokens of context on a 12 GB GPU. The post expresses interest in evaluating K2‑Horizon‑7B under the same criteria.
Read original
reddit/r/LocalLLM