Researchers show that most of the KV cache for Qwen3.8‑Flash‑Next can be moved to system RAM with minimal impact on decoding speed, allowing models that barely fit in VRAM to run at full context length without KV‑cache quantization. Implemented in vLLM on three RTX 3090 GPUs, the setup achieves ~1 million‑token context, delivering ~80 tokens/s at short context and ~60 tokens/s after the QSA budget of 2048 tokens is reached, after which speed remains stable.
Read original
reddit/r/LocalLLaMA