Running Qwen 3.8 27B with NVFP4 quantization on a dual RTX 5060 Ti 16 GB setup (32 GB VRAM) achieves ~130 tokens/s while using FP8 KV cache to support the full 262,144‑token context with ~10 k tokens of headroom. The model runs on stock vLLM 0.30.0 under headless Ubuntu, with GPUs power‑capped at 150 W, staying cool and silent.

Read original