The Qwen 3.8 Flash Next‑GSQ‑RCO‑IQ2_XS model achieves roughly 21 tokens per second on an RTX 3060 12GB with 16 GB DDR4 RAM, reaching 24+ tok/s with a warm cache, without gate pruning and while preserving 100 % bit‑exact output. The author had earlier explored predicting the next MoE expert to speed up CPU/GPU offloading but later shelved that project.
Read original
reddit/r/LocalLLaMA