The Qwen3.8‑Flash‑Next model, running on a 12 GB RTX 5070 with 64 GB DDR5 and a Ryzen 5 7600, now achieves ~65 tokens/s output and ~430 tokens/s prompt processing using IQ3_XXS quantization in a custom inference engine. Further speed gains are observed with 2‑bit RCO‑GSQ quantizations. These results improve upon earlier llama.cpp benchmarks of ~15 tokens/s output and 100‑120 tokens/s prompt processing on the same hardware.

Read original