A user reports improving Qwen3.8-Flash-Next decode speed from 25–29 t/s to 37–41 t/s on a dual RTX 3090 + DDR4 system by combining UD-Q4_K_XL quantization, expert cache, and
reddit/r/LocalLLM
A user reports improving Qwen3.8-Flash-Next decode speed from 25–29 t/s to 37–41 t/s on a dual RTX 3090 + DDR4 system by combining UD-Q4_K_XL quantization, expert cache, and