The R9V update achieves ~100 tokens per second text generation on Qwen3.8 Flash Next IQ4_XS using dual R9700 GPUs with 128 GB RAM, while fixing crashes caused by n‑gram SSD streaming and adding pinned‑image support. Q4_K_XL quantization is now supported, delivering ~50 tokens per second. The update was stress‑tested for ~12 hours with no observed instability.
Read original
reddit/r/LocalLLaMA