A developer achieved 310 tokens per second (t/s) inference speed running Qwen/Qwen3.8-Flash-Next-FP8 on a system with four RTX PRO 6000 GPUs. The performance was described as exceptionally fast, particularly for coding tasks, with the author stating it surpassed their previous experience with other models. The setup demonstrates significant gains in local LLM inference efficiency using FP8 quantization and multi-GPU scaling.
Read original
reddit/r/LocalLLM