FP8 quantization boosted serving throughput for Qwen3-8B on an RTX PRO 6000 Blackwell GPU from 1,725 to 2,597 tokens per second at concurrency 32 using vLLM, achieving