The user configured a server with four MI100 GPUs (32 GB each) interconnected via Infinity Fabric and forked vLLM to enable native support, achieving over 1000 tokens/s decode and over 5000 tokens/s prefill at 8‑way concurrency for a 27 B‑parameter model. This performance is a dramatic increase from the stock vLLM baseline of ~15 tokens/s on the same hardware. The optimization required developing numerous new kernels and applying advanced performance‑tuning techniques.

Read original