A $2,800 workstation equipped with eight Radeon Pro V620 GPUs (32 GB VRAM each, total 256 GB) and a custom vLLM fork runs the Qwen3.8‑Flash‑Next model at 60–100 tokens/s decode and over 3,000 tokens/s prefill. The cards are older RDNA2 enterprise cloud‑gaming hardware, offering high VRAM capacity at low cost.

Read original