A 10‑person startup built a workstation around a Threadripper 9970X CPU, 128 GB DDR5 ECC RAM, and four AMD Radeon R9700 AI Pro GPUs (32 GB each) powered by a 1600 W supply, with each GPU undervolted to under 210 W. The system runs a fork of Radiance to serve Qwen 3.8 27B and DSV4‑Flash models, supporting 16 concurrent sessions with ~6.5 k token/s prefill and 900 token/s aggregate decode (≈80 token/s per session) across 4‑128 k context lengths. Performance figures are for MXFP4‑quantized Qwen 3.8 27B, showing 6.3‑6.8k prefill tok/s and the stated decode throughput.

Read original