WHIRL is an open-source (Apache-2.0) native Windows inference engine for the AMD Radeon AI PRO R9700, written in pure C++ and HIP without requiring WSL, a Linux VM, or llama.cpp. It achieves up to 2.5× the prefill speed of llama.cpp, 107–328 tokens per second decode, and three times the server throughput when running a Qwen3.8‑27B model fine‑tuned in MXFP4.
Read original
reddit/r/LocalLLM