WHIRL v0.1.3 is an open‑source (Apache‑2.0) native Windows inference engine for the AMD Radeon AI PRO R9700 (RDNA 4, 32 GB) implemented in pure C++/HIP, requiring only the AMD driver. It outperforms llama.cpp b11214 on identical GGUF models, achieving up to 2.8× higher throughput (e.g., 182 tok/s vs 64 tok/s for a 4‑user server) and improved decode/prefill speeds on coding prompts and long contexts.
Read original
reddit/r/LocalLLM