Strix Llama provides a patched llama.cpp runtime and Windows installer for running the Qwen3.8‑Flash‑Next 125B MoE model (~6B active) on an AMD Strix Halo (Ryzen AI Max+ 395) with 128 GB RAM. The setup delivers 36–45 token/s generation and ~1200 token/s prefill using speculative decoding, with no Linux, ROCm, or build‑tool dependencies. Version 0.3.0 is the stable release.

Read original