The Gufo project's inference software for Qwen 3.8 Flash Next on Strix Halo delivers markedly higher performance than llama.cpp, especially at large context lengths. Benchmarks show 6204 chunks processed in 119 seconds with an encode rate of 1 239 tokens/s and a decode rate of 57 tokens/s, and the PP speed is roughly double that of the fastest Strix Halo‑specific llama.cpp fork.

Read original