An optimized fork of llama.cpp for dual AMD Radeon RX 7900 XTX GPUs achieves ~82 tokens/second decoding Qwen 3.8 Q8 at a 60k token context, up from ~28 tokens/second with the vanilla Vulkan build. The repository (github.com/nasone32/llama.cpp-RDNA3-7900xtx-opt) provides RDNA‑3 specific kernels and build tweaks for the two‑GPU configuration. Performance gains were demonstrated on a consumer‑grade PC running Linux, with the setup configured by Luna.

Read original