The commit adds an int8 cooperative matrix (coopmat1) matrix multiplication kernel using Vulkan for AMD RDNA3 and RDNA4 GPUs. Benchmarks on a Radeon RX 7900 XTX show the Gemma 4 26B model achieving ~3410 tokens/s in pp512 and ~136 tokens/s in tg128 workloads, representing a substantial performance gain. The implementation improves inference throughput for quantized LLMs on AMD hardware.
Read original
reddit/r/LocalLLaMA