The Radeon AI Pro R9700 running Qwen3.8-27B Q8 achieved a token rate of 90.8 tokens/s, as reported by a user on r/LocalLLM. This performance was enabled by the RDNA Boosts repository, with the user's own vLLM benchmarks slightly exceeding the result and showing practical daily usability. For details, see the original Reddit post and the linked boosts repository.
Read original
reddit/r/LocalLLM