Using two AMD Radeon AI PRO R9700 GPUs, the second card placed in a chipset‑connected PCIe slot lacked atomic operations, causing RCCL failures and crippling tensor parallelism. Relocating the GPU to CPU lanes via an inexpensive M.2‑to‑PCIe riser and switching from llama.cpp to vLLM raised inference speed for Qwen3.8‑27B from ~31 tokens/s on Windows to 72–113 tokens/s on Linux.

Read original