The author deployed two repurposed BC-250 mining APUs, each providing ~13.5 GB GPU memory, to run Qwen3‑6‑35B‑A3B Q4_K_M via llama.cpp with Vulkan and RPC on Bazzite, achieving ~60 token/s with a 64k context window over 1 GbE. The total hardware cost is about $300 including PSU, and they plan to scale to six units to test Qwen 3.8‑Flash.
Read original
reddit/r/LocalLLaMA