The author built a local inference machine around two Huawei Atlas 300I Duo Ascend cards, each offering 96 GB of device memory without CUDA compatibility. Initial attempts to run Qwen3.8‑Flash‑Next were incoherent and limited to roughly one token per second. After further optimization, the same dual‑card setup now yields coherent output with substantially higher token generation rates. The post details hardware notes, vLLM integration, benchmark results, and plans for future scaling.

Read original