Researchers packed two ~300B‑parameter mixture‑of‑experts models (GLM‑5.3‑Flash and MiMo‑V2.6‑Flash) onto a single AMD Strix Halo‑based 128 GB mini PC using the ROCm‑based Kyojin engine built on ExLlamaV3. GLM‑5.3‑Flash reaches about 580 tokens/s prefill (3.5 K context) and 26‑30 tokens/s decode, while MiMo‑V2.6‑Flash achieves up to 44 tokens/s speculative decode (code) and ~650 tokens/s prefill at 4 K. The models use EXL3 quantization, showing KLD divergences of 0.151 (GLM) and 0 (MiMo) relative to official FP8 references.

Read original