The post announces two open‑weight models derived from Qwen 3.8‑Flash‑Next: Victoria, which reduces expert count by 44% (512→288 per layer) via REAP, is retrained in 4‑bit NVFP4 and scores 70.0% on Terminal‑Bench 2.1; and Maple, a Canada‑first fine‑tune of the same base. Both were trained on a Dell B300 workstation, with Victoria’s NVFP4 build improving over the prior 62.5% benchmark.
Read original
reddit/r/LocalLLaMA