The pull request #29761 adds MTP support to Qwen Flash Next in the ggml-org/llama.cpp repository, allowing users to employ Mixture‑of‑Tokens with the model. Developed and merged in about 17 hours, the update follows prior work on Qwen Flash Next MTP. Quantized versions are available via the HuggingFace repo ggml-org/Qwen3.8-Flash-Next-GGUF.

Read original