A pull request #29928 to ggml-org/llama.cpp adds support for GLM 5 Flash MTP (Multi‑Token Prediction) and includes optimizations contributed by pwilkin. The change enables local use of GLM 5 Flash MTP models.
Read original
reddit/r/LocalLLaMA