PSA: llama.cpp now loads MTP tensors by default for any draft-mtp arch, even with MTP disabled
Recent llama.cpp builds now load MTP/NextN tensors by default for draft-mtp architectures (e.g., GLM-5.2, hy_v3, qwen35moe, step35) even when speculative decoding is disabled via --spec-type. Previously these tensors wer…
→ View original source