llama.cpp version 0.6.0 has been released, introducing MTP (mid‑tower prediction) speculative decoding for the Qwen4Exp model alongside various other enhancements. The update improves inference speed and efficiency while maintaining compatibility with existing quantized models. Users can now leverage the new speculative decoding feature to achieve lower latency on supported hardware.

Read original