A pull request by wjinxu adds DSpark speculative decoding support to llama.cpp (PR #25173). DSpark leverages multiple draft models for speculative decoding, with reference implementations including DeepSeek-V4-Pro-DSpark and Bonsai AntiDoom-1bit-DSpark. Community members are encouraged to benchmark and share throughput (pp/tg) improvements when using the feature.

→ View original source