A user discovered that upgrading to a specific llama.cpp build (b10270) doubled DeepSeek V4 Flash inference speed from 11 to 25.9 tok/s on an M4 Max, with the performance gap traced to the build itself rather than the model, prompt, or hardware.

Read original