A user discovered that upgrading to a specific llama.cpp build (b10270) doubled DeepSeek V4 Flash inference speed from 11 to 25.9 tok/s on an M4 Max, with the performance gap traced to the build itself rather than the model, prompt, or hardware.
Read original
reddit/r/LocalLLM