On an Apple M5 Max, ik_llama.cpp achieved a 20.8 % overall speedup over llama.cpp for the Qwen3.6‑35B‑A3B‑MTP model. Its gains were most pronounced in text generation, with a 31.4 % throughput increase, followed by 16.0 % faster vision processing and 14.1 % faster post‑vision generation. These results confirm that ik_llama.cpp delivers superior CPU/ARM performance for multimodal workloads on Apple silicon.
→ View original source
reddit/r/LocalLLM