Splash 1.1.0 introduces GGUF quantization support, MLX import, optimized kernels, speculative decoding, prefix cache, and mixed‑weight handling for efficient local LLM inference on Apple Silicon. On an M5 Pro with 64 GB RAM, the release enables ~50 tokens/s with the Qwen3.8 27B model using Unsloth UD‑Q4_K_XL quantization. These improvements make Splash a notable breakthrough for running large models locally.

Read original