DwarfStar 4 (ds4) is a purpose‑built C inference engine that runs large mixture‑of‑experts language models locally on high‑memory macOS (Metal), CUDA and ROCm systems. It employs asymmetric 2‑bit quantization to compress the routed experts of supported models while keeping the critical shared paths in higher precision, enabling DeepSeek V4/V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next to fit within 128 GB of RAM (with V4.1 Q2 streaming from SSD). The engine provides three interfaces—./ds4 for interactive chat, ./ds4‑server exposing OpenAI‑ and Anthropic‑compatible HTTP APIs, and ./ds4‑agent for persistent coding sessions—all sharing a single model state and KV cache. To avoid costly re‑prefills, long prefixes are cached to SSD and resumed via prompt hash lookup. Benchmarks on an M5 Max (128 GB) show 790.2 tokens/s prefill and 39.4 tokens/s generation for 2 k‑token contexts, dropping to 398.5 t/s prefill and 27.6 t/s generation at 65 k tokens; comparable numbers on a DGX Spark (128 GB) are 825.8/18.1 t/s and 823.0/13.8 t/s respectively. The project follows a strict validation flow: download the official GGUF weights, build ds4 for the chosen backend, then launch the CLI, server or agent, bypassing generic GGUF runners in favor of a tightly coupled, end‑to‑end verified stack.

Read original