The latest pull request to ggml-org/llama.cpp introduces the Maple 20B-A1B ternary mixture-of-experts (MoE) architecture optimized for CPU inference, aiming to reduce VRAM requirements. This addition enables running the 20‑billion‑parameter model on systems with limited graphics memory. A preview of the model is available on Hugging Face at deepgrove/maple-preview.
Read original
reddit/r/LocalLLaMA