A new pull request for llama.cpp introduces a GPU cache for mixture-of-experts (MoE) model experts that are stored in host memory. This optimization can yield substantial speed‑ups for MoE models that do not fully fit into VRAM, benefiting users with limited GPU resources. The change is part of PR #29887 in the ggml-org/llama.cpp repository.
Read original
reddit/r/LocalLLaMA