The article demonstrates how a 125‑billion‑parameter Mixture‑of‑Experts (MoE) language model can be run on a single consumer‑grade RTX 4090 (24 GB VRAM) by offloading the majority of expert weights to system RAM while keeping attention, KV cache and shared components on the GPU. Using llama.cpp compiled with CUDA (‑DGGML_CUDA=ON, ‑DGGML_CUDA_F16=ON, ‑DCMAKE_CUDA_ARCHITECTURES=89) and a GGUF model quantized to IQ4_XS (≈10‑15 % smaller than Q4_K_M), the author sets ‑ngl 99 to place all layers on the GPU initially, then applies either ‑‑n‑cpu‑moe N to move the first N expert layers to CPU RAM or ‑‑ot regex rules to fine‑tune placement. Flash Attention (‑fa on) and 8‑bit KV cache quantization (‑‑cache‑type‑k q8_0 ‑‑cache‑type‑v q8_0) halve the VRAM needed for context handling. With a Ryzen 9 7950X, 128 GB DDR5‑5600 RAM (dual‑channel recommended) and the model’s active parameter count per token dropping to roughly 10‑15 B, the configuration achieves usable token‑per‑second speeds after tuning: starting with ‑‑n‑cpu‑moe equal to the total layer count, VRAM usage is monitored via nvidia‑smi and reduced until ~22‑23 GB is consumed, leaving a 1‑2 GB headroom. Prefill performance is improved by increasing the batch size (‑‑ub) for long contexts, and throughput is measured with llama‑bench and a custom streaming script that reports TTFT and decode tok/s. The guide warns against swap usage, notes WSL2 memory limits, and advises beginning with ‑‑n‑cpu‑moe before moving to ‑‑ot for deeper tuning.
Read original
dev.to