Strata demonstrates that Mixture‑of‑Experts models can run using only the combined VRAM and RAM of a modest workstation, achieving real‑time inference on a system with four 16 GB P100 GPUs and 128 GB DDR4 RAM. With Flash‑Next Q4 quantization and a 262 k token context, the setup delivers roughly double the speed of a 27 B parameter model and five times the throughput of Flash‑Next when using a customized llama.cpp build. The hardware comprises an Intel Xeon E5‑2683 v4 CPU, 64 GB total GPU memory, and 128 GB system memory.

Read original