A team of 20 successfully self-hosted the Ornith-1.5-35B-A3B mixture‑of‑experts model on a single NVIDIA DGX Spark workstation with 128 GB unified memory. Despite its 35 B total parameters, the model activates only ~3 B parameters per inference, keeping compute demands low enough for one desktop box. The deployment has been running for about a week, demonstrating that a single machine can serve an entire team with decent output quality.

Read original