PolyStrata extends the Strata inference engine with four additional mixture-of‑experts models—GLM 5.3 Flash, Qwen3.6‑35B‑A3B, Ornith 1.5, and Gemma 4 26B‑A4B—each loaded from standard GGUF files. The Qwen3.6‑35B‑A3B model achieves 73‑75 tokens per second on a laptop equipped with an 8 GB GPU. The repository retains the original Strata engine for Qwen3.8‑Flash‑Next.

Read original