The user runs Qwen 3.6 35B on a modest laptop (16 GB DDR5, 13th‑gen i5, no GPU) via llama.cpp at ~15 output tokens per second. They desire a smarter Mixture‑of‑Experts model that is larger than Qwen 3.6 35B but smaller than GLM 4.5. However, they have found no publicly available models that fit this intermediate size range.

Read original