The user runs Qwen 3.8 27B (UD-IQ3_XXS) on an AMD Ryzen 9950 X3D CPU, Nvidia RTX 5060 Ti 16 GB GPU, and 64 GB DDR5 RAM via Unsloth Studio on Windows 11, achieving ~50 tokens / second with a 32 k token context. They are looking for a mixture‑of‑experts (MoE) model that matches or improves this speed while preserving quality for agentic applications, mentioning interest in the 35B‑A3B variant.
Read original
reddit/r/LocalLLM