Qwen 3.8 27b harness

u/Dingydongy007 2026-08-30

The poster reports that multiple local deployments of the Qwen 3.8 27B model, including Opencode, qwen CLI, and Claude CLI, have yielded unsatisfactory coding performance, indicating difficulty achieving useful results. …

→ View original source
Don't Sleep on EXL3 Quants

Don't Sleep on EXL3 Quants

u/PyaesoneP 2026-08-30

A user reports high performance using EXL3 quants, specifically running the Muse Glimmer 30B model at 3.00bpw with a 100K context and Q8_O KV cache on a 12GB VRAM GPU. The setup achieves approximately 30 tokens per secon…

→ View original source
Loading more articles...