A user reports high performance using EXL3 quants, specifically running the Muse Glimmer 30B model at 3.00bpw with a 100K context and Q8_O KV cache on a 12GB VRAM GPU. The setup achieves approximately 30 tokens per second with minimal perceived quality loss compared to larger K-quants. The user also noted that Qwen 3.8 27B remains usable at SC2.20bpw H3.

Read original