The user is evaluating which 16 GB VRAM Qwen3.8 27B variant delivers the best speed, maximum context length, and output quality for local use. They currently run an Unsloth IQ4_XS quant with 65 k Q8 KV context, achieving roughly 30 tokens/s but find it slow for agentic workloads, and are considering whether lower‑bit Q3 quants are safe enough for broader applications.
Read original
reddit/r/LocalLLaMA