Over a 30‑day period, the Unsloth Qwen3.8‑27B‑UD‑Q4_K_XL model was deployed locally as a coding agent and for extended overnight runs, achieving mean prompt‑processing throughput of 845.1 tokens/s and mean token‑generation throughput of 73.8 tokens/s. The model’s MTP acceptance rate was 0.481 (674/1401), indicating stable performance across varied workloads. Based on these results, the author considers it the first model they would recommend for production‑scale local LLM services.

Read original