Uniform GGUF quantization (Q4_K_XL, Q6_K_XL) of the hybrid Qwen3.8‑27B model causes deep‑thinking failures on long tasks, manifesting as non‑convergent generation, endless tail‑looping, or engine crashes, while short‑form chat remains unaffected. The issue was reproduced independently on both llama.cpp and vLLM (with the official vllm‑gguf plugin), indicating a runtime‑agnostic problem with uniform quantisation of the model’s Gated DeltaNet linear‑attention + full‑attention architecture.
Read original
reddit/r/LocalLLM