Adding a negative logit bias (‑2) to the token IDs for “wait”, “maybe”, and “perhaps” in Qwen models reduces hesitant language generation. When tested on 50 randomly selected MATH‑500 problems across multiple GGUF quantizations of Qwen‑Qwen3.5‑4B, this adjustment yielded higher answer accuracy. The technique complements the recent Meta paper by extending its findings to quantized models supported by llama.cpp.
Read original
reddit/r/LocalLLaMA