UkisAI team member Jovan proposes creating a "Swift" version of Bonsai 2 27B, applying the same token usage and overthinking error improvements they previously applied to Qwen3.8 27B. The team's testing indicates Bonsai 2 suffers significantly from overthinking loops and high token usage. They are asking the r/LocalLLaMA community whether such a model would be welcome and what quantization size (e.g., 1-bit, 2-bit) would be most relevant.
Read original
reddit/r/LocalLLaMA