The newly released Bonsai 2 27B model is a heavily quantized version of Qwen 3.8 27B that retains over 98 % top‑1 accuracy of its FP16 counterpart while consuming only 6–8 GB of VRAM. It was evaluated on an RTX 5090 alongside Gemma 4 12B and Qwen 3.5 9B using an identical Japanese voxel pagoda prompt, with all tests run in a one‑shot setting. The Bonsai model exhibited notably higher token usage compared to the smaller Gemma and Qwen counterparts.
Read original
reddit/r/LocalLLM