Developer tmballin has released Qwen3.8-27B with genuine 131,072-token context, vision support, and MTP-3 speculative decoding running on a single RTX 5080 16GB via NInfer v1.3 at ~3.95 BPW. The model and inference engine are now publicly available on GitHub with a production release tagged for this configuration. This demonstrates significant memory optimization enabling large-context multimodal inference on consumer-grade 16GB VRAM hardware.

Read original