Swift-1.5 Qwen3.8-27B, an RL+OPD post‑trained agentic/coding focused model, has been quantized to all‑NVFP4 (W4A4 gs16) and combined with z‑lab DFlash2 drafter for the NInfer v3 engine. The artifact fits on a single RTX 5090 32 GB GPU, occupying ~18.42 GiB (≈18.0 GiB VRAM) and supports a full 262,144‑token context with KV‑cache auto‑pooling up to 308,736 tokens. Decoding achieves roughly 160 tokens per second.
Read original
reddit/r/LocalLLM