The author reports that running the uncensored Qwen 3.8‑27B model with NInfer’s NVFP4/groupwise‑int and MTP optimizations on an RTX 5090 (32 GB VRAM) yields ~175 tokens per second. This performance level finally makes the GPU feel useful for local LLM workloads after previous QoL upgrades.

Read original