The author ran Qwen3.8-27B with a 100K token context window on an RTX 4080 (16 GB VRAM) using ExLlamaV3 1.5.0 via TabbyAPI, measuring VRAM consumption and generation speed. Results showed the model fits within the 16 GB limit while maintaining acceptable throughput, and enabling MTP (Mixture‑of‑Token‑Parallelism) provided a worthwhile trade‑off between extra VRAM use and performance. The experiment demonstrates that large‑context inference is feasible on consumer‑grade GPUs with appropriate optimizations.
Read original
reddit/r/LocalLLM