The user runs a Qwen 3.8‑27B model quantized to Q3 on a dual‑GPU setup (5060 Ti 16 GB + 3060 Ti 8 GB), achieving ~40 TPS generation with a 140k token context. They find the context window too limiting for complex agentic coding tasks and wish to extend it to 200k+ tokens using Qwen 3.8 Flash Next. They are evaluating either upgrading their current PC via Strata or swapping to a Strix Halo GPU to achieve the larger context.

Read original