The author reports that Qwen 3.8 Flash Next doubled Strata throughput on an RTX 3090 + RTX 5070 Ti system, achieving IQ3_S performance of 2466 pp/167 tps and UD-Q4_K_XL of 2341 pp/126 tps. Through a week of profiling and benchmarking with Opus 5.5, token generation speed rose from ~6 tps to over 50 tps on IQ3_XXS and reached up to five times the baseline llama.cpp performance on UD-Q4_K_XL. These gains were obtained by refining llama.cpp quantization parameters after initial low‑speed results.

Read original