A Reddit user benchmarked the Strata and Infernix inference engines on the Qwen Flash Next model using an RTX 5090 with 192 GB RAM on Windows. At a 512k token context, Strata achieved ~137 tokens/s generation and Infernix reached ~3,497 tokens/s prefill speed. The tests show both engines run smoothly on Windows with comparable performance across 262k and 512k context lengths.

Read original