The Strata inference engine reaches roughly 50 tokens/s TG and 1500 tokens/s PP on Qwen‑3.8 GGUF models using a laptop equipped with an RTX 5070 Ti (12 GB VRAM), 64 GB DDR5 RAM and an Intel 275HX CPU. Early KV‑cache and CPU‑throttling issues have been fixed, and the engine currently supports only Nvidia GPUs (AMD support is experimental).

Read original