Users report that running Strata with Qwen3.8 IQ3_XXS weights (~80 GB) on a DDR4‑based system paired with an AMD Radeon RX 7900 XTX yields stable generation speeds of 45–70 tokens per second, occasionally higher during coding workloads. This performance outperforms a tuned Llama CPP implementation on the same hardware, which peaks around 22.5 t/s, while delivering superior output quality. Systems equipped with DDR5 memory observe even faster rates, whereas lower‑bit Q2 weights are not recommended.

Read original