The author evaluated the locally‑run Qwen3.8‑Flash model on a DGX Spark against Claude Opus 5.5 using 21 graded tasks across three different harness configurations. Qwen3.8‑Flash matched Opus performance on everyday coding tasks and reached about 70 % of Opus’s score on harder tasks. The results showed that the choice of harness settings had a larger impact on outcomes than expected.
Read original
reddit/r/LocalLLM