After identifying instability in their long‑running evaluation workflow, the author re‑ran a swe‑verified Django 100‑task benchmark across local models and quantizations. The results confirm Flash Next as the leading model, while the performance gap between xhigh and medium reasoning effort is now correctly reflected. Similar clarification appears for the 3.8 B 27B variant, though everyday‑task nuances were noted.

Read original