On a specific real‑code test, a local 27B parameter model scored 98.0% correct, exceeding the average 96.6% of frontier cloud models. While frontier models achieve perfect scores on this task only about two‑thirds of the time, the local model’s result shows the performance gap is narrower than often claimed. Across the broader 113‑task benchmark, leading cloud models still lag in the 70‑74% range.

Read original