GPT‑6.1 Sol superseded GPT‑6 Sol just one week after its release, achieving an Intelligence Index score that is only one point below GPT‑6 Astra while operating at less than a quarter of Astra’s cost per task ($0.72 versus $3.26). The model retains the $2/$10 per‑million‑token pricing of its predecessor but improves the cache‑read discount from 90 % to 95 %, yielding a slightly lower blended price for agentic workloads. Relative to GPT‑6 Sol, GPT‑6.1 Sol gains four Intelligence Index points (five versus GPT‑5.6 Sol) and shows notable improvements in downstream benchmarks: +4 points in AA‑Briefcase v1.1, +5 points in GDPval‑AA v2.1, +12 points in Terminal‑Bench 4.0, +5 points in Humanity’s Last Exam, +6 points in GDP.pdf, and an +8‑point jump in AA‑Omniscience accuracy accompanied by a hallucination‑rate drop from 60 % to 54 %. Cost efficiency is further enhanced—31 % cheaper per task than GPT‑6 Sol ($1.05) and 64 % cheaper than GPT‑5.6 Sol ($1.99)—pushing the cost‑efficiency Pareto frontier outward. Token usage rises ~10‑30 % in output tokens versus GPT‑6 Sol, yet low and medium effort settings become Pareto‑optimal for token efficiency due to the intelligence gain. In the Coding Agent Index, GPT‑6.1 Sol adds three points over GPT‑6 Sol at max effort and, under the xhigh setting, exceeds GPT‑6 Astra by one point while consuming under 15 % of Astra’s cost per task.

Read original