A new controlled version of the CivBench benchmark evaluates large language models' ability to play full games of Civilization V. GLM‑5.3 outperforms Opus‑5.5, while Qwen‑3.8‑27B shows surprisingly strong performance. The benchmark is also being applied to newer models such as GPT‑6.1‑Sol and GPT‑6‑Astra.
Read original
reddit/r/LocalLLaMA