Nonobench v1.2 evaluated 43 language models on nonogram logic puzzles, introducing explicit reasoning‑effort tracking and fully public prompts/outputs. The open‑weight DeepSeek V4 Pro model tied for fourth place overall, while no open‑weight model succeeded at the newly added 20×20 Hard mode, which requires row‑by‑row answering to avoid miscounts.

Read original