Opera is a verbal critic framework designed for long‑horizon coding agents that treats each correction as a persistent note and monitors it until the underlying problem is solved. It determines when to intervene using both periodic and event‑driven triggers, diagnoses coding issues with typed operators, and audits the generated feedback against visible evidence before delivering it to the agent. After feedback is given, Opera tracks the agent’s subsequent actions to differentiate superficial compliance from genuine resolution. As a test‑time critic, Opera boosts the resolve rate of baseline non‑critic agents by up to 12.4 percentage points on Terminal‑Bench 2.1, 15.0 points on a SWE‑Bench Pro subset, and 8.9 points on DeepSWE v1.1, across four policy models, and attains the highest mean resolve rate among competing critic baselines on all three benchmarks. It also enhances performance when the policy itself acts as a critic. In addition, Opera‑guided rollouts serve as on‑policy training data: fine‑tuning Qwen3.5‑9B on these rollouts raises its resolve rate on held‑out SWE‑Bench Pro repositories by 10.2 percentage points without requiring a critic at inference time, matching the gain obtained from fine‑tuning on rollouts from a stronger model, while retaining performance when switching execution harnesses (e.g., from OpenHands to Terminus‑2), a transition that otherwise causes substantial degradation.
Read original
huggingface/daily-papers