UndoBench introduces a benchmark that isolates task competence from recovery capability in tool‑using AI agents by constructing 36 base workflows paired with 36 fault scenarios spanning eight enterprise domains. The design employs counterfactual paired trials executed under identical random seeds, complemented by wire‑level effect‑history and environment‑state oracles to observe exact side‑effects and state changes. Evaluation on 12 held‑out TEST workflows involved two open‑weight models, two agent frameworks, and three recovery strategies, yielding 5,760 individual executions (2,880 paired trials). Nominal task completion averaged 83.54%, yet the conditional recovery success rate (CRSR) was only 46.72%; naive retry caused duplicate external effects in 53.33% of trials. Extending the study to commercial API models reproduced the same competence‑recovery gap. Phase‑dependent analysis revealed that before any mutation, all methods behaved similarly without duplicate effects; during partial mutation, naive retry, per‑call idempotency, and zero‑privilege journaling failed, producing unsafe outcomes; however, after a commit but prior to acknowledgment, verification mechanisms and server‑side idempotency markedly improved safety. These results demonstrate that measuring nominal completion alone obscures critical, stage‑specific recovery vulnerabilities in autonomous agents.

Read original