The Agent Error Dataset (AED) introduces 50,228 error-diagnosis pairs derived from 9,961 source tasks across 33 environments, 19 harness families, and 23 policy models in text-based agent systems. Developed to enable failure analysis and error-aware post-training, AED retains source traces and execution metadata for cross-setting failure analysis without requiring repeated rollouts. The five-stage Agentic Error-to-Training (AET) pipeline systematically collects natural failures, generates diagnoses and proposed corrections, validates them against evidence, and constructs training views for diagnosis and actor recovery. In replay-supported scenarios, first-proposal corrections improved verifier pass rates from 18.4% to 51.1% (32.7 percentage points gain) across 3,062 matched replay pairs. Full-diagnosis fine-tuning on 1,656 source tasks increased Qwen3-8B’s exact-step agreement with internal teacher labels from 47.2% to 63.6% (averaged over three seeds on a 943-case holdout), outperforming the strongest prompted reference (54.7%) and showing incremental gains with larger training sets. In actor-training comparisons, action-only repair training achieved a 6.67 percentage point higher success rate than success-only training on WebShop-lite. These results highlight AED’s utility in enhancing agent robustness through targeted error correction and diagnosis.
Read original
huggingface/daily-papers