RobotWorld introduces a simulation testbed comprising 84 distinct tasks that cover manipulation, mobile manipulation, locomotion, driving, and aerial control, each equipped with explicit interaction budgets and executable success checks to evaluate multimodal agents’ ability to translate language instructions and sensor observations into physical robot actions. The benchmark records both task outcomes and detailed execution traces, enabling analysis of how agents construct perception‑and‑control pipelines such as image segmentation, camera calibration, spatial estimation, and dynamics‑based computation. Results show that while agents can generate sophisticated workflows, they frequently fail to translate these capabilities into reliable behavior: they lose track of task‑relevant object states despite achieving commanded poses, do not correct ineffective actions, recover from errors too late, or mistakenly declare unfinished tasks complete. Performance varies across models; Astra exhibits higher success on spatial and constrained‑contact objectives, whereas Opus 5.5 performs better on continuous‑balance and timed‑interaction challenges. By linking success patterns to specific execution failures, RobotWorld provides a rigorous proving ground and an empirical map of capability gaps, highlighting concrete targets for improving the reliability of general‑purpose agents in embodied, physical‑world settings.
Read original
huggingface/daily-papers