Qwen-Planner-Agent introduces a closed‑loop AI‑for‑AI framework that treats the planner agent itself as both the subject and the engine of development. The framework consists of three coupled modules. First, an AI‑driven data flywheel uses specialized agents to author mobile‑planning tasks, collect interaction trajectories on real devices, curate and balance the resulting dataset, and then employs training‑feedback signals to steer the next round of data generation, all under a human‑gate to ensure quality. Second, training begins with a supervised planning cold start and proceeds with hybrid‑environment online agentic reinforcement learning; the authors add Competence‑Aware Reward‑and‑Advantage Engineering (CARE) to shape rewards so that reasoning and tool‑use steps are cheaper without sacrificing task success. Third, at runtime an execution‑evidence loop orchestrates memory, skills, and tools, captures structured action feedback and preserved failure traces, and feeds this information back to jointly adapt the model and its harness. Evaluated on MobilePA‑Bench, Qwen‑Planner‑Agent outperforms all baseline models and systems, showing gains over its base model in tool use, memory utilization, skill execution, and sub‑agent coordination, while preserving general capabilities on non‑mobile agentic benchmarks.
Read original
huggingface/daily-papers