The article introduces InFlowOp, a label-free framework for optimizing multi-agent workflows in large language models (LLMs). Traditional workflow construction faces challenges in determining task granularity, agent assignment, and fault correction, often requiring costly re-execution or retraining with labeled data. InFlowOp addresses this by assigning a unified cost metric that evaluates agent competence against subtask demands relative to runtime efficiency. This cost guides bidirectional decisions on task decomposition and agent selection prior to execution, while enabling fault correction during runtime via the same cost function. To benchmark multi-agent coordination beyond single-agent capabilities, the authors propose Braid, a novel evaluation suite. Across multiple domains and model backbones, InFlowOp achieves up to +11.97% improvement over single-agent baselines, with +9.64% gains attributed to in-flow optimization. The approach eliminates dependency on external labels or assessors, offering a scalable solution for dynamic workflow refinement in complex AI systems.

Read original