The authors address the fragmented evaluation of LLM‑based agents for black‑box optimization (BBO) by creating AgenticBBO‑Bench, a unified finite‑budget benchmark that spans five heterogeneous domains: synthetic benchmark functions, hyperparameter optimization, database tuning, chip design, and molecular design. Under this protocol, agentic BBO pipelines consistently achieve higher family‑averaged performance than direct LLM optimization methods across all domains and surpass the best existing numerical optimizers in four of them. A systematic analysis isolates three key design influences: (1) the integration of numerical tools, which does not yield reliable gains; (2) task semantics and prior knowledge, where generic semantic cues improve results while domain‑specific priors prove less stable; and (3) the LLM’s role during search, showing that numerical optimizers can effectively incorporate agent‑generated trajectories. The work also introduces a five‑task frontier challenge, evaluating seven LLMs via the Codex agent harness; GPT‑6 Astra and DeepSeek‑V4.1‑Flash emerge on the Pareto frontier of performance versus computational cost. The benchmark and code are publicly released, enabling reproducible comparison of agentic BBO strategies.
Read original
huggingface/daily-papers