The paper defines "taste" as the ability of LLM agents to make good long‑horizon decisions during multi‑step tasks, such as choosing hypotheses or implementations. It argues that current benchmarks only assess end‑to‑end success and lack metrics for this decision‑making quality, proposing methods to measure and improve agent taste.

Read original