OmegaUse-OfficeVal is a new benchmark designed to evaluate Large Language Model (LLM) agents performing long-horizon office-suite tasks. The framework introduces task-level economic grounding to assess whether agents can complete complex workflows at a reasonable cost. The benchmark consists of 100 tasks derived from real-world practitioner requests.

Read original