Benchmark Radar is a living database and search engine that aggregates AI benchmarks and evaluation resources for large language models, agents, coding, reasoning, safety, and domain‑specific tasks. It enables researchers to discover benchmark datasets, code, and evaluation settings through daily updates and a searchable interface.

Read original