The paper introduces CausalDS, a benchmark designed to evaluate causal reasoning capabilities of data‑science agents that combine large language models with tool use. It addresses the gap between purely symbolic causal benchmarks and realistic data‑analysis tasks by providing datasets with grounded causal data‑generating processes. The benchmark aims to offer diverse, non‑templated evaluation scenarios for assessing integrated reasoning and analysis performance.

Read original