ClawProBench introduces a trace-aware benchmark for evaluating AI agents within stateful runtimes, addressing the limitation of benchmarks that only assess final answers. It evaluates model-plus-runtime configurations across evidence acquisition, runtime routing, safety boundaries, and repeated execution. The benchmark is instantiated on Open