Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints

Article automatically generated from technical news.

Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints Current agent benchmarks that rely on LLM judges are systematically unreliable—even on deterministic, replayed runs. This undermines almost every published leaderboard comparison. LLM Judge Endpoints Fail Basic Reliability: Rerunning Byte-Identical Inputs Flips Results LLM judging frameworks depend on the assumptio

Fonte originale