An analysis of 237 built-in LLM evaluation metrics across five open-source frameworks reveals that 126 scorers require a model to evaluate output, leaving only 94 capable of scoring without model dependency. The study quantifies the ongoing debate about AI evaluation reliability by categorizing metrics based on their operational requirements.

Read original