inspect_ai, the UK AI Safety Institute’s framework built with Meridian Labs, standardizes large‑language‑model evaluation by defining three primitives: a dataset of versioned samples (input with a target), a solver that generates answers (single model call, prompt chain, or tool‑using agent), and a scorer that converts the answer to a numeric metric (text match, custom function, or another model). A Task binds these components, and the surrounding infrastructure supplies 20+ provider adapters, sandboxed tool execution, parallel and resumable logging, and a web viewer for transcript inspection. The tutorial demonstrates three concrete examples. First, hello_eval.py creates four multiple‑choice questions, uses a multiple_choice solver, and a choice scorer, achieving 0.75 accuracy on a mock model (n=4). Second, a custom numeric_partial scorer yields a mean of 0.50 ± 0.29 (n=4) by rewarding exact answers and partial token matches. Third, agent_eval.py implements a calculator tool invoked by a ReAct agent, producing a deterministic score of 1.0 (n=1) with full tool‑call logs. All runs require only Python 3.12, pip install inspect‑ai, and no API keys; the mock provider guarantees reproducibility. The takeaway is that inspect_ai supplies scriptable, auditable evaluation pipelines that replace vague “vibe‑checks” with measurable, reproducible metrics.

Read original