ResearchAudio

Benchmark audit / reproducibility checklist

A benchmark score is not a protocol.

Put any model, agent, or coding benchmark claim through twelve checks. Expose the hidden versions, scaffolds, budgets, graders, task choices, and uncertainty before the number drives a decision.

Audit a benchmark claim →

Inspect one published result

Can somebody reproduce the number?

Count only details that are public, specific, and usable. This is a triage worksheet—not a certification, leaderboard, or substitute for rerunning the evaluation.

Protocol bands

0–3: headline only.
A number without enough protocol to investigate.

4–7: partial protocol.
Some controls are visible; comparison remains fragile.

8–10: decision-useful.
The result can inform a bounded decision with stated uncertainty.

Claim → protocol → decision

The number is the last line of the receipt.

Claim“Model A scores 92.4.”

A useful lead, not yet an operational conclusion.

ProtocolWhich 92.4?

Dataset, snapshot, scaffold, budget, grader, trials, and environment determine what the number means.

DecisionWill it survive your workload?

Use task distribution, failures, uncertainty, cost, and latency to bound the choice you can defend.

Primary-source guide

Small protocol changes can move the result.

OpenAI recommends task-specific evals, logging, representative datasets, continuous evaluation, and calibration of automated scoring with human judgment. Anthropic similarly emphasizes measurable criteria, real-world task distributions, edge cases, and reliable graders.

METR's Time Horizon 1.1 update is a concrete warning: changes to tasks, human-time estimates, scoring, and agent scaffolding changed estimates and left wide confidence intervals. A benchmark claim should expose these variables before it becomes a purchase or deployment decision.

OpenAI evaluation best practices → Anthropic: define success and build evaluations → METR Time Horizon 1.1 update → Terminal-Bench task and harness repository →

What makes an AI benchmark result reproducible?

Identify the task release, exact model snapshot, harness and prompts, resource limits, environment, grader, trial method, uncertainty, and artifacts. A third party should be able to reconstruct the run without guessing.

Why can two AI benchmark scores be incomparable?

Changing the scaffold, tools, budget, retry policy, grader, task mix, environment, or model snapshot changes the protocol. Compare scores only when those variables match or their effects are isolated.

How many trials should an AI benchmark use?

Enough to estimate variability for the decision at hand. Publish trial count, sampling method, spread, and uncertainty; a single average can hide unstable behavior.

Is this AI benchmark checklist a certification?

No. It is a fast evidence audit. Even a 12/12 receipt needs independent reproduction and review for the intended domain, risk, and workload.

Read the protocol behind the score

Get the next benchmark claim audited.

ResearchAudio turns primary sources, evaluation choices, failure cases, costs, and uncertainty into practical decisions for engineers and builders.