Evals.Health

Why do benchmark scores differ between sources?

updated August 17, 2026

A benchmark score is a property of a run, not of a model. Two sources publishing the same benchmark name can differ by several points because they used different graders, configurations, length adjustments, or model cohorts, and because vendor-run numbers use the vendor's own settings.

Published results fall into four bases. Vendor-reported numbers come from the model's own publisher under its own configuration. Independent-run numbers come from a third party that executes every model itself. Official-leaderboard numbers come from the benchmark's maintainers on a stated refresh cadence. Mixed boards combine rows of more than one kind. Vendor numbers often differ by points from independent runs of the same benchmark, so the basis label is part of the score.

HealthBench is the clearest worked example. Anthropic reports a raw score under its own protocol. OpenAI leads with length-adjusted numbers at maximum reasoning effort and reports production Instant settings separately. Baichuan ran its competitors itself. All three are labeled HealthBench, and none of the rows are directly comparable. The length adjustment alone removes roughly 2.99 points per 500 characters beyond 2,000 on HealthBench and 1.47 on HealthBench Professional, which is enough to reorder a table of verbose and terse models.

The grader is a second source of divergence. HealthBench and HealthBench Hard use GPT-4.1 by default; HealthBench Professional uses GPT-5.4 at low reasoning effort. GPT-4.1 was selected by meta-evaluation against physician grades, where it reached macro-F1 0.709 against 0.692 for o4-mini, 0.681 for o3, 0.661 for GPT-4.1-mini, and 0.580 for GPT-4.1-nano. Those spreads show that swapping the grader shifts scores independently of the model being graded.

Some metrics are relative by construction, and some boards score something other than a bare model. MedHELM ranks by mean win rate against the evaluated cohort, so every score moves when the model set changes. HealthAgentBench evaluates agent harnesses end to end rather than models. CHI-Bench accepts community submissions, so its rows mix author-run and submitted results across 45 harness configurations. The current frontier table for MedXpertQA (MM) is vendor-reported from Meta's Muse Spark launch evaluation and mirrored display-only elsewhere.

The practical rule is to compare rows only within one source under one configuration, and to treat differences across sources as unmeasured. This index records the basis, grader, and last update for every benchmark it tracks; current standings are on clinicalbenchmarks.ai.