Auditing Vendor Benchmark Claims
August 17, 2026
A vendor benchmark claim is an assertion that a named AI product reached a specific score on a named medical benchmark, and it is verifiable only when the vendor also states the benchmark version, who ran it, which grader and configuration were used, and when the run happened.
Most claims that reach health system procurement omit at least one of those. Reporting conventions differ by publisher, and a number lifted from a system card loses its qualifiers on the way into a slide. The checklist below is meant to be worked through on a single vendor call.
Nine questions to ask before accepting a score
1. Which benchmark, and which version. HealthBench, HealthBench Hard, and HealthBench Professional are three different sets. HealthBench holds 5,000 conversations graded against 48,562 physician-written criteria and was released in May 2025. HealthBench Hard is the bottom fifth of that set, the 1,000 conversations on which frontier models failed most at release. HealthBench Professional, released April 2026, is 525 tasks physicians selected out of 15,079 real clinician conversations. Leaderboards are versioned too: MedHELM stood at 5.0.0 as of May 14, 2026.
2. Which scale, and which variant. Some sites report 0 to 1, others 0 to 100. HealthBench results exist in length-adjusted and unadjusted forms; the adjustment removes roughly 2.99 points per 500 characters beyond 2,000 on HealthBench and 1.47 on HealthBench Professional. A vendor quoting one variant against a competitor's other variant is comparing nothing.
3. Who ran it. Establish whether the number is vendor-reported, produced by an independent runner, or taken from a leaderboard the benchmark's publisher maintains. On the parent HealthBench, Anthropic reports a raw score under its own protocol while OpenAI leads with length-adjusted numbers at maximum reasoning effort; those rows sit under one benchmark name and do not compare.
4. Which grader, at what setting. Rubric benchmarks are scored by a model. HealthBench's default grader is GPT-4.1, selected by meta-evaluation against physician grades, where it reached macro-F1 0.709 ahead of o4-mini at 0.692 and o3 at 0.681. HealthBench Professional's default grader is GPT-5.4 at low reasoning effort. A vendor that swapped in another grader has produced a number that does not belong on the published board.
5. What was tested, the model or the product. HealthAgentBench evaluates agent harnesses end to end across 54 tasks in 7 environments rather than bare models. CHI-Bench's August 12, 2026 update covered 45 harness configurations, and its authors found that harness choice matters as much as model choice. When the vendor sells a product built on a model, a bare-model score describes something the buyer is not purchasing.
6. Field size and run date. MedCode listed 85 models as of August 15, 2026; position in a field that size moves easily. MedHELM ranks by mean win rate, a metric defined against the evaluated cohort, so it shifts whenever the model set changes. WHBench froze its model set in March 2026, so its table describes that cohort only.
7. Repeats and variance. HealthBench's overall score is stable, with run-to-run standard deviation near 0.002 over 16 repeats, but stability on one benchmark implies nothing about another. Health Optimization Bench publishes 95 percent bootstrap confidence intervals beside every score. Ask for the interval or the repeat count.
8. The same-rubric anchor. A score carries little meaning without a reference point measured the same way. Physician-written responses score 0.437 on the HealthBench Professional rubrics. On the parent HealthBench, physicians writing unaided scored 0.13, and physicians working from April 2025 model references scored 0.48. At HealthBench Hard's May 2025 release, o3 led at 0.32, and several older models scored 0.00. Ask which anchor the vendor's number should be read against.
9. Whether the benchmark still separates systems. Saturated benchmarks produce high scores that carry no information. Exam-style medical QA including MedQA passed 95 percent by 2025 and was retired by most trackers. OpenAI has said the parent HealthBench is approaching a noise ceiling for frontier models and points to HealthBench Professional for continued measurement. A near-ceiling result on a near-ceiling benchmark is not evidence of an advantage.
Consistency is a separate question from capability
Capability benchmarks report whether a system can succeed once. A health system depends on whether it succeeds every time. Agentic benchmarks separate the two by reporting Pass^k, the rate of success across k identical consecutive attempts. EHR-Complex reports consistency below 50 percent at Pass^4 for nearly every model it evaluated. CHI-Bench's launch report led with reliability rather than capability: no agent stayed above 20 percent across three identical runs. When a vendor quotes pass@1 for an agentic product, the consistency figure is the one to request.
What the score cannot say about clinical use
A benchmark score measures response quality against a written rubric on a fixed task set. It does not measure patient outcome. It says nothing about the buyer's own population or workflow.
Rubric coverage is not a safety case. First, Do NOHARM v1 found potential for severe harm in up to 24.6 percent of directly applied consultation recommendations, and errors of omission accounted for more than 80 percent of the severe cases. Omission is the failure mode that a rubric built around expected content catches least reliably, and that a scripted demonstration does not catch at all.
Grader agreement is average agreement. GPT-4.1 exceeds the average physician's agreement with physician grades in 5 of HealthBench's 7 themes. That is a statement about aggregate behavior across tens of thousands of criteria, not about the single case a safety committee will want to examine.
Composite indexes carry weights the buyer did not choose. The Artificial Analysis Healthcare and Medical Index weights medical and health knowledge at 35 percent and draws on general-purpose benchmarks rather than purpose-built clinical tasks. MAST composites six component benchmarks, is marked preview, and publishes component-level detail for only one of them. A composite cannot be decomposed back into the capability a specific deployment needs. Related: can a benchmark score predict clinical safety.
Recording the answers
Each answer belongs in the procurement record beside the claim it qualifies, in the vendor's own words. A claim that survives the checklist can be re-checked when the vendor or the benchmark publisher issues an update.
Current standings sit outside this checklist by design. Scores move on publication cadences the buyer does not control, and a number transcribed into a contract exhibit is stale on arrival. Current per-model results across tracked benchmarks are published at clinicalbenchmarks.ai, with dedicated boards at healthbenchprofessional.com, healthbenchhard.ai, and healthoptimizationbench.com. The checklist establishes what a claim means; the boards supply what the number currently is.
Terms used here are defined in the glossary, and the questions carry the short answers.