The Healthcare AI Evaluation Reference
20 terms · 10 questions answered · updated August 17, 2026
Evaluation reports assume vocabulary most readers were never given. This site is the reference desk: a glossary of the terms healthcare AI evaluation runs on, and direct answers to the questions clinicians, researchers, and buyers bring to the results. Current standings live at clinicalbenchmarks.ai; the mechanics behind them are taught at clinicalevals.ai.
Questions, answered
- A vendor-reported benchmark result is a number the model's own vendor produced under a configuration it chose. Such a number is evidence about that system under stated conditions, and it is not comparable to a number another party produced on the same benchmark. A vendor claim becomes checkable only when the run's operator, grader, settings, and date are disclosed.
- Healthcare AI rankings change on each board's own refresh cycle, which ranges from roughly quarterly to never. Movement comes mostly from new models entering a cohort and from boards being re-run, not from measurement noise: HealthBench's repeat testing put run variability at a standard deviation of about 0.002 over 16 repeats. Any cited ranking should carry the date of the run that produced it.
- There is no single hardest healthcare AI benchmark, because difficulty is defined differently on each axis: conversational difficulty, expert-selected professional tasks, agentic consistency, and safety. HealthBench Hard is the clearest difficulty-tail construction, built from the 1,000 conversations where frontier models failed most in May 2025. Agentic benchmarks produce the lowest absolute numbers once repeated success is required.
- No tracked healthcare AI benchmark establishes that its scores predict clinical outcomes. Each measures performance on a defined task set, under a stated grading protocol, usually against physician-written rubrics rather than against patient results. Scores are best read as evidence about a specific construct, and they should be matched to the workflow a buyer intends to deploy.
- Five tracked benchmarks score agents rather than single responses: HealthAgentBench, CHI-Bench, PhysicianBench, EHR-Complex, and HealthAdminBench. They differ in environment, from terminal-based clinical artifacts and MIMIC-IV databases to real EHR APIs and computer-use administrative workflows. Most report a consistency metric alongside single-attempt success, and consistency is where scores fall furthest.
- On some rubric benchmarks, model responses score above the physician baselines published alongside those benchmarks. Those baselines measure a single written response graded against a rubric, not clinical care, and on safety and agentic benchmarks the pattern reverses.
- There is no single passing threshold. A HealthBench score is interpretable only against the variant it was run on and the published baselines for that variant: 0.13 to 0.48 for physicians on the parent set depending on what references they were given, and 0.437 for physician-written responses on HealthBench Professional.
- A benchmark score is a property of a run, not of a model. Two sources publishing the same benchmark name can differ by several points because they used different graders, configurations, length adjustments, or model cohorts, and because vendor-run numbers use the vendor's own settings.
- A benchmark is saturated when scores cluster near the top of its scale and stop separating systems. Exam-style medical question sets such as MedQA and MultiMedQA passed 95 percent by 2025 and were retired by most trackers, HealthBench Consensus is a near-saturated physician-consensus baseline, and OpenAI has said the parent HealthBench is approaching a noise ceiling for frontier models.
- Medical AI benchmarks use four grading methods: rubric scoring by an LLM grader against physician-written criteria, execution-grounded or deterministic verification, expert annotation of harm, and answer-matching accuracy. The method sets the limits on what a score can support.
Glossary
- An adversarial example is a benchmark item written to trigger a specific failure mode rather than to sample typical usage, such as a prompt carrying a false premise or a request that invites an unsafe omission.
- An agentic benchmark scores completion of multi-step tasks in an environment with tools and persistent state, rather than the quality of a single response. Success is usually verified by inspecting the resulting state rather than by grading prose.
- Benchmark saturation is the point at which top scores cluster so close to the ceiling that differences between leading models fall within run-to-run noise, and the ranking stops carrying information.
- A bootstrap confidence interval expresses how much a benchmark score depends on which tasks happened to be in the set, estimated by resampling the scored tasks with replacement and recomputing the score many times. It determines whether two models' scores are distinguishable at all.
- A composite index is a single score assembled from several underlying benchmarks under a fixed weighting scheme. The weights, rather than the components, determine what the index actually measures.
- Consensus criteria are rubric items retained only where multiple independent experts agree that the item is correct and required, producing a subset that measures uncontested behavior.
- An eval harness is the code and configuration that turns a benchmark dataset into a score: prompt construction, sampling settings, tool access, grader choice, and aggregation. Two harnesses running the same dataset can return different numbers for the same model.
- Grader bias is systematic error in the scoring step of an evaluation, where scores track a property of the response other than the quality being measured. Common sources are response length and the grader's resemblance to the model under test.
- A holdout set is the portion of a benchmark kept unpublished so that scores reflect general capability rather than exposure to the specific items being graded.
- An independent run is a benchmark result produced by a party that did not build the model under test, with one configuration applied to every model in the table.
- A length adjustment is a scoring correction that deducts points as a response grows past a set character count, so a model cannot raise its rubric score simply by writing more.
- Mean win rate scores a model by how often it beats the other models in the evaluated cohort, averaged across tasks. It is a relative ranking statistic, so a model's score changes when the cohort changes even if its underlying performance does not.
- Meta-evaluation measures the evaluator rather than the model: it scores how closely an automated grader reproduces expert judgments on the same responses. A benchmark that publishes no meta-evaluation offers no evidence that its grader is trustworthy.
- A model grader is a language model that scores another model's response against a rubric or reference answer, standing in for human graders so a benchmark can be run at scale and rerun on demand.
- A multi-turn benchmark scores a model's reply inside a conversation that already contains several exchanges, rather than a reply to an isolated prompt. Context handling and consistency across turns become part of what is scored.
- pass@k reports whether at least one of k attempts at a task succeeded; Pass^k reports whether all k attempts succeeded. The first describes capability, the second describes reliability.
- A physician baseline is the score that clinician-written responses receive when graded by the same rubric and the same grader as the model responses. It anchors an otherwise unitless rubric scale to human performance under identical scoring conditions.
- Red-teaming is the practice of constructing evaluation inputs designed to make a system fail, rather than sampling inputs that represent typical use. In healthcare evals it supplies the adversarial share of a task set and the stress cases that representative sampling rarely produces.
- A rubric criterion is a single scored statement in a grading rubric that specifies one thing a response should or should not do, judged as met or not met and carrying a weight that sets its contribution to the score.
- A vendor-reported score is a benchmark result published by the organization that makes the model, produced under that organization's own configuration and not independently reproduced.
Writing
Auditing Vendor Benchmark Claims A nine-point checklist health systems can use to audit an AI vendor's medical benchmark claim: benchmark version, who ran it, which grader, field size and date, consistency, and what the score cannot say clinically.