Evals.Health

Which healthcare AI benchmarks are saturated?

updated August 17, 2026

A benchmark is saturated when scores cluster near the top of its scale and stop separating systems. Exam-style medical question sets such as MedQA and MultiMedQA passed 95 percent by 2025 and were retired by most trackers, HealthBench Consensus is a near-saturated physician-consensus baseline, and OpenAI has said the parent HealthBench is approaching a noise ceiling for frontier models.

Saturation is a property of the task set, not of the models. Once the highest scores sit close to the maximum, the remaining gap is dominated by grader noise and rubric edge cases, and the ranking stops carrying information. Exam-style medical question answering reached that point first: MedQA and MultiMedQA saturated above 95 percent by 2025, and the trackers that hosted them archived the results. The Hugging Face Open Medical-LLM Leaderboard, which was built on those exam sets, has had no frontier submissions in 2026.

Rubric benchmarks are following the same path more slowly. HealthBench Consensus, the physician-consensus baseline released with HealthBench, is near-saturated, and frontier runs stopped reporting it separately. OpenAI has said the parent HealthBench itself is approaching a noise ceiling for frontier models and points to HealthBench Professional for continued measurement. Documentation benchmarks show partial ceilings rather than full saturation: on MedScribe, top scores cluster near 90, and the exact decimals below first place are not always displayed.

Several benchmarks that no longer appear in current comparisons are inactive rather than saturated, and the distinction matters when reading an old citation. AgentClinic and CRAFT-MD have published no frontier-model results since 2025. Microsoft's SDBench and MAI-DxO sequential-diagnosis study was not re-run on current models. MedAgentBench survives only as a component of the ARISE MAST composite, with no current frontier rows on its standalone board. MedArena's clinician preference pool is too thin on current frontier models to quote.

The benchmarks that still separate systems are the ones built on selected-difficulty or execution-grounded tasks. HealthBench Hard took the 1,000 conversations where frontier models failed most, and its best score at the May 2025 release was 0.320. HealthBench Professional enriched difficulty by roughly 3.5 times, made about one third of its tasks adversarial, and scores physician-written responses at 0.437 on the same rubrics. First, Do NOHARM measures harm frequency and severity instead of rubric credit, its first version finding potential for severe harm in up to 24.6 percent of directly applied recommendations. Agentic boards remain far from any ceiling: EHR-Complex reports consistency below 50 percent at Pass^4 for nearly every model, and CHI-Bench reported at launch that no agent stayed above 20 percent across three identical runs.

Saturation is a reason to change benchmark, not a reason to conclude a capability is solved: an exam set above 95 percent says nothing about consultation safety or multi-step workflow reliability. Which boards are currently live, and how recently each was refreshed, is tracked on clinicalbenchmarks.ai.