Evals.Health

What is the hardest healthcare AI benchmark?

updated August 17, 2026

There is no single hardest healthcare AI benchmark, because difficulty is defined differently on each axis: conversational difficulty, expert-selected professional tasks, agentic consistency, and safety. HealthBench Hard is the clearest difficulty-tail construction, built from the 1,000 conversations where frontier models failed most in May 2025. Agentic benchmarks produce the lowest absolute numbers once repeated success is required.

HealthBench Hard is the only tracked benchmark built explicitly as a difficulty tail. Five frontier models scored every conversation in the 5,000-conversation parent set; the roughly 1.5 percent where no model scored positive were removed, and the 1,000 lowest-average conversations became Hard. At its May 2025 release, the top score was o3's 0.320 against 0.60 for the same model on the full set, while GPT-3.5 Turbo, GPT-4o from August 2024, and Llama 4 Maverick scored 0.00. The dedicated board is healthbenchhard.ai.

HealthBench Professional is harder in a different sense: the tasks were chosen by physicians rather than by model failure. Its 525 tasks were selected from 15,079 real clinician conversations, with difficulty enriched about 3.5 times and roughly one third of tasks adversarial. Rubrics came from 190 physicians across 50 countries and 26 specialties, through three-stage adjudication with at least three physicians per criterion. Physician-written responses score 0.437 on those same rubrics, which sets a low human reference point on a scale of 0 to 1. The dedicated board is healthbenchprofessional.com.

On agentic benchmarks the binding constraint is consistency rather than single-attempt capability. EHR-Complex reports consistency below 50 percent at Pass^4 for nearly every model evaluated. CHI-Bench's launch report led with reliability rather than capability: no agent stayed above 20 percent across three identical runs. PhysicianBench reports Pass^3 alongside pass@1, and the lower half of its table falls steeply, with several agents near 1 percent on Pass^3.

Safety benchmarks measure a different kind of hardness. First, Do NOHARM v2 covers 1,100 consultation cases across 10 specialties with 12,747 expert annotations on 4,249 management options. Its v1 study found potential for severe harm in up to 24.6 percent of directly applied recommendations, and errors of omission accounted for more than 80 percent of the severe cases. A high score elsewhere does not carry over to this axis, since omission is invisible to benchmarks that score what a response contains.

Hardness is not fixed. MedQA and exam-style medical QA saturated above 95 percent by 2025 and were retired by their trackers, HealthBench Consensus stopped being reported separately, and OpenAI has said the parent HealthBench is approaching a noise ceiling for frontier models. Current scores for every benchmark named here, including which sets still separate models, are published on clinicalbenchmarks.ai.