Do AI models beat doctors on health benchmarks?
updated August 17, 2026
On some rubric benchmarks, model responses score above the physician baselines published alongside those benchmarks. Those baselines measure a single written response graded against a rubric, not clinical care, and on safety and agentic benchmarks the pattern reverses.
The comparison is possible only where a benchmark publishes a physician baseline measured on the same rubrics used to grade models. Two do. On HealthBench (OpenAI, May 2025), physicians writing responses unaided scored 0.13. Physicians given September 2024 model references scored 0.31. Physicians given April 2025 model references scored 0.48, against 0.49 for those references on their own, meaning the physicians were no longer improving on what the model had already written. On HealthBench Professional (OpenAI, April 2026), physician-written responses score 0.437 on the rubrics that grade the model runs.
A baseline of that kind measures one written reply to a conversation transcript, produced under study conditions and scored by a model grader against physician-written criteria. It does not measure examination, physical findings, longitudinal follow-up, or responsibility for the outcome. HealthBench criteria carry weights from -10 to +10, and negative criteria subtract points for harmful content, so an individual example can score below zero. A score above a physician baseline means the response earned more rubric credit on that task set.
Benchmarks built to measure harm rather than rubric credit report the opposite direction. The first version of First, Do NOHARM found that direct application of LLM consultation recommendations carried potential for severe harm in up to 24.6 percent of cases, with errors of omission responsible for more than 80 percent of the severe errors. Agentic health benchmarks show the same fragility under repetition: EHR-Complex reports consistency below 50 percent at Pass^4 for nearly every model evaluated, and the CHI-Bench launch report stated that no agent stayed above 20 percent across three identical runs.
Difficulty selection also decides whether the comparison means anything. HealthBench Hard is the 1,000 conversations where frontier models failed most, and at its May 2025 release the best score was o3 at 0.320, against 0.60 for the same model on the full set; GPT-3.5 Turbo, GPT-4o from August 2024, and Llama 4 Maverick scored 0.00. HealthBench Professional enriched difficulty by roughly 3.5 times over the parent set and made about one third of its tasks adversarial. The same model family can sit above a physician baseline on one slice and near the floor on another.
Current standings change with every model release and are not reproduced here. Per-model results for each tracked benchmark are on clinicalbenchmarks.ai, with dedicated boards at healthbenchprofessional.com, healthbenchhard.ai, and healthoptimizationbench.com. A comparison against a physician baseline holds only when the run configuration matches the one the baseline was measured under, as described in why benchmark scores differ between sources.