What is a good HealthBench score?
updated August 17, 2026
There is no single passing threshold. A HealthBench score is interpretable only against the variant it was run on and the published baselines for that variant: 0.13 to 0.48 for physicians on the parent set depending on what references they were given, and 0.437 for physician-written responses on HealthBench Professional.
HealthBench scores a model response against physician-written rubric criteria for that specific conversation. The example score is earned points divided by the maximum available positive points, and the overall score is the mean across examples, clipped to a 0 to 1 range. Some sites display the same figure on a 0 to 100 scale. Run-to-run noise is small at the benchmark level: OpenAI reported a standard deviation near 0.002 across 16 repeats of the full set, so differences of a few thousandths are not meaningful, while differences of several points are.
The parent set carries the widest set of anchors. Physicians writing unaided scored 0.13. With September 2024 model references they scored 0.31, and with April 2025 references 0.48, against 0.49 for those references alone. OpenAI has since said the parent HealthBench is approaching a noise ceiling for frontier models and points to HealthBench Professional for continued measurement, which means a high parent-set score now separates systems less than it did at release.
The harder variants have lower and more informative anchors. On HealthBench Hard, the 1,000 conversations where frontier models failed most, the best score at the May 2025 release was o3 at 0.320, against 0.60 for the same model on the full set. On HealthBench Professional, 525 tasks selected by physicians from 15,079 real clinician conversations, physician-written responses score 0.437 on the same rubrics. A score is best read as a distance from the baseline published for that variant rather than as a percentage of correctness.
Configuration changes the number more than most readers expect. The parent set and Hard use GPT-4.1 as the default grader; Professional uses GPT-5.4 at low reasoning effort with a length adjustment. Later OpenAI system cards report length-adjusted variants, where the adjustment removes about 2.99 points per 500 characters beyond 2,000 on HealthBench and 1.47 on Professional. Reasoning effort settings differ between reported rows as well. A HealthBench number without its variant, grader, and adjustment attached cannot be compared with another one.
Current per-model scores are maintained on clinicalbenchmarks.ai and on the dedicated boards for HealthBench Professional and HealthBench Hard. For procurement use, a benchmark score sets an upper bound on expected quality under study conditions and says nothing about performance on a specific deployment's case mix.