Evals.Health

Benchmark saturation

updated August 17, 2026

Benchmark saturation is the point at which top scores cluster so close to the ceiling that differences between leading models fall within run-to-run noise, and the ranking stops carrying information.

Saturation has a measurable threshold rather than a rhetorical one. Repeating HealthBench 16 times produced a standard deviation of about 0.002 on the overall score. When the spread across the top rows approaches that figure, the ordering is an artifact of sampling. OpenAI has said the parent HealthBench is approaching a noise ceiling for frontier models and points to HealthBench Professional for continued measurement.

The record is consistent. Exam-style medical question sets saturated above 95 percent by 2025 and were archived by their trackers. HealthBench Consensus was retired as a near-saturated physician-consensus baseline once frontier runs stopped reporting it separately. On MedScribe, top scores cluster near 90 and exact decimals below first place are not always displayed, which is what a board looks like as it approaches its ceiling.

Publishers respond by building harder sets from the same material. HealthBench Hard is the bottom fifth of HealthBench: 1,000 conversations selected because five frontier models scored them lowest in May 2025, after the roughly 1.5 percent on which no model scored positive were removed. At release, the top score on Hard was o3's 0.320, against 0.60 for the same model on the full set. HealthBench Professional took the other route, enriching difficulty about 3.5 times by selecting 525 tasks from 15,079 real clinician conversations; physician-written responses score 0.437 on its rubrics.

A saturated benchmark retains one use: as a floor. A model that fails a saturated set has a defect worth investigating. It is the ranking that expires, not the failure signal. Current standings, and the benchmarks retired for saturation, are listed at clinicalbenchmarks.ai, with the harder successor sets at healthbenchhard.ai and healthbenchprofessional.com.

See also: holdout set, consensus criteria, physician baseline, adversarial example, composite index.