How often do healthcare AI rankings change?
updated August 17, 2026
Healthcare AI rankings change on each board's own refresh cycle, which ranges from roughly quarterly to never. Movement comes mostly from new models entering a cohort and from boards being re-run, not from measurement noise: HealthBench's repeat testing put run variability at a standard deviation of about 0.002 over 16 repeats. Any cited ranking should carry the date of the run that produced it.
Three separate clocks govern how fast a ranking moves: the model release cycle, the board's refresh cadence, and the metric's dependence on the evaluated cohort. Run-to-run noise is the smallest of these on rubric benchmarks. HealthBench reports a run variability standard deviation of about 0.002 over 16 repeats, which is far below the gaps that separate models on that scale of 0 to 1.
Refresh cadence differs widely across the tracked set. MedHELM is run by its Stanford-led maintainers on a roughly quarterly cadence and sat at version 5.0.0 as of May 14, 2026, without rows for the Claude 5 family or GPT-5.6. Vals AI re-runs MedCode and MedScribe itself and last updated both on August 15, 2026, at 85 and 84 models. The ARISE MAST composite was last updated August 15, 2026 and is marked as a preview whose scores may move before full release. CHI-Bench was released May 20, 2026 and updated August 12, 2026 across 45 harness configurations. HealthAdminBench states no refresh cadence.
Some tracked results are static by nature. PhysicianBench and EHR-Complex report scores from their papers with no standalone public leaderboard, so their rows change only when a new paper appears. WHBench froze its model set in March 2026, before GPT-5.6 and the Claude 5 family shipped. Rankings drawn from these sources age against the model release cycle rather than being corrected by it.
Metric design also moves rankings without any model changing. MedHELM ranks by mean win rate, which is computed relative to the evaluated cohort, so every score shifts when the model set changes. Ceiling effects have a similar effect near the top: MedScribe's leading scores cluster near 90 and exact decimals below first place are not always displayed, which makes ordering at the top unstable. OpenAI has said the parent HealthBench is approaching a noise ceiling for frontier models and points to HealthBench Professional for continued measurement.
Boards also leave the tracked set entirely. MedQA and exam-style medical QA saturated above 95 percent by 2025 and were archived by most trackers. HealthBench Consensus stopped being reported separately, the standalone MedAgentBench board has no current frontier rows and survives inside the MAST composite, and MedArena's clinician preference pool is too thin on current frontier models to quote. Current standings, with the date of each run, are published per benchmark on clinicalbenchmarks.ai.