Holdout set
updated August 17, 2026
A holdout set is the portion of a benchmark kept unpublished so that scores reflect general capability rather than exposure to the specific items being graded.
Public benchmark items enter training data. Once they do, a score measures memorization as much as reasoning, and the benchmark stops tracking what it was built to track. A holdout set is the standard defense: the items are scored but never released, while a published portion documents what the tasks look like.
Health benchmarks differ in how much they hold back. HealthBench published its 5,000 conversations and 48,562 criteria in full in May 2025, which makes it reproducible and also exposes it. Health Optimization Bench describes its first version as 89 released tasks, wording that separates the published portion from the suite behind it, and its items are freshness-dependent, written against primary sources in preventive and optimization medicine, so they age as the evidence base moves.
Exposure is one of several forces that retires a benchmark. Exam-style medical question sets passed 95 percent by 2025 and were archived by most trackers, an outcome that capability gains and contamination both contributed to in proportions no one can recover after the fact. Periodic refresh is the other defense, since a benchmark that publishes new items on a cadence keeps a moving target even when past items are public.
For a buyer, the operative question is not whether a public benchmark holds items back but whether its score is being used as evidence about unseen work. Evaluation on a private set of the organization's own cases is the version of a holdout set that no vendor can have trained on, and it is the only one whose contents the buyer controls.
See also: benchmark saturation, adversarial example, independent run, eval harness, consensus criteria.