Consensus criteria
updated August 17, 2026
Consensus criteria are rubric items retained only where multiple independent experts agree that the item is correct and required, producing a subset that measures uncontested behavior.
Physician-written rubrics contain disagreement. Adjudication is how a benchmark separates items that reflect one author's practice from items most reviewers would endorse. HealthBench Professional built adjudication into authoring: a three-stage process with at least three physicians reviewing each criterion, across 525 tasks selected from 15,079 real clinician conversations.
HealthBench shipped a consensus subset alongside the parent set in May 2025. That subset has since been retired from tracking, because it sat near saturation as a physician-consensus baseline and frontier runs stopped reporting it separately. The pattern generalizes. Items every reviewer agrees on tend to be items every competent model handles, so a consensus subset loses the ability to distinguish models before the parent set does.
Consensus subsets remain useful as floors rather than as rankings. A model that misses uncontested criteria has a defect that no aggregate score should be permitted to average away. Read alongside benchmark saturation and physician baseline, a consensus subset answers a narrower question than a hard subset does: not how far a model can go, but whether it clears what experts do not argue about.
See also: rubric criterion, benchmark saturation, physician baseline, meta-evaluation, holdout set.