Evals.Health

Consensus criteria

updated August 17, 2026

Consensus criteria are rubric items retained only where multiple independent experts agree that the item is correct and required, producing a subset that measures uncontested behavior.

Physician-written rubrics contain disagreement. Adjudication is how a benchmark separates items that reflect one author's practice from items most reviewers would endorse. HealthBench Professional built adjudication into authoring: a three-stage process with at least three physicians reviewing each criterion, across 525 tasks selected from 15,079 real clinician conversations.

HealthBench shipped a consensus subset alongside the parent set in May 2025. That subset has since been retired from tracking, because it sat near saturation as a physician-consensus baseline and frontier runs stopped reporting it separately. The pattern generalizes. Items every reviewer agrees on tend to be items every competent model handles, so a consensus subset loses the ability to distinguish models before the parent set does.

Consensus subsets remain useful as floors rather than as rankings. A model that misses uncontested criteria has a defect that no aggregate score should be permitted to average away. Read alongside benchmark saturation and physician baseline, a consensus subset answers a narrower question than a hard subset does: not how far a model can go, but whether it clears what experts do not argue about.

See also: rubric criterion, benchmark saturation, physician baseline, meta-evaluation, holdout set.