Evals.Health

Rubric criterion

updated August 17, 2026

A rubric criterion is a single scored statement in a grading rubric that specifies one thing a response should or should not do, judged as met or not met and carrying a weight that sets its contribution to the score.

Each criterion is scored independently of the others. HealthBench, released by OpenAI in May 2025, carries 48,562 criteria across 5,000 conversations, written by 262 physicians from 60 countries covering 26 specialties and 49 languages. The median conversation has 11 criteria, with a range of 2 to 48. Weights run from -10 to +10, and negative criteria subtract points when a response contains the harmful content they describe.

The arithmetic matters for reading a score. On HealthBench, an example score is the points a response earned divided by the maximum positive points available for that example, so a response that trips negative criteria can score below zero on a given example. The overall score is the mean across examples, clipped to the range 0 to 1. A headline number therefore compresses thousands of independent judgments, and two models at the same headline value can be failing on entirely different criteria.

Criterion quality depends on who wrote and reviewed it. HealthBench Professional, released in April 2026, used a three-stage rubric adjudication with at least three physicians per criterion, drawing on 190 physicians from 50 countries and 26 specialties. Physician-written responses score 0.437 against those same rubrics, which places the human reference well below the top of the scale. Health Optimization Bench takes a different route to the same problem: every task is written against a primary source and audited by model families that did not author it.

For procurement, the criterion is the unit that transfers to internal evaluation. A vendor claim about a rubric benchmark is checkable only if the rubric version and the grader configuration are stated. Current scores for rubric benchmarks are published at clinicalbenchmarks.ai, and the dedicated board for HealthBench Professional is healthbenchprofessional.com. The judgment itself is usually made by a model grader.

See also: model grader, consensus criteria, physician baseline, length adjustment, grader bias.