Evals.Health

Length adjustment

updated August 17, 2026

A length adjustment is a scoring correction that deducts points as a response grows past a set character count, so a model cannot raise its rubric score simply by writing more.

Rubric grading rewards coverage. A longer response can satisfy more positive criteria without being more useful to a clinician, and in a clinical workflow it can be worse, because the reader has to find the answer inside it. The adjustment removes that incentive by subtracting a fixed amount for each unit of excess length.

The published parameters are specific. On HealthBench, the adjustment deducts about 2.99 points per 500 characters beyond 2,000, and about 1.47 points on HealthBench Professional. HealthBench Professional applies the adjustment by default with its GPT-5.4 grader at low reasoning effort. HealthBench itself is reported in both length-adjusted and unadjusted variants, and later OpenAI system cards lead with the adjusted figure at maximum reasoning effort while reporting production settings separately.

The variant is part of the number. An adjusted score and an unadjusted score from the same benchmark are not interchangeable, and a table that mixes them, or that sets an adjusted vendor figure next to an unadjusted independent one, is not measuring one thing. Any comparison across sources has to state which variant produced each row.

The adjustment approximates a real constraint, since response length has a cost in clinician time, but it is a blunt correction: it counts characters and does not evaluate what the extra characters contain. Adjusted and unadjusted rows are distinguished at clinicalbenchmarks.ai and on the dedicated board at healthbenchprofessional.com.

See also: rubric criterion, model grader, vendor-reported score, independent run, eval harness.