Evals.Health

Vendor-reported score

updated August 17, 2026

A vendor-reported score is a benchmark result published by the organization that makes the model, produced under that organization's own configuration and not independently reproduced.

Vendor numbers use the vendor's harness, grader, reasoning-effort setting, and scoring variant. Those choices move results by points on the same benchmark, so a vendor figure and an independent figure for one model on one benchmark are two measurements rather than one number reported twice. Mixing graders or configurations breaks comparability even when the benchmark name is identical.

HealthBench shows the problem directly. Anthropic reports a raw score under its own protocol. OpenAI leads with length-adjusted numbers at maximum reasoning effort and reports production Instant settings separately. Baichuan ran its competitors itself. Those rows share a benchmark name and little else, which is why they should not be stacked into a single ranking.

Whole tables can be vendor-reported. The current frontier table for MedXpertQA (MM), the 2,000-question multimodal subset, comes from Meta's Muse Spark launch evaluation and is mirrored display-only elsewhere. OpenAI's dynamic mental health evaluations are vendor-run by construction: they cover OpenAI models only, they are not independently runnable, and OpenAI states that the error rates are not representative of average production traffic.

A vendor-reported score is not therefore wrong. It is unverified, and it is usually produced under settings chosen to show the model at its best, which differs from the settings a buyer will deploy. The practical test is whether the configuration is disclosed in enough detail to be rerun by someone else. Basis labels for every tracked benchmark are published with the current standings at clinicalbenchmarks.ai.

See also: independent run, eval harness, length adjustment, grader bias, composite index.