Evals.Health

Can I trust vendor benchmark claims?

updated August 17, 2026

A vendor-reported benchmark result is a number the model's own vendor produced under a configuration it chose. Such a number is evidence about that system under stated conditions, and it is not comparable to a number another party produced on the same benchmark. A vendor claim becomes checkable only when the run's operator, grader, settings, and date are disclosed.

Vendor-reported results are benchmark numbers produced by the party that sells the model. This index labels every tracked result by basis: vendor-reported, independent-run, official-leaderboard, or mixed. The label determines what the number can be compared against. A vendor number supports a claim about that system under conditions the vendor selected. It does not by itself support a claim that the system beats a competitor whose number came from a different run.

HealthBench shows how quickly cross-source rows stop being comparable. Anthropic reports a raw score under its own protocol, OpenAI leads with length-adjusted numbers at maximum reasoning effort and reports production Instant settings separately, and Baichuan ran its competitors itself. All three describe the same 5,000-conversation benchmark. Because no two of those numbers were produced under the same protocol, sorting them into one table ranks configurations rather than models.

Two configuration choices move scores on rubric benchmarks: the grader model and the length adjustment. HealthBench's official grader is GPT-4.1, chosen by meta-evaluation against physician grades, where it reached macro-F1 0.709 against 0.692 for o4-mini, 0.681 for o3, 0.661 for GPT-4.1-mini, and 0.580 for GPT-4.1-nano; it exceeded the average physician's agreement in 5 of 7 themes. HealthBench Professional uses GPT-5.4 at low reasoning effort. Length adjustment subtracts about 2.99 points per 500 characters beyond 2,000 on HealthBench and about 1.47 on Professional, so an adjusted and an unadjusted score for the same response differ by construction.

Some evaluations are vendor-only by design. OpenAI's Dynamic Mental Health Evaluations cover OpenAI models, are not independently runnable, and carry OpenAI's own note that the reported error rates are not representative of average production traffic. Vendor-reported tables also appear for benchmarks the vendor did not publish: the current frontier table for MedXpertQA (MM) comes from Meta's Muse Spark launch evaluation and is mirrored display-only. Neither case is disqualifying. Both change what the number supports.

A procurement check on any vendor claim asks who executed the run, which grader and reasoning settings were used, whether the comparison rows were executed by the same party under the same settings, and when the run took place. Where those facts are absent, the claim stands only for the vendor's own system. Current standings for the benchmarks tracked here are published on clinicalbenchmarks.ai, with dedicated boards at healthbenchprofessional.com, healthbenchhard.ai, and healthoptimizationbench.com.