Evals.Health

Composite index

updated August 17, 2026

A composite index is a single score assembled from several underlying benchmarks under a fixed weighting scheme. The weights, rather than the components, determine what the index actually measures.

Building a composite requires two decisions that are usually made quietly: how each component is normalized onto a shared scale, and how much weight each one carries. Both encode a judgment about which capabilities matter for the intended use. A composite is legible only when its components and weights are published and its component scores can be inspected separately.

MAST (ARISE AI Research Network, August 2026) combines six curated clinical benchmarks: First Do NOHARM v2, SCT-Bench, MedAgentBench v2, PhysicianBench, ReXrank Mini, and CPC-Bench, across 11 models. The board is marked as a preview and was last updated August 15, 2026, and component-level breakdowns are published only for First, Do NOHARM v2. Where component scores are withheld, two models with the same index value may have opposite safety and agentic profiles, and the index cannot be audited against a buyer's own weighting.

The Artificial Analysis Healthcare & Medical Index publishes its weights: Medical & Health Knowledge 35 percent, Agentic Knowledge Work 25 percent, Non-Hallucination 15 percent, Reasoning 15 percent, and Agentic Customer Interaction 10 percent, drawn from AA-Omniscience, GDPval-AA v2, Humanity's Last Exam, and tau3-Banking, with 27 of 159 models scored. Its components are general-purpose benchmarks tilted toward health-relevant slices rather than purpose-built clinical tasks, which places it closer to a capability index than to a clinical one.

A composite is a shortlisting instrument. Acceptance criteria should be written against the component that matches the deployed workflow, since a weighted average can hide a failing component behind strong performance elsewhere. Per-benchmark and per-model pages are at clinicalbenchmarks.ai.

See also: mean win rate, benchmark saturation, vendor-reported score, independent run.