- Input
- Patient note and requested calculation
- Output
- Numeric or categorical answer
- Measure
- MedCalc accuracy
Healthcare task coverage analysis
Read the denominator behind the rank.
Independent analysis of MedHELM’s 37-benchmark journal snapshot, clinical task coverage, private data, LLM-jury scoring and published model comparisons.
37 benchmarks across clinical work
MedHELM ↗The final journal-paper suite, organized into five categories. Each square represents one benchmark. The taxonomy also names 121 tasks; those are not 121 independently measured datasets. [2]
Clinical decision support
12Patient communication
8Clinical note generation
6Medical research assistance
6Administration and workflow
5Counts sum to 37. The paper’s open/closed-ended prose split is internally inconsistent with this total, so it is not charted. [2]
What an independent team can access
Access categories in the final-paper inventory. A public-only rerun changes the suite denominator. [2]
Published results
Source record ↗Selected paper-reported measurements. Dates, model configurations and scoring conditions belong to each panel.
Paper-reported results / selected rows
Final-paper pairwise win rate
Nature Medicine 2026 final paper, Table 1; 37 benchmarks and nine historical models.
SD is variation across the paper’s comparisons, not a confidence interval. Ties count as wins under the stated method. Table 1 takes precedence over a swapped Claude macro-average sentence.
Source: Table 1 [2]
An original analytical tool
Map MedHELM to the work being evaluated
Explore concrete benchmarks inside the final 37-evaluation paper. Inputs and endpoints matter more than a broad category name.
9 of 9 evidence entries shown
- Input
- Clinical narrative
- Output
- Error detection or correction
- Measure
- Task-specific error metric
- Input
- Patient–doctor conversation
- Output
- Structured clinical note
- Measure
- LLM-jury assessment
- Input
- Findings text
- Output
- Impression summary
- Measure
- LLM-jury assessment
- Input
- Consumer medication query
- Output
- Long-form answer
- Measure
- LLM-jury assessment
- Input
- Procedure and case context
- Output
- Patient instructions
- Measure
- LLM-jury assessment
- Input
- Natural-language data request
- Output
- SQL query
- Measure
- Execution accuracy
- Input
- Medical note
- Output
- ICD-10 code set
- Measure
- F1 score
- Input
- Referral information
- Output
- Task-defined referral decision
- Measure
- Task-specific accuracy
This is an independently annotated coverage map, not a local validation plan or a reproduction of private datasets. [2]
The benchmark in detail
All dossiers →Final journal paper: 37-benchmark snapshot
MedHELM ↗
Coverage, access and scoring all shape a hospital benchmark.
What we examine
MedHELM supplies a structured view of clinical language-model tasks. This independent publication examines the final journal paper’s 37 evaluations, the evidence available for each task family and the aggregation behind its model rankings. Our analysis and coverage explorer make the unit of measurement visible. Published results belong to the cited study; we have not rerun its models or accessed its private datasets. The current community project is linked separately from this fixed paper snapshot.
- Inspect the task
- Trace the actual input, output and evaluation setting for each named benchmark.
- Read the evidence
- Explore cited cohort counts and selected historical results with their measurement boundaries.
- Make the inference explicit
- Separate the authors’ observations from our analytical interpretation and proposed evaluation questions.
Analysis & interpretation
All analyses →What do MedHELM’s 37 benchmarks and 121 tasks represent?
Distinguish the clinical taxonomy from the measured benchmark inventory in MedHELM’s final journal publication.
How should you read MedHELM’s win rate and LLM jury?
Understand relative ranking, equal benchmark weighting and the three-model jury before interpreting a MedHELM headline.
Why do MedHELM versions and dataset access matter?
Compare the 2025 preprint, 2026 journal snapshot and evolving project without mixing their inventories or results.
Questions, answered
How many benchmarks are in this MedHELM analysis?
We analyze the final January 2026 journal paper’s 37-benchmark inventory. Its earlier preprint used 35; the evolving community project is linked separately.
Are the 121 tasks 121 evaluated datasets?
No. The 121 tasks belong to a clinical taxonomy. The final paper’s measured suite contains 37 benchmarks mapped across its categories and subcategories.
What does a MedHELM win rate mean?
It is a comparison against the paper’s other models across the fixed benchmark suite. The stated method counts a score at least as high as a rival’s as a win, including ties.
Can the whole suite be reproduced from public data?
The final paper includes public, gated and private components. A public-only run is a subset and should be labeled with its actual coverage.
Prepare a comparison worksheet
Working tool / saved on this device
Prepare a reproducible benchmark reading
Use this secondary worksheet to record the version, task and evidence boundary of the benchmark you are considering. Checked items indicate documented decisions, not measured performance.
Every measurement has a source. Dossiers preserve benchmark versions, scoring conditions and access notes.
Download the evidence ↗