Healthcare task coverage analysis

Read the denominator behind the rank.

Independent analysis of MedHELM’s 37-benchmark journal snapshot, clinical task coverage, private data, LLM-jury scoring and published model comparisons.

Independent analysis by Arcophos · updated

37 benchmarks across clinical work

MedHELM ↗

The final journal-paper suite, organized into five categories. Each square represents one benchmark. The taxonomy also names 121 tasks; those are not 121 independently measured datasets. [2]

Clinical decision support

12

Patient communication

8

Clinical note generation

6

Medical research assistance

6

Administration and workflow

5

Counts sum to 37. The paper’s open/closed-ended prose split is internally inconsistent with this total, so it is not charted. [2]

What an independent team can access

16 public7 gated14 private

Access categories in the final-paper inventory. A public-only rerun changes the suite denominator. [2]

Published results

Source record ↗

Selected paper-reported measurements. Dates, model configurations and scoring conditions belong to each panel.

Paper-reported results / selected rows

Final-paper pairwise win rate

Nature Medicine 2026 final paper, Table 1; 37 benchmarks and nine historical models.

Pairwise win rate · 0–1
00.51
Reported
DeepSeek R1Win SD 0.11; macro-average 0.75.
0.66
o3-mini (2025-01-31)Win SD 0.15; macro-average 0.77.
0.66
Claude 3.5 Sonnet (20241022)Win SD 0.15; macro-average 0.74.
0.63
Claude 3.7 Sonnet (20250219)Win SD 0.15; macro-average 0.73.
0.63
GPT-4o (2024-05-13)Win SD 0.17; macro-average 0.74.
0.58

SD is variation across the paper’s comparisons, not a confidence interval. Ties count as wins under the stated method. Table 1 takes precedence over a swapped Claude macro-average sentence.

Source: Table 1 [2]

An original analytical tool

Map MedHELM to the work being evaluated

Evidence explorer

Explore concrete benchmarks inside the final 37-evaluation paper. Inputs and endpoints matter more than a broad category name.

9 of 9 evidence entries shown

Clinical decision support

Medical calculation

Read dossier ↗
Input
Patient note and requested calculation
Output
Numeric or categorical answer
Measure
MedCalc accuracy
Interpretation boundary

Exact or tolerance-based scoring varies by calculation.

[2]
Clinical decision support

Clinical error detection

Read dossier ↗
Input
Clinical narrative
Output
Error detection or correction
Measure
Task-specific error metric
Interpretation boundary

Correctness of detection differs from usefulness of a revised note.

[2]
Clinical note generation

Visit documentation

Read dossier ↗
Input
Patient–doctor conversation
Output
Structured clinical note
Measure
LLM-jury assessment
Interpretation boundary

A jury assesses outputs; actual user correction is outside this endpoint.

[2]
Clinical note generation

Radiology summarization

Read dossier ↗
Input
Findings text
Output
Impression summary
Measure
LLM-jury assessment
Interpretation boundary

Text summarization does not establish image interpretation.

[2]
Patient communication

Medication questions

Read dossier ↗
Input
Consumer medication query
Output
Long-form answer
Measure
LLM-jury assessment
Interpretation boundary

Reference and judge setup shape answer ratings.

[2]
Patient communication

Personalized instructions

Read dossier ↗
Input
Procedure and case context
Output
Patient instructions
Measure
LLM-jury assessment
Interpretation boundary

Patient comprehension is not directly observed.

[2]
Medical research assistance

Clinical research SQL

Read dossier ↗
Input
Natural-language data request
Output
SQL query
Measure
Execution accuracy
Interpretation boundary

Correctness is bound to the supplied schema and target.

[2]
Administration and workflow

Billing code assignment

Read dossier ↗
Input
Medical note
Output
ICD-10 code set
Measure
F1 score
Interpretation boundary

F1 is distinct from raw exact-match correctness.

[2]
Administration and workflow

Referral routing

Read dossier ↗
Input
Referral information
Output
Task-defined referral decision
Measure
Task-specific accuracy
Interpretation boundary

Private operational data limits independent reruns.

[2]

This is an independently annotated coverage map, not a local validation plan or a reproduction of private datasets. [2]

The benchmark in detail

All dossiers →
Dossier01

Final journal paper: 37-benchmark snapshot

MedHELM ↗

Coverage, access and scoring all shape a hospital benchmark.

UnitA benchmark-specific evaluation instance; 37 benchmark scores feed the aggregate.MeasurePairwise win rate and macro-average

What we examine

MedHELM supplies a structured view of clinical language-model tasks. This independent publication examines the final journal paper’s 37 evaluations, the evidence available for each task family and the aggregation behind its model rankings. Our analysis and coverage explorer make the unit of measurement visible. Published results belong to the cited study; we have not rerun its models or accessed its private datasets. The current community project is linked separately from this fixed paper snapshot.

Inspect the task
Trace the actual input, output and evaluation setting for each named benchmark.
Read the evidence
Explore cited cohort counts and selected historical results with their measurement boundaries.
Make the inference explicit
Separate the authors’ observations from our analytical interpretation and proposed evaluation questions.

Analysis & interpretation

All analyses →

Questions, answered

How many benchmarks are in this MedHELM analysis?

We analyze the final January 2026 journal paper’s 37-benchmark inventory. Its earlier preprint used 35; the evolving community project is linked separately.

Are the 121 tasks 121 evaluated datasets?

No. The 121 tasks belong to a clinical taxonomy. The final paper’s measured suite contains 37 benchmarks mapped across its categories and subcategories.

What does a MedHELM win rate mean?

It is a comparison against the paper’s other models across the fixed benchmark suite. The stated method counts a score at least as high as a rival’s as a win, including ties.

Can the whole suite be reproduced from public data?

The final paper includes public, gated and private components. A public-only run is a subset and should be labeled with its actual coverage.

Prepare a comparison worksheet

Working tool / saved on this device

Prepare a reproducible benchmark reading

Interactive worksheet

Use this secondary worksheet to record the version, task and evidence boundary of the benchmark you are considering. Checked items indicate documented decisions, not measured performance.

Identify the experiment

Every measurement has a source. Dossiers preserve benchmark versions, scoring conditions and access notes.

Download the evidence ↗