shareof.ai
Sign inStart free
Healthcare benchmark

MedHELM leaderboard

A healthcare-focused evaluation framework grounded in real clinical tasks, including reasoning and patient communication.

Version latest published editionMean win rateHigher is betterChecked August 13, 2026
Published results

Model rankings

These scores are reproduced from the linked primary source. They have not been re-evaluated or adjusted by Shareof.ai.

What this benchmark measures

MedHELM evaluates model performance across a taxonomy of real-world healthcare tasks using public, gated and private clinical datasets.

Reported metric: Mean win rate. A higher score represents stronger measured performance within this benchmark.

How to interpret the ranking

Benchmark performance is not medical approval and should not be used as a substitute for clinical validation.

Use the leaderboard to compare configurations evaluated under this exact methodology. Do not compare these values directly with a similarly named metric from another benchmark.

From model performance to brand performance

Does AI recommend your company?

Run a live visibility check and compare your brand with the companies appearing in the same AI answers.

Benchmark my business