MedHELM evaluates model performance across a taxonomy of real-world healthcare tasks using public, gated and private clinical datasets.
Reported metric: Mean win rate. A higher score represents stronger measured performance within this benchmark.
How to interpret the ranking
Benchmark performance is not medical approval and should not be used as a substitute for clinical validation.
Use the leaderboard to compare configurations evaluated under this exact methodology. Do not compare these values directly with a similarly named metric from another benchmark.