Stanford CRFM evaluates models on a versioned collection of capability scenarios and publishes prompts, responses, metrics and reproducible configurations.
Reported metric: Mean score. A higher score represents stronger measured performance within this benchmark.
How to interpret the ranking
Scores apply only to HELM Capabilities version 1.14.0 and should not be compared with scores from a different benchmark version.
Use the leaderboard to compare configurations evaluated under this exact methodology. Do not compare these values directly with a similarly named metric from another benchmark.