shareof.ai
Sign inStart free
General intelligence benchmark

HELM Capabilities leaderboard

A transparent, reproducible evaluation of general language-model capabilities across a curated collection of difficult scenarios.

Version 1.14.0Mean scoreHigher is betterChecked August 13, 2026
Published results

Model rankings

These scores are reproduced from the linked primary source. They have not been re-evaluated or adjusted by Shareof.ai.

What this benchmark measures

Stanford CRFM evaluates models on a versioned collection of capability scenarios and publishes prompts, responses, metrics and reproducible configurations.

Reported metric: Mean score. A higher score represents stronger measured performance within this benchmark.

How to interpret the ranking

Scores apply only to HELM Capabilities version 1.14.0 and should not be compared with scores from a different benchmark version.

Use the leaderboard to compare configurations evaluated under this exact methodology. Do not compare these values directly with a similarly named metric from another benchmark.

From model performance to brand performance

Does AI recommend your company?

Run a live visibility check and compare your brand with the companies appearing in the same AI answers.

Benchmark my business