HELM Capabilities
Stanford CRFM evaluates models on a versioned collection of capability scenarios and publishes prompts, responses, metrics and reproducible configurations.
Every published number should answer four questions: who measured it, what exactly was tested, which version was used, and when the result was checked.
Shareof.ai is a curator, not the original evaluator, unless a dataset is explicitly labelled as original Shareof.ai research.
Scores come from official research repositories, versioned leaderboards or the organization that conducted the evaluation.
Dates, reasoning modes, quantization and other configuration details remain attached to the published model label.
We do not rank results from incompatible versions, dataset splits, judges or evaluation setups in one table.
Automated judges, contamination risks, specialist scope and other material caveats appear alongside the leaderboard.
Each dataset keeps its source URL, version, metric, date checked and original organization.
Each source remains linked so readers can audit the original results.
Stanford CRFM evaluates models on a versioned collection of capability scenarios and publishes prompts, responses, metrics and reproducible configurations.
Stanford CRFM combines five safety benchmarks across six risk categories and publishes the evaluated prompts, responses and scores.
The benchmark combines five long-context tasks including RULER, InfiniteBench and multi-round co-reference resolution.
MedHELM evaluates model performance across a taxonomy of real-world healthcare tasks using public, gated and private clinical datasets.
Arena-Hard v2.0 uses 500 hard prompts and 250 creative-writing prompts, with GPT-4.1 and Gemini 2.5 used as automated judges in the published configurations.
Shareof.ai industry statistics are calculated from completed, measured scans. We aggregate results by category and never publish an individual prospect, domain, prompt or private report.
Public cohorts require at least 50 businesses and 200 sampled answers. Every page displays the number of companies and answers represented.
If an original source changes a result or methodology, we update the relevant page and its checked date. We do not silently combine the replacement result with an older benchmark version.
Corrections can be reported to hello@shareof.ai.
See the sampled answers, competitors and cited sources behind your company’s result.
Run a free visibility check →