shareof.ai
Sign inStart free
Safety benchmark

HELM Safety leaderboard

A standardized collection of safety evaluations spanning violence, fraud, discrimination, harassment, sexual content and deception.

Version 1.8.0Mean scoreHigher is betterChecked August 13, 2026
Published results

Model rankings

These scores are reproduced from the linked primary source. They have not been re-evaluated or adjusted by Shareof.ai.

What this benchmark measures

Stanford CRFM combines five safety benchmarks across six risk categories and publishes the evaluated prompts, responses and scores.

Reported metric: Mean score. A higher score represents stronger measured performance within this benchmark.

How to interpret the ranking

A high aggregate score does not establish that a model is safe for every deployment or risk category.

Use the leaderboard to compare configurations evaluated under this exact methodology. Do not compare these values directly with a similarly named metric from another benchmark.

From model performance to brand performance

Does AI recommend your company?

Run a live visibility check and compare your brand with the companies appearing in the same AI answers.

Benchmark my business