shareof.ai
Sign inStart free
OpenAI model

o3 (2025-04-16) benchmarks

Independent scores collected from compatible, versioned benchmark leaderboards. Each result links back to the evaluation and methodology that produced it.

3 benchmarksPrimary sources linkedChecked August 13, 2026
Measured results

Where this model was evaluated

Scores from different benchmarks measure different capabilities and are intentionally not averaged into one universal rating.

#3 of 11

HELM Capabilities

0.811

A transparent, reproducible evaluation of general language-model capabilities across a curated collection of difficult scenarios.

View full leaderboard →
#1 of 10

HELM Safety

0.982

A standardized collection of safety evaluations spanning violence, fraud, discrimination, harassment, sexual content and deception.

View full leaderboard →
#1 of 4

Arena-Hard v2.0 Preview

85.9%

A difficult set of open-ended real-world prompts evaluated with automated judges and style controls.

View full leaderboard →

What the scores mean

A benchmark score describes performance under one specific evaluation setup. It does not guarantee the same ordering for a different task, prompt, model version or deployment configuration.

Source policy

Shareof.ai preserves the published model name, benchmark version and source link. Provider-reported claims are not presented as independent results.

Read our methodology →
Models are measured. Brands should be too.

See if AI recommends your company.

Run a live sample and inspect the competitors and sources appearing in the answers.

Check my AI visibility →