shareof.ai
Sign inStart free
OpenAI model

o4-mini (2025-04-16) benchmarks

Independent scores collected from compatible, versioned benchmark leaderboards. Each result links back to the evaluation and methodology that produced it.

4 benchmarksPrimary sources linkedChecked August 13, 2026
Measured results

Where this model was evaluated

Scores from different benchmarks measure different capabilities and are intentionally not averaged into one universal rating.

#2 of 11

HELM Capabilities

0.812

A transparent, reproducible evaluation of general language-model capabilities across a curated collection of difficult scenarios.

View full leaderboard →
#6 of 10

HELM Safety

0.973

A standardized collection of safety evaluations spanning violence, fraud, discrimination, harassment, sexual content and deception.

View full leaderboard →
#2 of 10

MedHELM

0.697

A healthcare-focused evaluation framework grounded in real clinical tasks, including reasoning and patient communication.

View full leaderboard →
#4 of 4

Arena-Hard v2.0 Preview

74.6%

A difficult set of open-ended real-world prompts evaluated with automated judges and style controls.

View full leaderboard →

What the scores mean

A benchmark score describes performance under one specific evaluation setup. It does not guarantee the same ordering for a different task, prompt, model version or deployment configuration.

Source policy

Shareof.ai preserves the published model name, benchmark version and source link. Provider-reported claims are not presented as independent results.

Read our methodology →
Models are measured. Brands should be too.

See if AI recommends your company.

Run a live sample and inspect the competitors and sources appearing in the answers.

Check my AI visibility →