shareof.ai
Sign inStart free
OpenAI model

GPT-4.1 (2025-04-14) benchmarks

Independent scores collected from compatible, versioned benchmark leaderboards. Each result links back to the evaluation and methodology that produced it.

2 benchmarksPrimary sources linkedChecked August 13, 2026
Measured results

Where this model was evaluated

Scores from different benchmarks measure different capabilities and are intentionally not averaged into one universal rating.

#10 of 10

HELM Safety

0.963

A standardized collection of safety evaluations spanning violence, fraud, discrimination, harassment, sexual content and deception.

View full leaderboard →
#1 of 10

HELM Long Context

0.588

A reproducible evaluation of how reliably models reason over long documents, conversations and multi-hop evidence.

View full leaderboard →

What the scores mean

A benchmark score describes performance under one specific evaluation setup. It does not guarantee the same ordering for a different task, prompt, model version or deployment configuration.

Source policy

Shareof.ai preserves the published model name, benchmark version and source link. Provider-reported claims are not presented as independent results.

Read our methodology →
Models are measured. Brands should be too.

See if AI recommends your company.

Run a live sample and inspect the competitors and sources appearing in the answers.

Check my AI visibility →