shareof.ai
Sign inStart free
Open-ended prompts benchmark

Arena-Hard v2.0 Preview leaderboard

A difficult set of open-ended real-world prompts evaluated with automated judges and style controls.

Version 2.0 PreviewWin rateHigher is betterChecked August 13, 2026
Published results

Model rankings

These scores are reproduced from the linked primary source. They have not been re-evaluated or adjusted by Shareof.ai.

RankModelRelative scoreWin rate

What this benchmark measures

Arena-Hard v2.0 uses 500 hard prompts and 250 creative-writing prompts, with GPT-4.1 and Gemini 2.5 used as automated judges in the published configurations.

Reported metric: Win rate. A higher score represents stronger measured performance within this benchmark.

How to interpret the ranking

These are automated-judge results, not the human-vote LMArena leaderboard. Judge choice and style controls materially affect scores.

Use the leaderboard to compare configurations evaluated under this exact methodology. Do not compare these values directly with a similarly named metric from another benchmark.

From model performance to brand performance

Does AI recommend your company?

Run a live visibility check and compare your brand with the companies appearing in the same AI answers.

Benchmark my business