Arena-Hard v2.0 uses 500 hard prompts and 250 creative-writing prompts, with GPT-4.1 and Gemini 2.5 used as automated judges in the published configurations.
Reported metric: Win rate. A higher score represents stronger measured performance within this benchmark.
How to interpret the ranking
These are automated-judge results, not the human-vote LMArena leaderboard. Judge choice and style controls materially affect scores.
Use the leaderboard to compare configurations evaluated under this exact methodology. Do not compare these values directly with a similarly named metric from another benchmark.