shareof.ai
Sign inStart free
Long context benchmark

HELM Long Context leaderboard

A reproducible evaluation of how reliably models reason over long documents, conversations and multi-hop evidence.

Version latest published editionMean scoreHigher is betterChecked August 13, 2026
Published results

Model rankings

These scores are reproduced from the linked primary source. They have not been re-evaluated or adjusted by Shareof.ai.

What this benchmark measures

The benchmark combines five long-context tasks including RULER, InfiniteBench and multi-round co-reference resolution.

Reported metric: Mean score. A higher score represents stronger measured performance within this benchmark.

How to interpret the ranking

Advertised context-window size is not the same as measured long-context performance.

Use the leaderboard to compare configurations evaluated under this exact methodology. Do not compare these values directly with a similarly named metric from another benchmark.

From model performance to brand performance

Does AI recommend your company?

Run a live visibility check and compare your brand with the companies appearing in the same AI answers.

Benchmark my business