The benchmark combines five long-context tasks including RULER, InfiniteBench and multi-round co-reference resolution.
Reported metric: Mean score. A higher score represents stronger measured performance within this benchmark.
How to interpret the ranking
Advertised context-window size is not the same as measured long-context performance.
Use the leaderboard to compare configurations evaluated under this exact methodology. Do not compare these values directly with a similarly named metric from another benchmark.