shareof.ai
Sign inStart free
Evidence policy

How we curate AI benchmarks

Every published number should answer four questions: who measured it, what exactly was tested, which version was used, and when the result was checked.

Curation principles

Accuracy before volume.

Shareof.ai is a curator, not the original evaluator, unless a dataset is explicitly labelled as original Shareof.ai research.

  1. 01

    Use primary sources

    Scores come from official research repositories, versioned leaderboards or the organization that conducted the evaluation.

  2. 02

    Preserve model configurations

    Dates, reasoning modes, quantization and other configuration details remain attached to the published model label.

  3. 03

    Separate benchmark versions

    We do not rank results from incompatible versions, dataset splits, judges or evaluation setups in one table.

  4. 04

    Disclose limitations

    Automated judges, contamination risks, specialist scope and other material caveats appear alongside the leaderboard.

  5. 05

    Retain provenance

    Each dataset keeps its source URL, version, metric, date checked and original organization.

Current source registry

Where the data comes from

Each source remains linked so readers can audit the original results.

General intelligence

HELM Capabilities

Stanford CRFM evaluates models on a versioned collection of capability scenarios and publishes prompts, responses, metrics and reproducible configurations.

Safety

HELM Safety

Stanford CRFM combines five safety benchmarks across six risk categories and publishes the evaluated prompts, responses and scores.

Long context

HELM Long Context

The benchmark combines five long-context tasks including RULER, InfiniteBench and multi-round co-reference resolution.

Healthcare

MedHELM

MedHELM evaluates model performance across a taxonomy of real-world healthcare tasks using public, gated and private clinical datasets.

Open-ended prompts

Arena-Hard v2.0 Preview

Arena-Hard v2.0 uses 500 hard prompts and 250 creative-writing prompts, with GPT-4.1 and Gemini 2.5 used as automated judges in the published configurations.

Original visibility benchmarks

Shareof.ai industry statistics are calculated from completed, measured scans. We aggregate results by category and never publish an individual prospect, domain, prompt or private report.

Public cohorts require at least 50 businesses and 200 sampled answers. Every page displays the number of companies and answers represented.

Corrections and updates

If an original source changes a result or methodology, we update the relevant page and its checked date. We do not silently combine the replacement result with an older benchmark version.

Corrections can be reported to hello@shareof.ai.

Measured, not guessed

Benchmark your own AI visibility.

See the sampled answers, competitors and cited sources behind your company’s result.

Run a free visibility check →