Benchmark v2

How the 100 Questions AI visibility benchmark works

Each run is a time-stamped, directional comparison. It freezes one question set, asks it across four AI providers with web search, and keeps coverage and metric denominators visible beside the results.

Benchmark at a glance

25

questions frozen per run

20 + 5

discovery and diagnostic questions

4

AI providers tested with the same set

100

planned provider answers per run

01 · Question construction

One frozen test, two question cohorts

A run starts from the subject name, canonical domain, aliases, category description, market, locale, and optional competitors. The system first creates a neutral category brief, then generates the two cohorts separately.

20 discovery questions

Neutral category and use-case questions. A question is rejected or regenerated if it contains the subject, an alias, or the canonical domain.

5 diagnostic questions

Target-named questions that probe trust, comparisons, pricing, support, implementation, and factual knowledge for manual review.

The exact 25-question set, locale, prompt versions, model IDs, run timestamp, cohort mix, grounding mode, and scoring version are frozen. Every provider receives the same question text and neutral system instruction, without hidden target context.

02 · Grounding and eligibility

Sources are a scoring requirement

OpenAI, Anthropic, Google, and xAI requests use the same bounded web search through Vercel AI Gateway. There is no plain-text fallback. A provider result enters score-eligible denominators only when the call succeeds and returns valid web sources.

Eligible

A successful provider answer with normalized, valid HTTP or HTTPS source URLs.

Excluded, but visible

Unsupported search, missing sources, and failed calls stay visible in coverage instead of silently entering a score.

03 · Metric definitions

Every number keeps its denominator

Discovery visibility
Eligible discovery answers that mention the target divided by all eligible discovery answers.
Conservative visibility floor
Discovery answers that mention the target divided by all planned discovery answers, including failures and ungrounded results in the denominator.
Coverage
Score-eligible answers divided by planned answers. Results are marked provisional when coverage is below 90%.
Prominence
The mean position of the target in eligible discovery answers: lead = 1, shortlist = 0.67, incidental = 0.33, and absent = 0.
Share of voice
Target mention events divided by target and selected-competitor mention events, with each entity counted at most once in an answer.
Claimed-domain citation rate
Eligible answers that cite the submitted canonical domain or one of its subdomains divided by all eligible answers.
Sentiment
Positive, neutral, and negative labels among answers that mention the target. It is not treated as a factual-accuracy score.

04 · Interpretation

Treat the result as a directional snapshot

  • Search results and model responses vary, so a run measures one provider and web snapshot rather than a permanent rank.
  • The API-grounded benchmark does not claim parity with consumer chat products, where routing, personalization, prompts, and search can differ.
  • With 25 questions, a simple worst-case sampling interval is roughly ±20 percentage points. That is a scale cue, not a claim that the generated set is a random or representative statistical sample.
  • Coverage should be read beside visibility. Below 90% coverage, the product labels metrics provisional.

05 · Evidence and retention

Inspect the answers behind the metrics

Runs are private to their authenticated owner. The product stores the frozen questions, normalized answers and sources, model and prompt versions, usage, metric labels, and exclusion reasons needed to explain a result. Complete raw provider payloads are not retained. The default answer-retention window is 30 days.

Run the same test across four providers

See the questions, sources, coverage, and calculations for your brand.

Start a benchmark

Still evaluating? Read common questions.