01 · Question construction
One frozen test, two question cohorts
A run starts from the subject name, canonical domain, aliases, category description, market, locale, and optional competitors. The system first creates a neutral category brief, then generates the two cohorts separately.
20 discovery questions
Neutral category and use-case questions. A question is rejected or regenerated if it contains the subject, an alias, or the canonical domain.
5 diagnostic questions
Target-named questions that probe trust, comparisons, pricing, support, implementation, and factual knowledge for manual review.
The exact 25-question set, locale, prompt versions, model IDs, run timestamp, cohort mix, grounding mode, and scoring version are frozen. Every provider receives the same question text and neutral system instruction, without hidden target context.
02 · Grounding and eligibility
Sources are a scoring requirement
OpenAI, Anthropic, Google, and xAI requests use the same bounded web search through Vercel AI Gateway. There is no plain-text fallback. A provider result enters score-eligible denominators only when the call succeeds and returns valid web sources.
Eligible
A successful provider answer with normalized, valid HTTP or HTTPS source URLs.
Excluded, but visible
Unsupported search, missing sources, and failed calls stay visible in coverage instead of silently entering a score.
03 · Metric definitions
Every number keeps its denominator
- Discovery visibility
- Eligible discovery answers that mention the target divided by all eligible discovery answers.
- Conservative visibility floor
- Discovery answers that mention the target divided by all planned discovery answers, including failures and ungrounded results in the denominator.
- Coverage
- Score-eligible answers divided by planned answers. Results are marked provisional when coverage is below 90%.
- Prominence
- The mean position of the target in eligible discovery answers: lead = 1, shortlist = 0.67, incidental = 0.33, and absent = 0.
- Share of voice
- Target mention events divided by target and selected-competitor mention events, with each entity counted at most once in an answer.
- Claimed-domain citation rate
- Eligible answers that cite the submitted canonical domain or one of its subdomains divided by all eligible answers.
- Sentiment
- Positive, neutral, and negative labels among answers that mention the target. It is not treated as a factual-accuracy score.
04 · Interpretation
Treat the result as a directional snapshot
- Search results and model responses vary, so a run measures one provider and web snapshot rather than a permanent rank.
- The API-grounded benchmark does not claim parity with consumer chat products, where routing, personalization, prompts, and search can differ.
- With 25 questions, a simple worst-case sampling interval is roughly ±20 percentage points. That is a scale cue, not a claim that the generated set is a random or representative statistical sample.
- Coverage should be read beside visibility. Below 90% coverage, the product labels metrics provisional.
05 · Evidence and retention
Inspect the answers behind the metrics
Runs are private to their authenticated owner. The product stores the frozen questions, normalized answers and sources, model and prompt versions, usage, metric labels, and exclusion reasons needed to explain a result. Complete raw provider payloads are not retained. The default answer-retention window is 30 days.
Run the same test across four providers
See the questions, sources, coverage, and calculations for your brand.
Start a benchmarkStill evaluating? Read common questions.