What does an AI visibility benchmark measure?
It measures whether web-grounded AI answers mention a target brand, how prominently it appears, how its answer-level mentions compare with selected competitors, and whether answers cite the submitted domain. 100 Questions presents those signals with their denominators and source evidence.
Which AI providers are included?
A benchmark sends the same frozen 25-question set to OpenAI, Anthropic, Google, and xAI through Vercel AI Gateway. The exact model IDs are configurable and are frozen with each run so the result records what was tested.
Why is it called 100 Questions if each run uses 25 questions?
Each run uses 25 unique questions and asks all 25 across four providers. That creates 100 planned provider answers: 25 × 4. The shared set makes the provider comparison more consistent than asking a different set to each model.
What is the difference between discovery and diagnostic questions?
Twenty discovery questions ask neutral category or use-case questions without naming the target, its aliases, or its domain. Five diagnostic questions name the target to examine trust, comparisons, pricing, support, implementation, and factual knowledge.
What makes an answer eligible for scoring?
The provider call must succeed and return valid HTTP or HTTPS web sources. Missing-source, unsupported-search, and failed results do not enter eligible-score denominators. They remain visible in coverage so a missing answer cannot silently improve or weaken a metric.
Are these results identical to ChatGPT, Claude, Gemini, or Grok apps?
No. This is an API-grounded benchmark. Consumer chat products can use different prompts, personalization, routing, models, and search behavior. A run should be treated as a time-stamped directional comparison, not a prediction of every consumer session.
Is the visibility score a statistically representative market ranking?
No. The generated question set is not claimed to be a random or representative sample. With only 25 questions, a simple worst-case interval is roughly plus or minus 20 percentage points. The benchmark is designed to expose directional evidence and gaps, not manufacture false precision.
Are benchmark runs private, and how long are answers retained?
Runs are private to the authenticated owner. The default answer-retention window is 30 days. The product retains normalized evidence and the versions needed to explain results, rather than complete raw provider payloads.
Is 100 Questions a subscription?
No. The first benchmark is $9, three benchmarks are $39, and ten are $99. After the introductory purchase, a single benchmark is $15. Every credit buys the same complete benchmark, remains valid for 12 months, and is purchased through Stripe-hosted Checkout. Taxes may apply.
What information is needed to run a benchmark?
You provide a subject name, canonical domain, category and use-case description, market, and locale. You can also provide aliases and selected competitors. Those inputs are frozen with the run and used to construct and analyze the shared question set.
Does this replace SEO analytics or prove why a model answered a certain way?
No. It complements search, content, and brand research by showing answer evidence at one point in time. Citations show sources returned with an answer, but they do not prove a model's internal reasoning or establish that one page caused a mention.
What is generative engine optimization (GEO)?
Generative engine optimization is the practice of making a brand and its expertise easier for AI answer systems to understand, retrieve, and cite. It complements technical SEO and useful content with clear entity information, consistent category language, source-worthy pages, and measurement across multiple AI providers.
How can a brand improve its visibility in AI-generated answers?
Start with clear product and category language on crawlable pages, publish original evidence that directly answers buyer questions, keep company facts consistent, earn relevant third-party references, and fix technical crawl barriers. Then rerun the same benchmark after meaningful changes. No tactic guarantees a mention, so improvements should be evaluated as directional evidence over time.