Decision relevance
Would the answer influence discovery, a shortlist, risk assessment, or purchase?
Build a controlled question set for ChatGPT, Claude, Gemini, Grok, or another answer engine. Use the selection framework, then draw from 100 prompts covering discovery, fit, comparison, proof, and brand accuracy.
Question selection framework
A useful question represents a real decision, exposes a meaningful competitive set, and can be judged consistently. Start with customer language from sales calls, support logs, site search, paid-search terms, interviews, and conventional keyword research.
Would the answer influence discovery, a shortlist, risk assessment, or purchase?
Does the discovery prompt avoid naming or praising the target brand?
Does it add a new audience, use case, constraint, comparison, or proof requirement?
Can reviewers label mentions, prominence, competitors, citations, and accuracy consistently?
Use the bank, not the whole bank
Replace the bracketed fields with language your buyers actually use. Remove questions that do not affect discovery or a buying decision. Then freeze the final wording, market, locale, providers, and collection date.
Do not mix brand-named prompts into a discovery score. If the question already supplies the brand, the answer cannot prove that an AI system would have discovered it independently.
Remove near-duplicates that ask the same decision in slightly different words. Balance broad category questions with high-value use cases, constraints, comparisons, and proof checks so one prompt family cannot dominate the score.
Recommended starter benchmark
The downloadable CSV marks a balanced 25-question starter set. It is a practical scale for a directional baseline, not a statistically representative sample of everything buyers ask.
The complete library
Tests whether the brand appears before the buyer names a vendor.
Tests whether answers connect the brand to the situations it is built for.
Tests competitive framing, substitution, and shortlist position.
Tests whether answers surface credible evidence instead of unsupported claims.
Tests accuracy and positioning after the brand is explicitly named. Keep these separate from discovery visibility.
Put the library to work
Run 25 frozen buyer questions across OpenAI, Claude, Gemini, and Grok.
Turn observed answer counts into a transparent, component-level score.
Run a focused 10-prompt manual test and preserve the evidence.
Track provider conditions, answers, citations, competitors, and reruns.
Map cited pages to claims, brand effects, gaps, and actions.
Summarize scope, results, sources, limitations, and priorities.
Translate the benchmark into a concise client decision narrative.
Questions about the questions
No. Treat the list as a research bank. Select a smaller set that represents real buyer decisions, freeze it, and run the same questions across providers and reruns. The 25-question starter set uses 20 neutral discovery questions and five brand-named diagnostics.
A brand-named question tests recognition and factual accuracy after the user supplies the brand. A neutral question tests discovery: whether the brand appears before it is named. Combining them inflates visibility and hides the more important discovery gap.
No. They are structured test questions, not a claim about prompt frequency or representative sampling. Validate them with customer interviews, sales calls, site search, support logs, paid-search data, and conventional keyword research.
Keep mention rate, prominence, selected-competitor share of voice, citations, sentiment, and provider coverage separate. Preserve the answer and source evidence behind every metric, and report failed or unsupported runs as coverage rather than silent misses.
From worksheet to evidence
Review the stored answers, citations, competitor mentions, coverage, limitations, and prioritized actions before running your own benchmark.