How to benchmark AI citations vs competitors?
Run a fixed set of 20–30 buyer questions across ChatGPT, Perplexity, and Gemini. Record every brand named to calculate your mention rate and share of voice. This comparison highlights exactly where competitors win citations and where your visibility gaps exist.
Why benchmark AI citations?
AI models are non-deterministic, so repeating tests monthly is essential and answers vary between runs due to temperature settings and platform updates.
Mention Rate Formula: (Times your brand is named / Total prompts) × 100. Citation Rate Formula: (Times your brand is cited as a source / Total prompts) × 100. Share of Voice Formula: (Times your brand is named / Total brand mentions across all competitors) × 100.
Hypothetical Example: If you run 20 prompts and your brand is named 10 times, your mention rate is 50%. If the engines named 50 brands total, your share of voice is (10 / 50) × 100 = 20%. This arithmetic is for illustration; actual data requires logged benchmarking.
How many prompts do I need for a benchmark?
Everything depends on a good, fixed prompt set. Write 20 to 30 prompts that reflect how buyers really ask about your category, spanning comparison, fit, and recommendation questions. Freeze the list so every run is comparable. This range is a practical framework for starting, not a statistically validated universal threshold. Larger sample sizes may be necessary to detect smaller shifts in visibility.
This is the same foundation used for measuring AI share of voice. The difference here is emphasis: you are watching competitors as closely as yourself, so your prompts should include the queries where rivals are most likely to appear.
How should I calculate AI citation share of voice?
Benchmark against the brands you actually compete with, not aspirational giants. Include three to six real rivals plus any surprise names that keep appearing in answers. To calculate, divide the number of times your brand is named by the total brand mentions across all competitors, then multiply by 100. This metric clarifies your relative visibility.
- Direct competitors you lose deals to
- Category leaders buyers compare you against
- Unexpected brands the engines name in your space
That last group is often the most useful. If an engine keeps citing a brand you dismissed, it is seeing something you should understand.
What are the limitations of AI benchmarking?
Generative AI is non-deterministic; the same prompt can yield different outputs on different days. Geography, personalization history, and model updates also affect visibility. These observations track relative visibility, not statistically significant change.
| Prompt | Platform | Date | Geography | Brand Mention | Citation URL | Type |
|---|---|---|---|---|---|---|
| Example query | ChatGPT | 2026-08-21 | US | Brand A | example.com | Comparison |
How to maintain consistency
Run prompts in logged-out or fresh-session windows to limit personalization. Use a fixed test date and geography. Re-run the full set on a monthly cadence to track trends, as generated answers vary.
Do this on a regular schedule. A single run is a snapshot; a repeated one shows whether the gap between you and each rival is widening or closing.
Turn the gap into action
Once you have the table, read it prompt by prompt. Where a competitor is named and you are not, ask why. Usually it comes down to a page they have and you lack, proof they publish that you do not, or corroboration they have earned.
This is where the benchmark connects to strategy. Studying why your competitor gets cited and you don't on your weakest prompts tells you exactly what to build next.
When two brands look identical
Sometimes you and a rival cover the same ground and one still gets cited more. In those close calls, small edges decide it: clearer structure, stronger proof, better third-party mentions. Understanding how AI picks between two similar brands helps you find the edge that moves the benchmark in your favor.
Because answers are non-deterministic, treat the benchmark as a trend, not a verdict. The goal is to become steadily more likely to be named than your rivals, run after run.
Record more than just who was named
| Prompt | Engine | Date | Geography | Brand Named | Rank/Framing | Cited Source |
|---|---|---|---|---|---|---|
| Example query | ChatGPT | 2026-08-21 | US | Brand A | Positive, first | Link X |
Capture these fields for each answer. This detail helps identify patterns in why competitors are chosen and how to adapt your content.
A benchmark gets far more useful when you capture a little extra detail on each answer, not just the yes-or-no of whether a brand appeared. The richer record turns a scoreboard into a diagnosis.
For every prompt and engine, note which brands were named, in what order, and how they were framed. Being listed first as the clear recommendation is different from being mentioned last as an also-ran. Record which sources the engine cited, since those pages are the ones doing the work, and note whether the mention was positive, neutral, or a caveat. Keep the raw answer text so you can revisit the wording later.
With that detail, patterns emerge quickly. You might find a rival is cited because one specific review platform keeps surfacing, or because they own a comparison page you lack. Those are actionable findings. A bare mention count tells you the score; the fuller record tells you why the score is what it is, which is what actually lets you close the gap.
What are the limits of this methodology?
This benchmark is an observed-visibility workflow, not a statistically validated standard. Treat the 20–30 prompt range as a practical starting framework. Generated answers vary between runs — OpenAI documents that AI models are non-deterministic — so a single snapshot is weaker than a repeated schedule. Personalization, logged-in history, and geography can change who gets named. A brand mention is not a citation; track them separately. Platform differences matter: the same prompt can surface different sources on ChatGPT, Perplexity, Claude, Gemini, and Copilot.
Q: When is a change in mention rate meaningful?
A: Only after you freeze the prompt set, control session and geography as best you can, and compare like-for-like runs over time. Small swings on one day are noise. Look for sustained gaps versus the same competitors on the same prompts.
What should a crawlable benchmark log include?
Use a simple table (spreadsheet or HTML) with one row per prompt × platform × run. Do not invent sample percentages — leave measured cells blank until you log real runs.
| Prompt | Platform | Run date | Geography | Brand mentioned? | Cited source URL | Query type |
|---|---|---|---|---|---|---|
| ChatGPT | recommendation | |||||
| Perplexity | comparison | |||||
| Claude | fit |
Questions people ask
Three to six is usually right. Include the rivals you actually lose deals to, the category leaders buyers compare you against, and any surprise brands the engines keep naming in your space. Too few misses the real picture; too many dilutes your focus. The unexpected names are often the most instructive, since the engine is seeing something you may have overlooked.
On a regular cadence, such as every two to four weeks, using the exact same prompt set. AI answers are non-deterministic, so a single run is only a snapshot. A repeated schedule reveals whether the gap between you and each competitor is widening or closing, which is the information that tells you if your work is paying off.
Work prompt by prompt. Wherever a competitor is named and you are not, identify what they have that you lack, usually a specific page, published proof, or third-party corroboration. Then build or earn that one thing and re-test. Closing gaps one at a time, starting with the prompts closest to a buying decision, moves the benchmark fastest.
Want this done for you?
Want to know if AI is recommending you or your competitor? Get an AI visibility audit from Acromatico.
Get a free AI Visibility Audit →