Updated September 7, 2026
Short answer: To benchmark your brand's AI citations vs competitors, run the same 20 buyer queries across ChatGPT, Perplexity, Gemini, and Claude, then score every mention, position, and accuracy for you and your rivals. The outcome is a clear share-of-voice gap you can track monthly and close with focused content fixes.
A single AI answer is noise. A fixed query set, scored the same way every month, is the signal.
An AI citation is any time an assistant names your brand — with or without a link — inside an answer to a buyer question. A bare mention, a linked source, and a ranked recommendation are not the same strength, so log them separately. Your benchmark is the share of a fixed query set where you appear versus each rival, not a vanity count of one-off screenshots.
That distinction matters because AI Overviews and chat answers pull from a wider pool than classic top-10 blue links. Ahrefs' 2026 AI Overviews analysis found only about 38% of cited pages ranked in Google's top 10 — down from roughly 76% months earlier — so a brand can win citations without winning the SERP, and lose them while ranking well.
Start with 15 to 30 questions a real customer would actually type or ask out loud — not keyword-stuffed search terms. Mix intent types so the sample mirrors how people use AI engines:
Keep the list in a simple spreadsheet with one row per query. Reuse the same rows every cycle so month-to-month moves reflect real change, not a new sample.
Run every query across at least four engines: ChatGPT, Perplexity, Gemini, and Claude. Each pulls from different training and retrieval behavior, so a brand can show up consistently in one and vanish in another. Skipping any of the four leaves a real gap in the picture, not just a smaller sample.
Run each query at least twice per engine when you can. Answers are not fixed — the same question asked twice can return a different list, especially on engines that pull live web results. One pass shows what happened once; two or three passes start to show a pattern.
For every response, record three things for your brand and for each competitor that appears:
That turns a vague sense that "we don't show up much" into a structured dataset you can compare run to run.
AI share of voice is the percentage of sampled queries where a brand is mentioned, set against how often each competitor is mentioned on that same query set. If you appear in 6 of 20 queries and your top rival appears in 14, that gap is the clearest single number to track — more useful than any one answer, because answers vary run to run.
Break share of voice down by query type. A brand might dominate direct brand questions while staying invisible on comparison questions, which points to a specific gap — usually a missing or weak comparison page — rather than a general visibility problem.
You do not need a paid GEO dashboard to benchmark your brand's AI citations vs competitors. Sit down with a spreadsheet, 20 questions a real buyer in your city would ask out loud, and three to five rivals a customer would actually consider — not every name in the category. Run each question in a fresh ChatGPT, Perplexity, Gemini, and Claude chat. Mark mentioned or not, first-named or buried, and whether the facts are right. Do the same 20 next month. That monthly gap is the number. Software can wait; the first useful read is whether you show up at all against the shops next door. When you do want tools later, start with the GEO tools for small business comparison, then the company SEO retainers if you want the audit run for you.
Monthly is a solid default for most small businesses — frequent enough to catch a real trend, light enough to sustain without tooling. Categories with fast-moving rivals, or teams that just shipped a GEO fix and want an early read, can run a smaller core of 5 to 10 queries every two weeks.
Keep the query list mostly stable between benchmarks. Changing the questions every time makes it impossible to tell whether a shift reflects a real change or just a different sample.
Acromatico's company GEO and SEO retainers start with this citation benchmark, then close the gaps and re-measure so you can see the share-of-voice move.
See company SEO retainersLimit the list to three to five real alternatives a customer would actually consider — not every business in the category, and not aspirational names far outside your size or price point. Too many rivals dilute the comparison; a focused list makes the share-of-voice gap obvious.
Include at least one competitor known to show up well in AI answers, if you can spot one, plus one or two closer peers. Comparing against a clear leader shows what a strong result looks like, rather than only comparing among brands that may share the same gap.
Yes. Run the whole method in a plain spreadsheet: one row per query, one column per engine, and simple yes/no/position notes in each cell. You do not need a paid AI-tracking tool for a useful first benchmark — the manual version produces a real, comparable dataset, and it is the same method paid tools automate behind the scenes.
Date each run and keep every prior version rather than overwriting it. The second and third benchmarks are what reveal a trend; a single snapshot cannot separate a real shift from normal answer-to-answer variation. For tool options later, see best ChatGPT visibility trackers and GEO tools for small business.
A low share of voice on comparison and best-of questions usually points to missing or weak comparison content. A low score on direct brand questions — or worse, inaccurate facts about the brand itself — usually points to inconsistent information across the web that the model is pulling from. Either finding is a concrete starting point, not a vague sense that "AI doesn't know about us." For the difference between this work and traditional SEO, see the SEO vs GEO guide; for the full GEO picture, see GEO for Small Business.
No. Answers vary between sessions even for the same query, which is why a single screenshot is not a benchmark. Run each query more than once per engine, or track the same fixed list across several monthly cycles, so you are reading a pattern instead of one pass.
Use a spreadsheet: 20 buyer questions as rows, one column per engine (ChatGPT, Perplexity, Gemini, Claude), and a yes/no plus position note in each cell. Reuse the same rows next month. That is a real benchmark. Paid trackers automate the same method; they are not required for the first read.
Fifteen to thirty is the useful band. Fewer than ten swings too hard when one answer changes. More than thirty is hard to re-run by hand. Mix best-of, comparison, how-to, and direct brand questions so the sample matches how people actually ask AI engines.
Three to five alternatives a customer would actually consider. Include one name that already shows up well in AI answers if you can spot one, plus closer peers. A long list of every business in the category hides the gap you can actually close.