A client called me in a mild panic last month. Her keyword tool had just told her that "affordable geo tool for a small dtc brand" — the exact phrase she'd built a whole page around — had a monthly search volume of zero. She wanted to know if she should pull the page down.
I pulled up our own Search Console instead. That phrase, on our own site, sits at an average position of 4.3 with real impressions this month. Zero volume, real position, real traffic. The tool wasn't lying — it genuinely can't see that query. It just isn't built to.
Ignore the volume number on question-form long-tail, hyper-local attribute combos, and comparison/alternative queries — keyword tools chronically underreport these, and they're exactly the phrasing an LLM retrieves when it answers a specific question. Trust the volume number on broad category terms with real commercial intent and on competitor gap analysis, where tools still do genuinely useful work.
What a keyword tool is actually measuring
Ahrefs, SEMrush, and Google's own Keyword Planner all build their volume estimates from the same rough family of inputs: aggregated clickstream panels and, for Keyword Planner specifically, historical Google Ads auction data. That's a deliberate design choice, not a bug — Google Ads exists to sell advertisers a forecast of how often an ad might show for a given term, and Keyword Planner's numbers are a byproduct of that forecasting engine, rounded and bucketed for millions of simultaneous users rather than reported exactly per query. Google's own developer documentation on how Search works describes ranking as a process built around matching a huge, constantly shifting set of real user queries to content — the ranking system itself does not report per-query volume to publishers at all. Volume estimates are a third-party reconstruction of demand, not a direct read of it.
That reconstruction works reasonably well at scale, for terms with enough historical ad spend and enough repeated, near-identical phrasing to average into a stable number. It works badly — often reporting zero — for anything specific enough that the same three or four attributes rarely got typed in exactly that order before. Which is precisely the shape of question a person now types into a chat window instead of a search box.
Where keyword tools still earn their subscription
I'm not telling you to cancel Ahrefs. Two jobs still belong to a keyword tool, and neither has a good substitute:
- Broad category terms with real commercial intent. "Miami wedding photographer," "AI visibility audit," "geo agency" — head terms with genuine volume and buyer intent still deserve a page, and the tool's relative sizing between them is still directionally useful for prioritizing which broad category to build authority in first.
- Competitive gap analysis. Pulling a competitor's ranking pages and finding the terms they rank for that you don't is a job keyword tools do better than almost anything else — it's structural comparison, not demand forecasting, and the data holds up fine for that use.
Both of those jobs share a trait: they're asking the tool a question it was actually built to answer — "how big is this bucket relative to that bucket," and "what does this specific competitor rank for." The trouble starts when you ask it a third question it was never built to answer: "does anyone care about this exact phrase."
Three query types where the tool actively misleads you
1. Question-form long-tail
"What does an AI visibility audit cost" reads like a low-value long-tail term in a volume column. It's also close to verbatim how someone asks ChatGPT or Perplexity the same question when they're three steps into deciding whether to hire anyone. We built a full page around exactly that question because the phrasing itself — not a rounded volume estimate — told us it mattered. A keyword tool sees a low bucket. A person with a real, unresolved question sees the fastest way to get an answer, and increasingly types it straight into a model instead of a search box.
2. Hyper-local attribute combos
Stack a business type, a neighborhood, and a specific service together — "brickell chiropractor," for instance — and volume tools flatten it toward zero because too few historical searches match that precise string. On our own site that exact phrase currently sits at position 11.6 with real impressions in the last 28 days — striking distance, not nonexistent demand. A generic tool built for national or even metro-level averages structurally cannot represent a query that only matters within a six-block radius, and a chat assistant answering "who's a good chiropractor near Brickell" doesn't care what the volume tool thinks — it's retrieving whatever page actually names the neighborhood and the service together.
3. Comparison and alternative queries
"X vs Y" and "alternatives to X" queries used to route a click to a listicle. Now the engine frequently answers the comparison directly in the chat window, and the page that gets cited is the one that names both entities clearly and states a real difference — not the one that ranked for a keyword-tool-approved head term. We cover the mechanics of this shift in geo tools vs. a geo agency: tools show you what exists, they don't tell you which comparison a buyer is actually deciding between this month.
Run this five-minute check before you kill a query
Before you drop a page or a topic because a tool showed zero or near-zero volume, run this instead of trusting the number:
- Type the exact phrase into ChatGPT or Perplexity the way a customer would ask it — not a keyword-tool version, the real sentence, attributes and all. If the engine gives a confident, specific answer, real demand exists for that phrasing even if no ad ever bid on it.
- Check your own Search Console for that phrase or close variants under Performance → Queries. A query already earning impressions and a mid-page position, even with zero clicks, is proof of demand a third-party volume tool can't see.
- Ask your sales or support team whether that specific phrasing — business type, location, budget tier, feature — comes up in real conversations. If it's a recurring objection or question, it belongs on the page regardless of what the tool says.
- Only after those three checks come back empty should a zero-volume reading actually mean "skip it." At that point you're not ignoring the tool — you've replaced it with better evidence and it happened to agree.
A concrete example: the query a keyword tool would have killed
Take a hypothetical, but entirely typical, small business: a two-chair hair salon in a specific suburb, $2,000/month marketing budget, deciding whether to build a page around "best hair salon for curly hair near [suburb] under $80." A keyword tool returns zero volume — too many attributes stacked together, no ad history to draw from. Run it through ChatGPT instead, phrased the way a real person would ask it, and the model answers directly, pulling from whatever local pages actually name the neighborhood, the specialty, and a price range together. That's the situational, attribute-rich phrase this whole piece is about: exactly the kind of query an LLM retrieves because it's specific enough to match confidently, and exactly the kind a volume column will always underreport because it's too specific to have accumulated ad-auction history. The fix isn't a better tool. It's writing the page for the actual sentence instead of the bucketed version of it.
Why this gap is widening, not closing
OpenAI's own documentation on GPTBot and its retrieval crawlers, Anthropic's published crawler documentation, and Perplexity's bot and crawler guides all describe systems built to fetch and index full pages for retrieval at query time, not to report aggregate search volume back to publishers the way Google Ads does. None of these engines publish a public "how many people asked this" number, which means there is currently no direct replacement for a keyword tool's volume column in the AI-answer world — only proxies: your own Search Console impressions, direct prompting, and what your team hears from real customers. Treat those three as the actual instrument now, and the keyword tool as one input among several rather than the deciding vote.
The position this puts us in
Here's the part competitors selling keyword-tool subscriptions won't say out loud: a tool built to sell advertisers a forecast of ad auction behavior was never going to be the right instrument for measuring whether an AI model will retrieve your page. It's a useful, still-relevant tool for two specific jobs — sizing broad categories and mapping competitor gaps — and a genuinely misleading one for the situational, attribute-rich queries that increasingly decide whether you get cited at all. Running both checks, the volume number and the direct-prompt test, on every query before you commit a page to it is the only honest process. We build that dual check into every content plan we run, including our own — see the real content structure that wins citations in the content that gets quoted.
Where this fits with the rest of your GEO work
Keyword judgment is one input into a bigger system, not the whole system. Once you know which situational queries actually matter, the next question is whether your pages are structured to answer them directly — see FAQ pages vs. FAQ schema for AI for how to format the answer itself so it's liftable. And if you're not sure whether you're already showing up for the queries that matter, run a DIY AI visibility audit first — free, and it tells you in an afternoon whether this whole exercise is even your biggest gap right now, or whether a different metric entirely deserves the attention first.
Questions people ask
No. Keep them for broad category terms with real commercial intent and for competitive gap analysis — seeing which pages a competitor ranks that you don't. Stop trusting the volume number on question-form long-tail, hyper-local attribute combos, and comparison queries. Those are exactly the phrases keyword tools underreport and LLMs retrieve.
Keyword tools estimate volume from aggregated, historical Google Ads and clickstream data, bucketed and rounded for tools that serve millions of users. A specific, attribute-rich phrase — a business type, a location, a budget tier, all in one query — rarely has enough historical ad-auction data to register above zero, even when real people type close variants of it into a chat window every day.
Your own Search Console queries, your sales team's actual objections and questions, and direct prompting of ChatGPT and Perplexity with the situational phrasing a real customer would use. These sources surface the attribute-rich, low-volume-but-real queries a keyword tool was never built to see.
Not sure which queries actually matter?
We run the full dual check — volume tools plus direct prompting plus your own GSC data — as part of every AI visibility audit.
Get a free AI Visibility Audit →
