AI Visibility / GEO · The Darkroom

The content that gets quoted — the pattern behind it

We run our own playbook articles through the same AI visibility checks we sell. Some paragraphs get lifted verbatim into answers. Most don't. Here's the structural pattern that keeps separating the two, with a before/after you can apply today.

2026-07-25 · 9 min read · by Italo Campilii
Sedimentary rock layers on a coastline, a natural structure built one visible layer at a time

A few weeks into publishing daily on this site, I started doing something a little obsessive: after each Darkroom post went live, I'd feed the URL and a handful of likely buyer questions into ChatGPT, Perplexity, and Google AI Mode and watch what came back. Not to check if we ranked — to check whether any actual sentence of ours showed up, word for word or close to it, inside the answer.

The pattern that emerged wasn't about topic, and it wasn't about length. Two posts on nearly identical subjects, published the same week, with the same research behind them, got treated completely differently. One had three separate paragraphs quoted across different engines within days. The other, arguably better-researched, got cited zero times in a month. The difference wasn't what we said. It was how the sentence carrying the answer was built.

That's the thing worth writing about, because most advice in this space is still "write good content" repeated with more buzzwords. Good content is table stakes. What actually determines whether a passage gets lifted is a narrower, more mechanical thing, and once you see it, you can't unsee it on your own site.

The short answer

AI answer engines don't cite pages, they cite passages — small, self-contained chunks of text scored against a query independently of the page around them. The passages that win share four traits: the answer comes in the first sentence, the paragraph makes sense with zero surrounding context, the claim is specific and attributed rather than vague, and the heading above it is phrased the way a person would actually ask the question. You can audit any page against those four traits in about ten minutes.

Why passages, not pages, are the unit that matters

This is the part most business owners haven't internalized yet: modern answer engines are built on retrieval, not "reading the page." A query comes in, the system finds candidate chunks of text across many sources, scores each chunk for relevance, and stitches the highest-scoring chunks into an answer with citations attached. The page is just the container the chunk happened to live in. That architecture — retrieval-augmented generation, first formalized in the 2020 Facebook AI Research paper that coined the term (Lewis et al., "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks") — is the mechanical reason a mediocre page with one great paragraph can out-cite a great page with no standout paragraph.

Google's own documentation on how Search understands content reinforces the same unit-of-meaning idea from a different angle: structured data guidance exists specifically because search and AI systems need unambiguous signals about what a specific piece of content on the page actually is and means, not just that the page exists. Schema.org's own getting-started documentation describes the vocabulary as a way to mark up "the little pieces" of a page — a person, a price, an answer — precisely because machines increasingly work at that granularity, not the page level.

Once you accept that the passage is the unit, a lot of SEO instinct stops applying cleanly. A 3,000-word comprehensive guide doesn't win by being comprehensive. It wins if, somewhere in those 3,000 words, there's a paragraph that answers one specific question so cleanly it can be lifted whole. The rest of the page still matters for ranking and trust — but the citation itself is won or lost at the paragraph level.

Paragraph 1Paragraph 2Paragraph 3Paragraph 4ONE PAGERetrieval scorerranks chunks by querySource BCited answerSource CAnswergiven
One paragraph out of four gets scored, cited, and stitched into the answer — the other three, and the page around them, are along for the ride.

The four traits of a passage that gets lifted

1. The answer is in the first sentence

Retrieval scoring rewards density. A paragraph that opens with three sentences of throat-clearing before it gets to the point buries the exact span a system would need to pull cleanly. Lead with the answer, then explain it.

2. It makes sense with zero surrounding context

If a paragraph depends on "as mentioned above" or an unresolved "it" or "this" to make sense, it can't be lifted out of the page and still read correctly — and answer engines lift out of the page constantly. Every citable paragraph should survive being read completely alone.

3. The claim is specific and attributed, not vague

"Many businesses struggle with this" is not citable. "OpenAI documents GPTBot's crawling behavior directly at platform.openai.com/docs/gptbot" is citable, because it's a checkable, specific fact with a named source attached. Attribution reads as trustworthy to a system built to avoid hallucinating.

4. The heading above it mirrors a real question

Answer engines are frequently matching a user's literal question against your heading and the paragraph beneath it. A heading like "Overview" scores worse than "What is Google AI Mode and how is it different from AI Overviews?" — the second is close to verbatim what someone actually typed, and Google's FAQPage structured data documentation exists precisely because question-and-answer pairs are a recognized, favored content shape.

A real before/after, from our own drafts

Here's an early draft paragraph from one of our own posts, next to the version we shipped after applying the four traits above. Nothing here is a client result or a fabricated statistic — it's our own editing process, shown honestly.

Before — not citable

When it comes to making sure your content can be found by AI systems, there are a lot of factors that come into play, and robots.txt is definitely one of the more important ones that people tend to overlook, especially since it's been around for a long time and a lot of the old rules don't necessarily apply anymore in this new landscape.

After — citable

Robots.txt controls whether GPTBot, ClaudeBot, and PerplexityBot can fetch a page at all. OpenAI, Anthropic, and Perplexity each publish and honor their own user-agent token, so a blanket bot rule written years ago can silently block AI crawlers a site owner never intended to exclude.

The "after" version answers a specific question in its first sentence, names the actual crawlers, cites the fact that each publishes its own token, and would still make complete sense if it were the only sentence a system pulled. The "before" version is 90% filler wrapped around a real point that never gets stated plainly.

What the research backs up, and what it doesn't

Academic interest in this exact question is recent but real. A 2023 paper out of Princeton, Georgia Tech, IIT Delhi, and the Allen Institute, "GEO: Generative Engine Optimization" (Aggarwal et al.), was among the first to formally study which content-editing strategies change how often a source gets cited by generative answer engines, testing structural changes like adding statistics, citations, and quotations against a benchmark of real queries. Their central finding lines up with what we're describing here: structural and attribution changes to existing content measurably shifted citation behavior, independent of the underlying facts changing.

Where I'd push back on some GEO commentary is the implication that this is a trick you can layer on top of thin content. It isn't. The four traits above only work on a paragraph that's already true, specific, and useful — they're a filter for whether that good information is packaged in a way a retrieval system can use, not a substitute for having good information in the first place. For the wider structural rules beyond individual passages, our guide on how to structure content for AI extraction covers page-level architecture, and schema markup: the language AI actually reads covers the machine-readable layer on top of the prose.

The 10-minute audit — run it on your own best page

Step-by-step
  1. Pick your best-trafficked page and copy out every paragraph into a plain list. You're about to score each one independently, so strip the page's visual context out of your own head too.
  2. Read each paragraph as if it were the only sentence pulled into an answer. Does it still make sense? Cross out any that lean on "as above," an unresolved pronoun, or the previous paragraph to land.
  3. Check where the answer sits in the paragraph. If the direct answer is in sentence three or later, move it to sentence one and let the explanation follow.
  4. Circle every vague claim — "many," "often," "studies show" with no source — and either attach a real, checkable attribution or cut the sentence.
  5. Read your H2s and H3s as if a stranger typed them into a search bar. Rewrite any heading that reads like an internal label ("Overview," "Our Approach") into the actual question it's answering.
  6. Re-test with the same buyer questions in ChatGPT, Perplexity, and Google AI Mode after you ship the edits, and compare. This is the same loop we run on our own posts, and it's the fastest way to see whether a structural change actually changed anything.

Ten minutes is enough to find the two or three paragraphs on a page most worth rewriting. You don't need to touch the whole page — you need the one passage that would actually answer someone's question if it were the only thing they read.

Why this matters more than another keyword

Chasing another keyword assumes the old game: rank a page, win a click. The passage game is different — a system reads your paragraph, decides it's the cleanest answer available, and puts your name next to it whether or not anyone ever visits the URL. That's a lower-effort, higher-leverage target than most people treat it as, precisely because almost nobody is editing at the paragraph level yet. Most sites are still optimizing pages. The ones getting quoted are optimizing sentences. See also write extractable answers AI can lift and the anatomy of an AI-citable page for the adjacent pieces of this system, and how AI summarizes long content if your pages run long and you want to know which sections survive the cut. The full library of these teardowns lives on The Darkroom.

Questions people ask

What makes a paragraph more likely to get quoted by an AI engine?

A quotable paragraph answers one question directly in its first sentence, makes sense if it's the only sentence a reader sees (no unexplained pronouns or "as mentioned above"), and attaches a specific, attributable fact instead of a vague claim. Passage-retrieval systems score chunks of text on their own, not the page as a whole, so a paragraph has to stand on its own to survive that scoring.

Do AI engines cite whole pages or specific passages?

Specific passages. Retrieval-augmented systems break a page into chunks, rank each chunk against the query, and pull the highest-scoring chunks into the answer, a pattern documented across the retrieval-augmented generation research this generation of engines is built on. A page can have one excellent, citable paragraph and nine mediocre ones and still get quoted, or it can be well-written overall and never get picked because no single paragraph answers a specific question cleanly.

Does schema markup help my content get cited by AI?

Schema markup, applied correctly per the Schema.org vocabulary, gives crawlers a structured, unambiguous description of what a page is and what its parts mean, which reduces the chance of misclassification. It doesn't rewrite bad prose into good prose. Clear FAQPage, Article, and HowTo markup around content that's already well-structured tends to help; markup slapped onto vague or self-referential paragraphs won't fix the underlying extractability problem.

— Italo & Ale
written from the studio floor · developed in the darkroom

Want this done for you?

Want to know which of your own paragraphs are actually getting cited, and which are invisible to AI? Get an AI visibility audit from Acromatico.

Get a free AI Visibility Audit →

Want to know if AI actually recommends your brand?

Run your free AI visibility audit →Start a 90-day GEO sprint →