How do I get my content cited by AI?
Once your pages are readable and retrievable, citation stops being a technical problem and becomes a writing one. This is the craft of producing the sentence worth quoting.
Once your pages are crawlable and readable without JavaScript, getting cited stops being a technical problem. What remains is a writing problem, and it is a specific one with a specific test.
The unit is the passage
A search engine ranks pages. A generative engine quotes sentences.
That single distinction reorganises everything about how you write. When a model assembles an answer, it is not choosing a page to send someone to — it is choosing two or three sentences to reproduce and attribute. Your page might be the most thorough treatment of a subject on the internet and still lose to a competitor who wrote one clean paragraph, because the competitor’s paragraph could be lifted and yours could not.
If a search engine ranks pages, a generative engine quotes sentences. Citation is the craft of writing the sentence worth quoting.
Everything below is a test you can apply to a passage in under a minute.
Test 1: does it stand alone?
Copy any three consecutive sentences from the middle of your page. Paste them into a blank document and read them cold.
Do they still answer something? Or do they depend on the previous paragraph, a pronoun with no antecedent, a chart above them, or a definition established two sections earlier?
If they cannot survive the copy-paste, they cannot be extracted. A model assembling an answer will use someone else’s self-contained paragraph instead, and it will do so silently — you will never see the near-miss.
The fix is mechanical. Name the subject rather than referring back to it. Restate the context in a clause rather than assuming it. Put the answer in the same sentence as the thing it answers. This produces prose that reads slightly more repetitive to someone consuming the whole page and dramatically better to everyone else, which is the correct trade in 2026.
Test 2: is there anything to attribute?
Read your passage and ask what a model would credit to you.
“Many companies find that shorter forms improve conversion” contains nothing attributable. A model can restate that idea from a dozen sources, so it will paraphrase and cite none of them — or cite whoever said it most concretely.
“In our 2026 analysis of 4,000 signups, reducing the form from nine fields to four raised completion by 31%” is unattributable to anyone else. There is no way to use that sentence without naming where it came from.
The Princeton and IIT Delhi GEO study — Aggarwal et al., ACM SIGKDD 2024 — tested optimisation methods across roughly 10,000 queries and found the three strongest were Statistics Addition, Quotation Addition and Cite Sources, delivering relative visibility improvements of roughly 30–40 percent. Lower-ranked sources benefited most: a source ranked fifth that added citations saw a reported relative lift above 100%.
Two honest caveats: those are maximum figures under favourable conditions rather than averages, and the study predates the current generation of engines. The direction has held up well regardless.
Test 3: is it backed?
This one is counterintuitive: citing other people makes you more citable, not less.
A model assembling an answer carries its own credibility risk. A page that names its sources, links to primary research and dates its claims is safer to quote than one making equivalent assertions from nowhere. When a model uses your passage, it inherits your citations as supporting evidence.
There is a reputational mechanism too. Pages that cite sources read as analysis; pages that assert read as marketing. Models are demonstrably sensitive to that distinction, and so are the humans who decide whether to link to you — which then feeds back into the corroboration layer.
Practically: name the organisation, date the figure, and link where possible. “Research suggests” is worth nothing. “Pew Research, tracking approximately 68,000 queries, found clicks fell from 15% to 8% when an AI summary appeared” is worth a great deal.
See which of your pages get cited
A free scan runs your category’s questions across ChatGPT, Gemini and Perplexity and records every URL they cite.
Run my free scanTest 4: is it yours?
The hardest test and the most durable advantage.
Ask what your page contains that the current top five results do not. If the honest answer is nothing, the page is a restatement — and restatements are not cited, because there is always an older, better-linked version of the same claim.
Four kinds of genuinely original material work, in rough order of durability:
- Original data. A survey, a benchmark, an analysis of your own operational data. Nobody else has it and everyone writing about the topic has to cite it.
- Primary expertise. Things only someone doing the work would know — real failure modes, actual costs, what goes wrong at scale. Generic content cannot fabricate this convincingly.
- A stated position. Most category content is interchangeable because everyone hedges. A clear, defensible view — including where your own product is the wrong choice — is memorable and quotable.
- Structured comparison. A genuine side-by-side you actually ran, with criteria stated. Comparisons are among the most frequently extracted formats.
If you have no original data, the cheapest source is your own operations: support ticket categories, a survey of fifty customers, an internal comparison you already ran for a purchasing decision. Most companies are sitting on publishable data they have never thought of as content.
Formats that get lifted
Some structures are disproportionately likely to be extracted, because they map cleanly onto how an answer gets assembled.
Definition paragraphs. “X is Y that does Z” in one sentence, immediately under a heading asking what X is.
Numbered procedures. Steps with a verb first. These are lifted almost verbatim for how-to prompts.
Comparison tables. Criteria down the side, options across the top. Models read these reliably and reproduce them as prose.
Direct question-and-answer blocks. A question as a heading, two to three sentences beneath it, no preamble. This is the single highest-yield structure available and it is why FAQ sections earn citations far out of proportion to their length.
Explicit criteria lists. “Three things decide whether X works: A, B and C” followed by a paragraph on each. Legible to skimming humans and trivially extractable.
The hedging problem
The most common reason good writing fails this test is professional caution.
“It depends on your circumstances, but in many cases it may be advisable to consider whether a shorter form could potentially improve outcomes” is unquotable. There is nothing in it. A model cannot extract a claim from a sentence that carefully avoids making one.
The fix is not to become reckless. It is to state the answer, then qualify it in the following sentence. “Shorter forms convert better. In our analysis, cutting from nine fields to four raised completion 31% — though the effect narrows for high-intent audiences who would have completed either version.”
The claim survives extraction. The qualification is there for anyone reading the page. Nothing has been overstated, and the passage has become usable.
Five patterns that prevent citation
The buried lede. The answer exists and arrives in paragraph six, after context, history and a definition nobody asked for. The model uses someone else’s paragraph one. This is the single most common cause of a good page never being quoted.
The everything page. One URL attempting to cover a topic, its history, its variants, its tooling and its future. It ranks acceptably and gets cited for nothing, because no passage in it answers any single question completely. Splitting it usually produces three pages that each outperform the original.
The unattributed assertion. Strong claims with no source, no date and no number. These read as marketing to a model exactly as they do to a person, and they are paraphrased away rather than quoted.
The pronoun chain. Paragraphs beginning “This means…”, “It also…”, “They typically…”. Perfectly readable in sequence and completely unextractable, because lifted out of context they refer to nothing.
The JavaScript answer. Everything above done correctly, inside a component that only renders client-side. Invisible to every AI crawler regardless of quality.
A workflow that produces it
Four passes over a draft, each quick, each catching a different failure.
- The answer pass. Read only the first two sentences under each heading. Do they answer that heading? If a section opens with context or background, move the answer up.
- The specificity pass. Find every “many”, “often”, “can help”, “may improve”. Each is replaceable with something concrete or deletable.
- The standalone pass. Copy three random paragraphs out and read them cold. Fix the pronouns and name the subjects.
- The subtraction pass. Delete anything that restates what is already on the first page of results. This usually removes fifteen to twenty percent and improves the piece, because what remains is the part that was actually yours.
One last discipline: check that the passage exists in raw HTML. AI crawlers do not execute JavaScript — Vercel and MERJ recorded zero JavaScript execution across more than 500 million GPTBot fetches — so a perfectly crafted passage inside a client-side-rendered component is invisible regardless of how quotable it is. The best writing in the world does not survive being unreadable.
Sources
- Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande — “GEO: Generative Engine Optimization”, ACM SIGKDD 2024 (arXiv:2311.09735); Statistics Addition, Quotation Addition and Cite Sources as top-performing methods, ~30–40% relative visibility improvement across ~10,000 queries.
- Vercel and MERJ — analysis of 500M+ GPTBot fetches finding zero JavaScript execution; same behaviour for ClaudeBot and PerplexityBot.
- Pew Research Center — approximately 68,000 real search queries; clicks fell from 15% to 8% of visits when an AI summary was present.
Frequently asked questions
How do I get my content cited by AI?
Write passages that stand alone, make claims specific enough to attribute, back them with named sources, and include something nobody else has. Models lift two or three sentence chunks rather than pages, so the working unit is the passage. The Princeton and IIT Delhi GEO study found adding statistics, quotations and cited sources were the highest-performing methods, improving AI visibility by roughly 30 to 40 percent.
What makes a passage quotable by an AI engine?
It answers a question completely without needing the surrounding paragraphs, contains a concrete claim a model can reproduce without inventing anything, and comes from a source the model has reason to trust. Test yours by copying three sentences out of context and reading them cold — if they no longer make sense, they cannot be lifted.
Does adding statistics really improve AI citations?
The Princeton and IIT Delhi study tested optimisation methods across roughly 10,000 queries and found Statistics Addition, Quotation Addition and Cite Sources were the three strongest, with relative visibility improvements of roughly 30 to 40 percent. Those are maximum figures under favourable conditions rather than averages, but the direction has held up consistently in practice.
Should I cite my sources if I want to be cited?
Yes, and it is counterintuitive. Citing authoritative sources makes your page safer to quote, because a model assembling an answer inherits your citations as supporting evidence. Pages that name their sources read as analysis rather than assertion, and that distinction affects selection.
Does long content get cited more?
No. Length correlates with thoroughness rather than causing citation, and padding actively hurts — it dilutes the extractable passage and buries the answer. A 600-word page that answers one question completely is more citable than a 3,000-word page that answers six partially.
How do I know if my content is being cited?
Run your category’s real questions through ChatGPT, Gemini and Perplexity and record which URLs appear as sources. Perplexity is the easiest place to see this because it cites openly and updates fast. Store the results, because answers vary between identical runs and only a trend across several runs is meaningful.