How to get cited by ChatGPT
Buyers now ask a model for "the best tool for X" and get the same five names back. Here is how that list is actually assembled - and what genuinely moves you into it.
The short answer
To be cited by an AI assistant you have to clear three gates, in this order: be retrievable (the machine can fetch and read your content), be attributable (it can tell what you are and who you're for), and be quotable (you published something worth extracting). Most people skip straight to the third and wonder why nothing happens.
The uncomfortable part is that the third gate is rarely about your own pages at all. Models synthesise how the web describes your category. If nobody independent has written about you, there's nothing to retrieve, and no amount of on-page work fixes that.
Classic SEO asks "does Google trust this domain enough to rank it?" Citation asks a different question: "is this the clearest available source for the specific thing I was just asked?" Those questions have different answers, which is why sites with almost no authority get cited next to household names.
How citation actually works
When someone asks an assistant for a recommendation, roughly three things happen. It's worth understanding because each stage fails differently, and the fix depends entirely on which stage you're failing at.
Stage one - the query gets rewritten
Your buyer types "best tool for tracking freelance invoices." The model doesn't search for that string. It expands it into several more specific queries, and those rewritten queries are what actually get retrieved against. This is why targeting the literal phrase your buyer types is less useful here than it is in classic SEO: you're optimising for the machine's paraphrase of the question, not the question. Covering the concept thoroughly beats matching a keyword exactly.
Stage two - candidates get retrieved
Sources come back from a search index, from the model's training data, or from a live fetch, depending on the assistant and the question. Three practical consequences:
- If your content isn't in the raw HTML, you may not exist at this stage. Client-rendered pages are a common and completely silent failure - an alarming share of AI-built sites are invisible here from birth.
- If you block AI user agents in
robots.txt, you've opted out of the surface. Plenty of sites did this in 2024–25 and are now quietly wondering why they're never named. - Recency matters more than in classic search. Undated content is discounted; a dated, recently updated page is a stronger candidate.
Stage three - the answer gets synthesised
The model writes an answer and attaches citations to the sources it leaned on. It favours passages that are self-contained: a paragraph that makes sense lifted out of your page, with the claim and its support in the same place. This is the single most actionable insight on this page, and it's why the writing style here looks the way it does. A brilliant argument spread across six paragraphs with pronouns pointing backwards is nearly impossible to quote. One tight paragraph that answers the question completely is trivially quotable.
The prerequisites nobody checks
Before any content strategy is worth doing, these have to be true. They take an afternoon and they're the reason a lot of otherwise-good sites are invisible.
| Check | How to test it in a minute | If it fails |
|---|---|---|
| Content in raw HTML | View source, or load your page with JavaScript disabled. Can you read the actual content? | Server-render or pre-render. Nothing else matters until this passes. |
| AI agents allowed | Read your robots.txt. Look for GPTBot, OAI-SearchBot, PerplexityBot, ClaudeBot, Google-Extended. | Decide deliberately. Blocking is a legitimate choice - but then stop wondering why you're not cited. |
| Category stated plainly | Does one sentence on your homepage say what kind of thing this is, for whom? | Write it. "We help teams move faster" tells a model nothing it can use. |
| Dates visible | Does the page show when it was published and updated? | Add them, and mean them. Undated content gets discounted. |
| Structured data | Does the page carry Organization, Product or FAQ schema? | Add it. It's not magic, but it removes ambiguity about what you are. |
What actually gets quoted
Once you're retrievable and attributable, the question is what to publish. Some formats are dramatically more quotable than others, and the pattern is consistent enough to plan around.
Original numbers beat opinions
A model synthesising an answer needs something specific to attach to a source. "Onboarding is important" is unquotable - every page says it. "Across four sites we launched in 2026, impressions accumulated within days while click-through rates sat between 0.1% and 0.6% for the first three weeks" is quotable, because it's a fact that exists nowhere else. Academic work on generative-engine optimisation has found that adding statistics and citations measurably lifts a source's visibility in AI answers, which matches what practitioners see. If you have data nobody else has - and every operating business does - publishing it is the highest-leverage GEO move available to you.
Direct answers beat build-ups
Put the answer immediately after the question as a heading, in one self-contained paragraph, then elaborate. Journalism's inverted pyramid, for a reader that stops reading as soon as it has what it needs.
Comparisons beat claims
"X vs Y" and "alternatives to X" pages are cited heavily because buying questions are comparative by nature. Being genuinely fair about where you lose is not a weakness here: a page that only says you win reads as marketing and gets treated as such, by models and humans alike.
Being described elsewhere beats describing yourself
The hardest one, and the most important. Models weight how the web talks about you, not only what you say. Getting genuinely mentioned in the roundups, forum threads and comparisons that already get cited does more than any on-page change. This is unglamorous work - being useful in public where your category gets discussed - and it's the part most people skip because it doesn't feel like optimisation.
What doesn't work
- Keyword stuffing for a machine that reads meaning. Repetition was a signal for lexical matching. Embeddings don't reward it, and it makes your page worse to quote.
- "llms.txt" as a magic file. Proposals exist for machine-readable summaries and there's no harm in publishing one, but no major assistant currently commits to consuming it as a ranking input. Treat it as cheap insurance, not strategy.
- Publishing AI-written content at volume. You're competing to be the source a model prefers over its own generation. Text indistinguishable from what it would have written anyway gives it no reason to cite you.
- Checking once and declaring victory. Citation is probabilistic. The same prompt returns different sources on different days. A single lucky check tells you almost nothing.
- Blocking AI crawlers and hoping. Worth stating plainly: you can't be quoted from a page nothing is allowed to read.
How to measure it
This is where most people give up, because a citation produces no click and therefore never shows up in analytics. You have to go and look, deliberately and repeatedly.
- Fix a prompt set. Twenty to forty prompts a real buyer would actually type - "best X for Y", "alternatives to Z", "how do I…" - written once and then left alone. Changing the prompts between runs destroys comparability, which is the only thing that makes the exercise worth doing.
- Run them on a schedule across the assistants your buyers use, and record who gets named, not just whether you did. Competitor share-of-voice is the more useful number.
- Watch referral traffic separately. Visits from assistant domains are small in volume and unusually high in intent - one 2026 analysis put AI-referred conversion at several times classic organic. Segment it or you'll never see it.
- Expect noise. Run the same prompt several times before drawing any conclusion. Treat it as sampling a distribution, not reading a rank.
This site is a brand-new domain running the exact process on this page, with the results published as they happen - including the failures. Predictions are logged with dates before they're known and graded when they come due, on the receipts page. If it doesn't work, that will be written there too.
An honest note on certainty
Anyone claiming a settled playbook here is selling something. This discipline is roughly two years old, the systems change without notice, and nobody outside the labs can see the ranking mechanics. What's written above is drawn from how retrieval-augmented systems demonstrably behave, published research on generative-engine optimisation, and first-party testing - and where the evidence is thin, this page says so.
The reason to act now anyway is that the cost of being wrong is low and the cost of being late is high. Almost everything here - crawlable content, plain statements of what you are, original data, honest comparisons - is work that pays off in classic search regardless. You're not making a bet on GEO. You're doing good technical marketing that happens to also work on the surface everyone is moving toward.
Want to know if you're cited today?
The GEO Audit runs your category's real buying prompts across the major assistants, records who gets named instead of you, and returns the specific fixes - with a 90-day recheck so you find out whether they worked.