Muhammad Basim
SEO

SEO When AI Answers the Question First

By Muhammad Basim·
SEO When AI Answers the Question First

For twenty-five years, the deal was simple. Google found your page, showed a blue link, and someone clicked it.

That deal is being renegotiated without you in the room.

Ahrefs analysed 300,000 keywords and compared click-through rates before and after AI Overviews rolled out. For keywords without an AI Overview, the position-one click-through rate fell from 7.6% to 3.9% over two years. For keywords with one, it collapsed from 7.3% to 1.6%. After controlling for the general downward trend, AI Overviews correlated with roughly a 58% reduction in clicks for top-ranking pages — up from the 34.5% they measured a year earlier.

Chartbeat, tracking more than 2,500 news sites, found Google search referrals down about a third across 2025. And Pew found that when an AI Overview is present, people click a traditional result 8% of the time versus 15% when it isn't.

So the question isn't whether this is happening. It's what still works — and the honest answer is that a lot does, but not the parts most people are optimising.

The short version

Ranking and being cited are two different games now. You can rank first and never be quoted. You can rank on page three and be quoted constantly.

What the research consistently points at:

  1. Answer the question in the first paragraph — a large share of citations come from the opening third of a page
  2. Use specific, named things — entities, proper nouns, real tools and people, not vague categories
  3. Include original data and statistics — the single most reliable citation lever in the published studies
  4. Write definitively — "X is defined as" beats "X can sometimes be thought of as"
  5. Keep it fresh — dated content gets skipped for anything time-sensitive
  6. Stay crawlable by the retrieval bots — which is a different decision from the training bots

And one thing that doesn't work, despite the noise: llms.txt. More on that below.

First, a warning about the numbers

Every stat in this space comes with an agenda attached, and you should know that before you read another article about it — including this one.

Most published GEO research comes from companies selling GEO tools. Sample sizes vary from 1,200 pages to 680 million citations. Methodologies are rarely comparable. And the findings genuinely conflict: one study reports 38% of AI Overview citations come from top-10 organic results, another reports 68% don't. (Those are actually the same finding stated from opposite ends, which tells you something about how this gets reported.)

The figures I'm using here come from the more defensible sources — Ahrefs, Pew Research Center, Chartbeat, SparkToro, and Princeton's GEO study — and I've named the source and sample every time. Where the research is thin or self-interested, I've said so rather than quoting a confident number.

Treat any GEO statistic without a named methodology as marketing.

Ranking and citation are different games

This is the conceptual shift, and everything practical follows from it.

Ranking is a competition for position on a results page. Citation is a competition to be the passage a model finds most useful when assembling an answer.

They overlap — ranking well makes you more likely to be retrieved in the first place, and that relationship is strong enough that abandoning traditional SEO would be daft. But they're not the same, and the gap between them is where the opportunity is.

A page that ranks first with a chatty 400-word introduction before it answers anything is a bad citation candidate. A page ranked eighth that answers the question in its first two sentences, with a named statistic attached, is an excellent one.

The corollary matters too: citation doesn't guarantee traffic. Pew found only about 1% of AI Overviews led to a click on a cited source. Wikipedia is the most-cited domain in AI Overviews and still saw human pageviews decline. You can win the citation and get nothing for it.

Which is why the strategy below isn't "chase citations." It's "be the source, and structure your business so being the source is worth something even when the click doesn't come."

What answer engines actually select for

Across the more credible studies, a consistent picture emerges.

Answers near the top. Analysis of ChatGPT citations found a large majority of cited passages come from the first third of a page's content. The intro is where citation is won or lost, which inverts the classic SEO habit of building up to the answer.

Statistics and original data. Princeton's GEO study tested nine optimisation methods across 10,000 queries and found adding statistics was among the two strongest. Independent analyses put original data tables among the highest-citation content types. If you have proprietary numbers, they are your single biggest asset here.

Definitive language. Pages that commit — "X is," "X is defined as" — get cited noticeably more than pages that hedge. Models are looking for statements they can repeat without qualification, and a sentence full of "may," "could," and "in some cases" isn't repeatable.

Named entities. Cited content has measurably higher proper-noun density. Name the tool, the framework, the researcher, the company. The full argument is here.

Quotations and expert attribution. Cited pages carry more expert quotations than uncited ones, and the Princeton study found quotation addition especially valuable in some domains.

Freshness, where it's relevant. A large share of AI citations come from content updated recently, and pages updated within the past year are markedly more likely to be cited. But this is query-dependent — a question about yesterday's news needs current sourcing; a question about the history of Rome doesn't.

Clean structure. Proper heading hierarchy, direct-answer formatting, and schema markup all correlate with citation. Models cite what they can parse cleanly.

Chunking: writing so you can be quoted

Here's the practical technique underneath all of that, and it's the one thing I'd change first.

Write so any single section can be lifted out and still make sense.

An answer engine doesn't read your article. It retrieves passages. If your section on pricing only makes sense because of context established four headings earlier, that section can't be quoted — the model will find someone whose section stands alone.

What that means in practice:

  • Each H2 section opens by stating its own point rather than continuing a thread
  • Definitions are self-contained: "Entity SEO is…" rather than "This approach is…"
  • Pronouns resolve within the section, not across the whole page
  • The specific detail — the number, the name, the threshold — sits inside the passage rather than in a paragraph above it

Analysis of what gets pulled verbatim reinforces this. Numeric formats — statistic lines and table rows — are copied most often, because numbers can't be safely reworded. Plain prose gets paraphrased. So a claim that matters should carry a number in the same sentence, where it travels with the claim.

This is genuinely different from writing for a human reader who arrives at the top and reads down. You're writing for a reader who arrives in the middle and takes one paragraph.

First-hand experience: the thing that can't be synthesised

Every model has read every general explanation of your topic that exists. Writing another one competes with an infinity of identical content, and adds nothing a model couldn't generate itself.

What it can't generate is what happened to you.

Your own client results. The number from your own audit. The mistake you made in 2023 and what it cost. The thing that works in your market that the general advice gets wrong. Your survey of 200 customers.

This is the one durable moat, and it's why the advice "publish original data" keeps appearing at the top of every citation study. Not because models love statistics as a genre, but because a specific verifiable number that exists nowhere else is the one thing a synthesiser genuinely needs a source for.

For a personal-brand site this is your structural advantage over larger competitors. A big content team can out-publish you on volume. They can't have your experience.

Structured data's actual role

Schema markup correlates with citation, and FAQPage, Article, and BreadcrumbList are the ones that come up repeatedly.

But be careful how you read that correlation. Sites with clean schema tend to be sites that are technically competent, well-organised, and professionally maintained — all of which independently predict citation. Schema is probably doing some direct work by helping models parse and verify your content, and probably also standing in for general quality.

Either way the action is the same, because it's cheap: implement Article, FAQPage, and BreadcrumbList properly and validate them. Just don't expect schema alone to move anything if the underlying content buries its answer on screen three. Implementation detail here.

Crawler access: the decision that's easy to get wrong

There's a genuine choice here, and it's more nuanced than "block the AI bots" or "let them all in."

Training crawlers and retrieval crawlers are now separate. GPTBot, ClaudeBot, and Google-Extended are largely about training. OAI-SearchBot, Claude-SearchBot, and PerplexityBot are what fetch content to answer questions with citations.

Which means you can block training while staying eligible for citation, if that's your position. The catastrophic error is blocking the retrieval bots by accident — OpenAI's documentation is explicit that sites opted out of OAI-SearchBot won't appear in ChatGPT search answers, regardless of what GPTBot crawled previously.

Blocking GPTBot doesn't affect your Google rankings. Blocking OAI-SearchBot removes you from ChatGPT's answers entirely. People conflate these constantly.

Check your robots.txt and your CDN, since Cloudflare and similar can be blocking bots at a layer your robots.txt never sees. Full breakdown and a working template.

Measuring any of this

The honest position: you can measure some of it, and a lot is genuinely invisible.

GA4 added a native AI Assistant channel in May 2026, which recognises ChatGPT, Gemini, DeepSeek, Copilot, and Grok — but notably not Perplexity, which still lands in generic Referral. It's also forward-only, so historical sessions keep whatever classification they originally got. And Google's own AI Overviews and AI Mode stay inside Organic Search rather than being broken out.

Beyond that, a substantial share of AI referral sessions arrive with no referrer header at all and land in Direct. Mobile app traffic passes nothing. Zero-click citations — where you're the source and nobody clicks — are invisible by definition.

So build the tracking you can, and hold the numbers loosely. Setup walkthrough.

What still works exactly as before

Worth ending here, because the panic in this space obscures how much is unchanged.

Ranking still matters enormously. Retrieval correlates strongly with organic position. Being invisible in Google means being unavailable to the systems reading Google's results.

Technical SEO still matters. Crawlability, indexation, and site speed decide whether anything else is possible.

Topical depth still matters. Covering a subject thoroughly is what makes you retrievable across the many related queries a model fans out into when assembling an answer.

Links and authority still matter. They feed the ranking that feeds the retrieval.

Genuinely useful content still wins. The formats that survived — original research, tools, strong opinion, deep first-hand guides — are the ones that were always the good stuff. The six formats, in detail.

The uncomfortable summary is that AI search didn't invalidate SEO. It raised the floor. Thin content that ranked on technique alone is being eaten first, and the people complaining loudest are frequently the people who were publishing it.

Frequently asked questions

How do I get cited by ChatGPT?
Answer the question in your opening paragraph, include specific statistics and named entities, write definitively rather than hedging, and make sure each section stands alone when read in isolation. Rank well organically, since retrieval correlates strongly with organic position. And confirm OAI-SearchBot isn't blocked in your robots.txt or at your CDN — that single misconfiguration removes you from ChatGPT's answers entirely.

Is GEO different from SEO?
Partly. The foundations are identical — crawlability, topical depth, authority, and ranking all still matter, and generative engines largely retrieve from conventional search results. What differs is structure: answer engines quote passages rather than sending traffic to pages, so front-loading answers, self-contained sections, original data, and definitive phrasing matter more than they used to. Treat GEO as an additional layer on competent SEO, not a replacement for it.

Should I block AI crawlers?
It depends which ones, and the distinction matters more than the decision. Training crawlers like GPTBot and Google-Extended can be blocked without affecting your Google rankings or your eligibility for AI citations. Retrieval crawlers like OAI-SearchBot, Claude-SearchBot, and PerplexityBot are what make citation possible — blocking those removes you from AI answers. Most businesses seeking visibility should allow retrieval bots regardless of their position on training.

Are AI Overviews killing SEO?
They're killing a particular business model — publishing informational content to collect ad-supported clicks. Ahrefs measured roughly 58% lower click-through for top-ranking pages when an AI Overview appears, and Chartbeat found publisher search referrals down about a third across 2025. But search itself is growing, AI referrals convert at notably higher rates than organic in several analyses, and content that offers something a model can't synthesise still earns visits. It's a redistribution rather than an extinction — though it's genuinely an extinction for thin content.

What to do next

Take your best-performing page and read only its first paragraph. Does it answer the question the page is about, with a specific number or named thing in it?

For most pages the answer is no — the first paragraph sets up the topic and the answer arrives on screen two. That's the single highest-value edit available to you right now, and it takes ten minutes per page.

Then check your robots.txt for OAI-SearchBot. If it's blocked, you're invisible in ChatGPT and it's probably an accident.

Free: The SEO audit checklist.


Related guides

Join the Newsletter

Get practical marketing tactics delivered straight to your inbox.

Muhammad Basim

Written by

Muhammad Basim

Related Articles

Newsletter

Free: The 60-Minute
Email Authentication Fix

A no-fluff checklist from the Deliverability Playbook. In one hour: set up SPF, DKIM & DMARC correctly, check your domain against blocklists, and pass Gmail & Yahoo's 2026 sender requirements.

No spam — that would be ironic. Unsubscribe anytime.