Guide 1 of 9 · Draft Preview

How AI Actually Finds You

What ChatGPT, Google AI Overviews and Perplexity actually check before they'll recommend a business over its competitors

Guide 1 of [series name TBD] — content draft, no design applied. Oliver builds the page from this.

The short version

When someone asks an AI system a question, they don't get ten blue links. They get one answer, and one or two businesses get mentioned in it. Everyone else doesn't exist for that question.

That answer isn't based on Google rankings. It's based on which business the AI system can describe clearly, trusts, and can confidently answer questions about. This guide explains what "trust" means to an AI system in practical terms, and what to do about it. Guide 2 in this series turns it into a checklist you can action today.

Worth checking before you assume you know where you stand: search a query your site ranks page one for on Google, then ask ChatGPT or Google's AI Overview the same question. It's common for the two to return completely different businesses. Ranking well in one system tells you very little about the other.

Also worth checking: your own analytics are probably lying to you about this. ChatGPT and Google's AI Mode frequently don't pass a normal browser referrer when someone clicks through from an answer, so Google Analytics dumps that visit into "Direct" traffic — indistinguishable from someone typing your URL in by hand. If you're judging how much AI search matters to your business by what GA4 shows you, you're working from an undercount, not a real number. First-party attribution captured at the point of a form submission or purchase — not a browser referrer — is the only way to see this accurately.


Two different systems

Traditional SEO is the practice of influencing a document-retrieval system built on link graphs, term frequency, and crawl accessibility. It runs a three-stage pipeline: Google crawls your pages, decides which to index, then ranks them against a query using signals like backlinks, keyword relevance, technical health, and user behaviour.

AI search works differently. Most AI answer systems use retrieval-augmented generation (RAG): they retrieve a set of candidate documents, then generate a synthesised answer citing some of them. There's no ranked list — being "cited" means your content was extracted and used as a source, or it wasn't. The systems also don't share one index: ChatGPT's web search runs on Bing's index, Google's AI Overviews run on Google's own data, Perplexity runs its own crawler. Optimise for Google alone and you have real gaps elsewhere.

Traditional SEO AI search
What you're optimising for Position in a ranked list Being cited in a synthesised answer
How relevance is judged Keyword matching + link authority Semantic understanding + content quality + entity trust
Links Core ranking signal — more quality backlinks, higher rankings Indirect — links build domain authority, which improves retrieval odds, but there's no link graph in the AI answer itself
Measurement Rankings, impressions, clicks — mature tooling Citation frequency, brand search uplift — tooling is still immature
Crawler Googlebot GPTBot, PerplexityBot, ClaudeBot, Google-Extended (Gemini training only — Google's AI Overviews still run on Googlebot)

They're not the same job, but the good news is most of the work overlaps — see below.


The five things AI checks before it trusts you

(Formal label: entity signals)

Every AI platform sources its answers differently, but they all evaluate a business against the same underlying question: can I describe this clearly, and do I trust what I'd be saying? Five things determine that.

  1. Identity clarity. One consistent description of what you do, in the same language, everywhere — your site, your directories, your socials. Inconsistent descriptions create uncertainty, and uncertain businesses don't get recommended.
  2. Subject authority. Clear specialisation in a defined area. AI systems favour specialists over generalists — authority comes from depth on a topic, not volume of content across many.
  3. Structure (formal label: meaning architecture). Schema markup, site architecture, content hierarchy. Good content sitting in a poorly structured site is harder to extract and cite, even when it's genuinely the best answer.
  4. Independent confirmation (formal label: ecosystem validation). Other sources across the web saying the same thing about you that you say about yourself. One voice — your own site — is weak evidence. Several independent, consistent voices are strong evidence.
  5. Consistency over time (formal label: signal consistency). Positioning that doesn't drift. Stability compounds; changing your story resets what you've built.

These aren't a checklist to tick once. They reinforce each other — strengthening one tends to lift the others, and they compound the way a reputation does, just faster than word-of-mouth alone.


Where the two systems overlap (do this once, it serves both)

  • Content quality. Deep, genuinely expert content performs in both. Thin or templated content gets penalised by Google's Helpful Content classifiers and ignored by AI citation systems — there's no trade-off to manage here.
  • Named, credentialed authors. Consistent brand identity, demonstrable expertise — this feeds Google's quality signals and AI systems' citation preferences at the same time. Building an author entity in Google's Knowledge Graph also strengthens how AI systems represent you.
  • Schema markup. Organisation/Person schema establishes entity identity for both the Knowledge Graph and AI retrieval. FAQ schema drives rich results in Google and gets extracted directly by AI answer systems. Treat this as baseline hygiene to get right once, not as a lever you keep pulling — schema doesn't move citation on its own.
  • Cited, verifiable sources. Content that references data and attributes claims ranks better under E-E-A-T evaluation and gets cited more by AI. Same work, both outcomes.
  • Crawlability. If Googlebot can't reach your content, neither can most AI crawlers. This is a genuine prerequisite, not a nice-to-have — and it cuts the other way too: content that only renders after JavaScript runs is a real risk, because most AI crawlers don't execute JavaScript. Static, server-rendered HTML is the safer default.
  • Topical depth. Both reward genuine depth over shallow coverage of many topics.

What actually moves the needle, ranked

Most guides to this stop at "do the things above." In practice, some of them matter far more than others. In order of impact:

  1. Getting named on other people's pages beats backlinks, which beats schema. A mention — even unlinked — is a trust signal AI systems can act on. A backlink is stronger still. Schema is table stakes, not a differentiator; treat it as hygiene you fix once, not a growth lever.
  2. Being properly indexed on Bing matters more than most businesses realise, because ChatGPT's web search and Perplexity both draw on Bing's index. A site that's never been submitted to Bing Webmaster Tools is invisible to a meaningful slice of AI search regardless of content quality.
  3. Original data and research is the strongest asset you can build. Proprietary figures, first-party datasets, genuinely new analysis — this is what earns both an editorial link (SEO) and a citation (AI search), because it's the one thing a competitor can't just rewrite.
  4. Static, crawlable pages beat JavaScript-rendered ones, because most AI crawlers don't run JS. If your key content only appears after client-side rendering, you may be invisible to these systems no matter how good the content is.
  5. Freshness matters more for AI answers than for classic SEO on time-sensitive topics — AI systems with live retrieval weight recency heavily; classic rankings are more forgiving of a stale-but-authoritative page.

The six platforms, briefly

Each platform sources differently, so a single-platform strategy leaves real gaps.

Platform How it sources
Google AI Overviews Google's own index and infrastructure
Gemini Google's index, Knowledge Graph, Google-owned platforms
ChatGPT (web search) Bing's index — building Bing visibility is the most direct route here
Copilot Bing's index, plus a closer relationship with LinkedIn content via Microsoft's ownership
Perplexity Its own real-time web crawler, plus other indices — explicitly cites its sources
Claude Indexed web content and training data — clear entity definition and consistent independent description matter most

Where to go next

This guide covers how AI decides who to trust. Guide 2 turns it into an actionable technical checklist — the specific, ordered list of things to fix on your own site starting today.


Guide 1 of 9 · internal draft preview, not for search engines.