Guide 7 of 9 · Draft Preview

Are You Actually Being Cited?

How to find out whether ChatGPT, Perplexity and Google AI are really mentioning your business — because your normal analytics probably can't tell you

Guide 7 of [series name TBD] — content draft, no design applied.

The short version

Guide 4 covered what to build. This guide covers a question most advice on this topic skips entirely: how do you actually know whether any of it worked?

The honest starting point is that visibility in AI search isn't something that shows up automatically in the tools you already use. It doesn't behave like classic SEO, where a rank tracker gives you a number every day. Retrieval pipelines differ by platform, they change without notice, and — this is the part most measurement advice misses — your existing analytics setup is very likely undercounting how much AI search is already sending you. Not by a little. Potentially by a lot, and silently.

This guide covers four things: how to monitor whether you're actually being cited, how to check your basic search and index health, how to measure whether any of it converts into something commercial, and — the practical core of this guide — how to build first-party attribution that catches the traffic your analytics tool is currently losing.


The blind spot nobody mentions (formal label: referrer-based attribution failure)

Here's the mechanism, plainly: when someone clicks through to your site from a normal Google search result, their browser sends a referrer header — a piece of information saying "this visit came from google.com." Google Analytics and every other session-based analytics tool reads that header and files the visit under "Organic Search."

ChatGPT and Google's AI Mode frequently don't send that header. When someone clicks a link inside an AI-generated answer, a meaningful share of those visits arrive with no referrer information at all. Your analytics tool has nothing to file them under except "Direct" — the same bucket as someone who typed your URL into their browser from memory, or came in from a source your tool can't identify.

The practical consequence: if you're judging how much value AI search sends your business by looking at a "referral" or "AI search" channel in GA4, you are working from an undercount. Some of what you're crediited with as "Direct" traffic is actually AI-referred traffic wearing a mask. Session-based, referrer-dependent analytics simply cannot see it.

This is the strongest, most practical argument in this whole guide series for building your own first-party attribution rather than trusting a dashboard. It's also — as far as we can tell — something almost nobody writing about AI search measurement actually mentions. Most guides to "tracking AI citations" stop at prompt-based spot checks and citation-monitoring tools. Very few mention that the traffic itself is being systematically miscategorised before it ever reaches your reporting.


The fix: capture attribution at the point of conversion, not from the browser

The reliable way around this is to stop relying on the referrer header entirely for the moment that matters most — when someone actually converts (a form submission, a lead, a purchase, a signup) — and instead capture attribution data directly into that submission, at the point it happens.

This means: when a visitor lands on your site, capture what you can see about how they got there and store it. When they eventually convert — which might be on their first visit, or their fifth, days or weeks later — attach that stored attribution data to the lead record itself. You're no longer dependent on a browser passing along information it might withhold; you're recording what you can see, when you can see it, and keeping it.

Here's a concrete field set worth implementing. None of this is complicated to build — it's mostly hidden form fields and a small amount of client-side storage — but almost nobody does all of it, which is exactly why it's a differentiator.

Field What it captures Why it matters
source Your own classification of the traffic source, resolved from whatever signals are available The single field you'll actually report from day to day
landing_path The specific page the visitor first arrived on Tells you which page is doing the work — an AI-cited page looks very different from a homepage landing
referrer / referrer_host The raw referrer string, when one exists, and the domain it resolves to Your fallback signal — useful when it's present, and its absence is itself informative
utm_source, utm_medium, utm_campaign, utm_term, utm_content The standard five UTM parameters, captured if present in the landing URL Catches anything you've deliberately tagged — including any links you place in places you control
gclid / fbclid / msclkid The click IDs Google, Meta and Microsoft ads attach to paid traffic Lets you separate genuine paid performance from everything else without relying on UTM discipline alone
attr_first_* (first-touch) The visitor's original attribution data, captured on their very first visit and preserved Answers "what actually introduced this person to us" — see below, this is the field that recovers AI-referred value
attr_last_referrer (last-touch) The referrer on the visit immediately before conversion Answers "what brought them back this time" — often a different answer to first-touch, and both are useful

Why first-touch and last-touch need to be separate fields

This is the detail that makes the whole system worth building properly rather than half doing it. Consider a realistic sequence: someone asks ChatGPT a question, gets your business named in the answer, clicks through, looks around, and leaves without converting. Two weeks later, having half-remembered your name, they type it directly into Google, click your listing, and fill in a form.

A tool that only records last-touch attribution will credit that lead to "Organic Search — Brand." That's not wrong, exactly, but it's incomplete in a way that actively misleads you: it makes it look like Google search generated the lead, when the actual introduction — the thing that put you in this person's head in the first place — was an AI citation two weeks earlier. If you're deciding whether the work in Guide 4 and Guide 8 is worth the effort based on what your attribution shows, and your attribution only captures last-touch, you will systematically undervalue AI citation as a channel, on top of the referrer problem above.

Capturing first-touch separately (store it once, on the visitor's first-ever visit, and don't overwrite it) and last-touch separately (refresh it every visit) means you can answer both questions honestly: what originally brought this person into your world, and what brought them back today. For a channel like AI citation, where the gap between "cited" and "converts" is often measured in days or weeks rather than minutes, first-touch is frequently the more honest number.


Citation monitoring — are you actually being cited, or just listed

Separate from the attribution work above, you need a direct check on whether your content is being surfaced at all. Two things matter here, and they're not the same thing:

  • Being listed among sources. An AI system names your domain somewhere in a "sources" list at the bottom of an answer. Better than nothing, weak on its own.
  • Being cited for the key claim. Your business is the thing the answer actually attributes the specific fact or figure to, in the body of the answer itself. This is the outcome that actually matters — it's the difference between "we were in the list" and "we were the answer."

Build a monthly routine, per flagship asset, that checks:

  • Google's AI Overviews / AI Mode, ChatGPT, Perplexity, Copilot and Claude, using a consistent set of prompts each time (your own head query plus the sub-questions you mapped when you built the asset — see Guide 4's query fan-out step).
  • Bing Webmaster Tools' AI-performance reporting where it's available — this is currently one of the few semi-official windows into AI-referral behaviour any platform gives you.
  • Referral tagging on any links you control that point into AI-facing surfaces, so you can see clicks even where a citation itself isn't directly measurable.

Treat all of this as a working hypothesis you're testing, not a settled measurement. Nobody outside the AI platforms' own teams knows their actual selection criteria, and it changes without notice — this monitoring loop is what lets your understanding catch up with reality instead of running on last year's assumptions.


Search and index health — the unglamorous prerequisite

Before any of the above can work, the basics have to be true: your content needs to be indexed, and indexed everywhere that matters. Track, per platform:

  • Indexation — is the page actually in the index at all, on both Google and Bing (Bing matters disproportionately here, since ChatGPT's web search and Perplexity both draw on it).
  • Rankings and snippet capture — classic SEO metrics, still worth tracking, because a page that can't rank generally can't get cited either.
  • Crawl coverage — is your crawler policy actually letting the AI platforms' crawlers in, and are they visiting.
  • Independent-index engines — platforms like Perplexity that run their own crawler rather than inheriting Google or Bing's index need checking separately; good performance on the big two tells you nothing about them.

Commercial measurement — does any of this actually matter to the business

Citation volume on its own is a vanity metric until it's connected to something commercial. Once your first-party attribution (above) is in place, the questions worth asking monthly are:

  • Source-specific conversion rate. How do AI-referred visitors convert compared with organic search visitors? These are genuinely different audiences with different intent, and treating them as one number hides the difference.
  • Assisted conversions. AI-referred visitors often don't convert on the visit where they were first introduced to you — they come back later, often via a direct brand search, which is exactly the first-touch/last-touch problem covered above. Track how often that pattern shows up.
  • Lead quality, fed back into what you build. If AI-referred leads convert at a different rate, or a different value, than other channels, that's a real input into deciding what to build more of — not just a number to report.

The point of all this measurement is to make sure citation volume is never mistaken for commercial value on its own. A page can be cited constantly and generate nothing; another page might be cited rarely but convert every visitor who arrives through it.


The lifecycle decision — what you actually do with the numbers

This isn't a one-off audit; it's a loop. Once you have real numbers, each asset gets one of three decisions on a regular schedule: refresh it, consolidate it into something stronger, or retire it. An asset that isn't earning its place — not getting cited, not converting, not ranking — is a maintenance cost with no return, and the honest move is usually to fold it into something that is working rather than let it sit there contradicting more current information elsewhere on the site.

The whole loop, in order: build, prove, distribute, measure, decide. Guide 4 covers build. Guide 8 covers the distribution step in detail for research assets specifically. This guide covers prove, measure and decide.


Where to go next

This guide covers whether what you've built is working. It doesn't cover the highest-value thing you can build in the first place — original research that nobody else owns the answer to, which is both the strongest citation asset in the whole taxonomy and the one most guides get wrong. Guide 8 covers the three ways to produce it and walks through a full worked example.


Guide 7 of 9 · internal draft preview, not for search engines.