EchoRanked
8 min readEchoRanked Team

How to track AI Overviews (and why one check tells you nothing)

ai-overviewsai-visibilitymethodologygoogle

Short answer: Tracking AI Overviews means recording, for a fixed set of queries on a repeating schedule, four things: whether an Overview appeared at all, whether your brand was named in it, whether you were cited with a link, and which sources were cited instead. The part that separates useful tracking from theatre is repetition — AI Overviews do not fire consistently for the same query, so a single check confounds "we lost visibility" with "no Overview appeared this time", and those need different responses. Below is what to record, how to read it, and the three mistakes that make most AI Overview tracking unfalsifiable.

What makes this different from rank tracking

A rank tracker asks a stable question: at what position does this URL appear? The answer is an integer, it is comparable day to day, and a change of two positions is a real change.

AI Overviews break all three properties.

They do not always appear. The same query can return an Overview one day and not the next. Google has adjusted how broadly they trigger more than once, and it varies by query type, and it varies by whether the SERP is judged to need one.

There is no position inside them. You are named or you are not. Several brands can be named in one Overview. There is no second place.

The cited sources shift. Two Overviews for the same query, generated a day apart, can lean on different pages entirely.

So the unit of measurement is not a position, it is a rate: over N checks of a query, in how many were you named? That reframing is the whole method. Everything below follows from it.

What to record on every check

Four fields, and the third and fourth are the ones people skip:

FieldWhy it matters
Overview present?Without this you cannot distinguish "we lost the Overview" from "there was no Overview". These are entirely different problems.
Brand named?Being mentioned in the prose, with or without a link.
Cited with a link?A link is worth more than a mention, and the two move independently.
Which sources were citedThe only field that tells you why you are absent, and who took the slot.

Track the first field separately from the rest or your denominator is wrong. An appearance rate of "2 out of 10 checks" means something very different when eight of those ten checks returned no Overview at all — in that case you were named in 2 of the 2 Overviews that existed, which is a completely different fact.

This is the same reason our own measurement excludes failed and no-answer calls from the share denominator rather than counting them as absences. A missing surface is not the same as a lost citation, and averaging them together produces a number that is wrong in a direction you cannot predict.

How many checks before a change means anything

More than you would like. Here is the arithmetic, using the interval that is actually appropriate for "named in k of N checks" — a Wilson score interval:

ObservedPoint estimate95% interval
3 of 560%23% – 88%
6 of 1060%31% – 83%
30 of 5060%46% – 72%

Same rate, three sample sizes. At N=5 your true appearance rate could plausibly be anywhere from a quarter to nearly nine tenths. Reporting "60%" from five checks is not wrong so much as it is missing the only part of the sentence a decision needs.

The practical rule: do not act on a change until the intervals stop overlapping. Say you checked 30 times last month and were named 12 times — 40%, interval 25–58%. This month it is 15 of 30 — 50%, interval 33–67%. The midpoint moved ten points and the intervals still overlap across a 25-point span. You have not measured an improvement; you have measured noise with a nicer midpoint. We went through this in depth in why your AI visibility score is probably noise, and it is the single most common way AI visibility reporting misleads the person reading it.

Use the Wilson interval rather than the normal approximation, incidentally. At small N and at extreme proportions — 0 of 10, 10 of 10 — the normal approximation degenerates: it can produce bounds outside 0–100% and it collapses to zero width at the extremes, claiming perfect certainty at exactly the point where you have the least.

The three traps

Trap 1: tracking queries nobody asks. An appearance rate computed over a query set you invented is a precise measurement of nothing. Start from evidence — your Search Console queries are the best available proxy for what buyers ask, because the questions people typed into Google are the questions they are now putting to an AI. Then add the shapes that generative answers specifically attract: "best X for Y", "X vs Y", "alternatives to X".

Trap 2: checking from one location and calling it universal. AI Overviews are personalized and localized. A check from one IP in one country is one sample from one context, and generalizing it to "our visibility" imports a bias you have not measured. If your market spans countries, sample per country and report per country.

Trap 3: mistaking a screenshot for data. The screenshot in the Slack message showing a competitor named instead of you is a single draw from a distribution. It might be representative. It might be the one run in twelve where that happened. There is no way to tell from the screenshot, and organisations reorganise quarters around these.

Doing it by hand, if you want to start today

You do not need a tool to begin, and starting manually is a good way to learn what the data looks like before you pay for anything:

  1. Pick 10 queries from Search Console, weighted toward commercial intent and question shapes.
  2. Check each one three times a week, in an incognito window, ideally from a consistent location.
  3. Record the four fields above in a spreadsheet — one row per check, not one row per query. The per-check grain is what lets you compute a rate later; if you only store the latest state you have thrown away the measurement.
  4. After four weeks, compute for each query: Overviews seen ÷ checks, and named ÷ Overviews seen. Now you have a baseline with enough behind it to compare against.

The honest cost of this is about twenty minutes a week and it will produce a more defensible picture than most vendor dashboards, because you will know exactly how it was built.

Where it breaks down is scale and consistency. Ten queries is a thin sample of a category, three checks a week is a wide interval for months, and manual checking degrades the moment someone is on holiday. That is the point at which automation stops being a luxury.

What we do

EchoRanked runs this loop for you across five engines rather than one — Google AI Overviews alongside ChatGPT, Perplexity, Gemini and Copilot — because a brand strong in one and absent in three has a problem no single-surface tracker will show them.

Three things we do differently that are worth stealing whether or not you use us. Every share carries its Wilson interval on the screen next to the number, not in a footnote. Every answer is tagged with the measurement surface it came from — the consumer product a buyer actually uses, versus a provider API, which can run different retrieval and is different evidence. And every number is auditable back to the verbatim answer that produced it, so a figure you doubt can be checked rather than argued about.

We also publish our own AI visibility weekly, losses included. A vendor asking you to trust a measurement should be willing to show you theirs.

Reading the results

Once you have a few weeks of per-check rows, the questions worth asking are:

  • Which queries return an Overview at all? Categories vary enormously. If Overviews rarely fire for your queries, this is a smaller problem than your board thinks.
  • Where are you named but not linked? That gap is usually a structural fix — the engine knows about you but is citing someone else's page about you.
  • Who is cited instead, and what page? If it is a competitor's comparison page, the response is to write yours. If it is a third-party listicle, the response is outreach. These are different projects and the citation data is what tells you which one you are in.
  • Did the intervals move, or just the midpoints? Only the first is a result.

Frequently asked questions

Can you track AI Overviews in Google Search Console?

Not separately. Search Console does not break out AI Overview impressions and clicks as their own dimension, so an Overview that names you is folded into your ordinary numbers rather than reported as an event. That gap is the reason dedicated tracking exists at all.

Why does an AI Overview appear for a query one day and not the next?

Because triggering is not deterministic. Google has changed how broadly Overviews fire more than once, and it varies with query type and how the SERP is judged. This is exactly why "Overview present?" has to be recorded per check rather than assumed — otherwise a triggering change reads as a visibility collapse.

How often should I check?

Frequently enough that your confidence interval is tighter than the change you want to detect. Daily is comfortable for a small query set. Three times a week is workable. Monthly is close to useless: the interval will still be wide enough to swallow any realistic month-over-month movement.

Does being cited in an AI Overview send traffic?

Some, and less than the equivalent organic position — the answer is displayed rather than clicked. That is a real trade rather than a reason to ignore it: on a query where the answer is shown, being the cited source is the only visibility available.

Where to go next

Before you track anything, make sure the crawlers can read you — how to check if ChatGPT can read your website covers the four files that decide it. For the levers that change whether an engine names you, SEO for ChatGPT separates what is established from what is being sold as fact. And where ChatGPT gets its information explains why a cited answer and an uncited one call for completely different responses.

Keep reading

  • Why your AI visibility score is probably noise

    Most AI visibility tools hand you one number. But LLM answers vary run to run, so a single score is a sample, not a measurement. Here's why bare numbers mislead, what run-to-run variance actually looks like, and how a confidence interval fixes it.

  • SEO for ChatGPT: what actually changes whether it recommends you

    An honest account of the levers that move whether ChatGPT names your product — separating what is established, what is inferred, and what is being sold to you as fact.

  • Where does ChatGPT get its information?

    ChatGPT draws on three separate pipelines — training data, a live retrieval index, and on-demand page fetches. They behave differently, and only two of them are things you can influence.