Once your site is AI-ready and you've started paying attention to AI search, a new problem shows up: too many numbers, most of them impossible to check. A dashboard says "AI Visibility: 40%." A vendor's case study says brands "using our platform" saw citations jump by some large percentage. A competitor's homepage claims they're "the category leader." None of these are wrong to look at, but only some of them are measurements. The rest are marketing dressed up as data.
This post is a practical answer to a simple question: if you're going to keep a running eye on how your business shows up in AI search, what actually belongs on that scorecard, and what should you stop paying attention to? Not a definition of AI visibility itself, that's covered in detail in AI visibility metrics. This is about building a measurement habit, layer by layer, and knowing which layer produces a real number and which one tends to produce noise dressed up as insight.
Think in Layers, Not One Score
The mistake most dashboards make is collapsing everything into one number: an "AI SEO score," an "AI readiness index," a blended percentage that's supposed to tell you how you're doing in AI search overall. The problem isn't that a single number is convenient, it's that the things being blended together don't behave the same way. Some of what you'd want to track is deterministic, entirely inside your control, and checkable by hand if you wanted to. Some of it is a sample of something probabilistic that a company you don't work for decides. Mixing them into one figure hides which kind of problem you actually have, as covered in more depth in readiness vs. visibility.
A cleaner way to think about it is four separate layers, each answering a different question, each with its own real signal and its own vanity trap sitting right next to it.
| Layer | The real question | The real signal | The vanity trap |
|---|---|---|---|
| Site | Is my own site legible to AI, and is that changing? | Readiness score over time, with regression alerts | A single point-in-time score treated as a permanent grade |
| Query | Am I actually named when a relevant question is asked? | A fixed, self-run set of questions checked repeatedly | A "visibility score" you can't reproduce yourself |
| Traffic | Did any of this bring a real visitor? | GA4 referral data, Search Console's AI report | Treating "impressions" or "exposure" as if they were visits |
| Market | How does this compare to competitors or the industry? | A benchmark audit run against a named competitor, on demand | A vendor's aggregate case-study percentage with no visible method |
The rest of this post walks through each layer.
Layer 1: The Site Itself
This is the layer you fully control, and it's the one with the least excuse for a bad number, because everything on it is either objectively true or objectively false about your own pages.
The real signal is your readiness score changing over time, plus an alert when it drops. A score isn't useful as a single reading, it's useful as a trend line: did the last set of fixes actually move it, and did something later quietly undo them. That second half matters more than people expect. Once a site is connected on a paid plan, a weekly re-crawl compares the new result against the last one and sends an alert if the score drops past a threshold (five points by default, adjustable per site), naming specifically what regressed rather than just reporting a lower number. A redesign that drops a canonical tag, a plugin update that disables schema output, a theme switch that strips a script tag, all of these show up as a dated, specific alert instead of a mystery you'd only notice by accident months later.
The vanity trap here is treating one screenshot of a high score as proof of anything ongoing. A score is a snapshot of the page as it existed the moment it was crawled. A site that scored 92 in March and hasn't been re-checked since isn't necessarily still a 92, it's a site nobody has looked at in months. The number is only doing work if something is re-running the check.
Cadence: this layer mostly runs itself. A weekly automatic re-crawl catches drift on a connected site without you doing anything. Layer on a manual re-audit after any redesign, migration, or CMS change, the events most likely to break something silently, and treat a fully clean pass as a reason to move on to other work, not as a permanent state.
Layer 2: What AI Actually Says When Asked
This is the probabilistic layer: whether your business gets named, and by whom, when a real question gets asked. It's real data, but only within the boundaries of how it was collected.
The real signal is a fixed set of questions, checked the same way, repeatedly, over time. Ask the same ten or so questions a real customer might ask, put them to the same AI tools, and track two plain yes-or-no marks each time: were you named, was your site listed as a source. What makes this a measurement and not a guess is that the method is fixed and disclosed, you know exactly what was asked, of what, and how many times, so a change in the result means something changed, not that the questions themselves shifted under you.
The exact mechanics of what "mention" and "citation" mean, why they're two different events, and why a single-digit sample tells you almost nothing, are covered in full in the visibility-metrics breakdown linked above. One detail worth repeating here because it's a design choice, not just a definition: a well-built version of this check has a real "not checked yet" state, separate from "checked and not named." Collapsing those two into one silently inflates or deflates whatever percentage comes out the other end, since a question that was never actually asked shouldn't count against you the same way a question that was asked and answered without you does.
The vanity trap is any version of this number you can't reproduce. If a tool reports your "visibility" and won't say how many questions it asked, how many times, or where the questions came from, you're being handed someone else's model of an AI system, not a measurement of your business. A second, quieter trap: treating one good answer as a trend. AI answers to the same question vary run to run, which is exactly why the method has to be "many questions, run more than once," not "I asked ChatGPT this morning and it mentioned us."
Cadence: monthly is a reasonable rhythm for a self-run check, more often than that mostly adds noise rather than signal, since the underlying answers don't move that fast. Re-run it sooner after you've made a specific, deliberate change, the entity work, an update to your sameAs links, a fix to something the site layer flagged, so you're testing whether that change did anything rather than re-confirming a number you already had.
Layer 3: Real Traffic
This is the layer people skip because it's the least exciting, and it's also the only one that answers "did a human being actually visit."
The real signal is your own analytics. AI referral traffic shows up in GA4 the same way any other referral does, once you know which sources to look for, walked through in our GA4 tracking guide. Search Console adds a second, Google-specific number: its generative AI performance report counts impressions and clicks tied specifically to AI Overviews and AI Mode, separate from your regular Search performance numbers, covered in our guide to Search Console's report.
The vanity trap is reading "impressions" or "exposure" as if it meant "visits." These are related but not interchangeable ideas. An impression means your content was shown or used somewhere in an AI answer. A visit means someone clicked through afterward. A tool that reports a large "exposure" number and lets you assume it means traffic is letting a favorable-looking word do work a click-through report would have to earn honestly.
Cadence: check this alongside whatever analytics review you already run, monthly or quarterly is normal. There's no separate dashboard worth building for this, it's a filter on the analytics you already have.
Layer 4: What the Market Says About Itself
This is the layer with no consistent unit of measurement at all, and it's the one most likely to sneak a made-up number into a real conversation.
The real version of this layer is a benchmark you run yourself, against a named competitor, on demand, not a claim you read somewhere. That means running the same audit against a competitor's site that you'd run against your own, comparing the same pillars the same way, so the two numbers were produced by the same method and are actually comparable. Once a month, or right after you've made changes and want to see whether the gap moved, is a sensible rhythm for that.
The vanity trap is everything else in this category: a vendor's homepage claiming to be "the category leader" with no named methodology behind it, a case study reporting that brands "saw citations increase" by some percentage with no baseline, no sample size, and no way to reproduce the result. This industry has a specific, recurring version of this problem: a widely repeated "answers of 40 to 60 words get cited most" rule that traces to no actual provider and no actual study, already debunked after checking whether any AI company has ever published anything resembling it. None have. The rule keeps circulating anyway, because it sounds specific enough to be true, and specificity is exactly what makes an invented number persuasive.
The test that separates the two: can you reproduce this yourself, with your own data, using a disclosed method? A benchmark audit passes that test, because it's the same tool running the same check on two sites. A marketing claim almost never does.
Cadence: don't schedule this one at all. Treat a competitor's or a vendor's public claim as something to check only when you're actually evaluating it, a sales conversation, a competitor comparison you're about to publish, not as a number to track over time, since there's usually no consistent method behind it to track in the first place.
A Quick Test for Any Number You're Handed
Before a number goes into a report, a slide, or a decision, run it through three questions:
- Could I reproduce this myself? If reproducing it requires trusting someone else's black-box model of an AI system, it's a sample, not a fact. If it requires re-running a check you fully understand, it's real.
- Is this about my site specifically, or an aggregate claim about "brands like mine"? A number about your own site, checked with a method you know, is worth acting on. A number describing an industry average or a case study's blended result almost never applies cleanly to your specific situation.
- Would this number move if I did nothing? A readiness score won't move unless the site changes or someone re-crawls it. A citation check can shift on pure sampling noise between runs. A market claim can shift because a vendor updated its marketing copy. Knowing which kind of movement you're looking at keeps you from reacting to noise.
Common Mistakes
Blending a deterministic check and a probabilistic guess into one score. An "AI SEO score" that mixes real crawlability facts with a guess about citation likelihood produces a number that looks precise and explains nothing, because you can't tell which half moved when it changes.
Treating a single screenshot as a trend. One good answer from ChatGPT this morning is a data point, not a pattern. The layer 2 section above covers why the method has to be repeated questions, not a single lucky run.
Comparing your number to a competitor's self-reported number. Two different tools, two different question sets, two different formulas. The only comparison that means anything is one you ran yourself, the same way, against both sites.
Confusing "we track five AI engines" with "we get a real answer from five AI engines." Coverage claims and working coverage are two different things. Some engines return a real, searched answer with sources; others return an answer built entirely from training data with nothing to check, or nothing at all without the right access configured. A tool that's honest about which engines are actually working right now, instead of quietly treating "couldn't check" the same as "not mentioned," is giving you a real number. One that blurs that distinction is giving you a bigger-looking number.
Checking obsessively instead of on a rhythm. Daily re-checks of a probabilistic layer mostly measure the AI's own run-to-run variation, not your business. The cadence guidance above exists so you're spending attention where it produces signal.
Frequently Asked Questions
Is your site ready for AI?
Get a free readiness score in under a minute. No signup, no card.
Run the free check