Reading the numbers

Every AI-visibility number
undercounts. Ours too.

A common complaint from people doing this work: the tool reports a handful of citations a month, while their actual inbound from ChatGPT is several conversations a week. The tool is not broken. It is measuring the only thing it can measure, and that thing is a fraction of what happens. Here is why, and what to do about it.

The part no tool can see

When a person asks ChatGPT for a recommendation, the answer they get is shaped by their history, their memory settings, their location, and the exact words they used. A monitoring tool asks the same question from a clean session with none of that. So a tool measures the public, context-free answer, and a real user gets a private, context-loaded one.

Those are different questions with different answers, and the private one cannot be sampled from outside. This is not a gap a better crawler closes. It is structural: the measurable surface is smaller than the real one, always, so every honest number here is a floor, not a count.

One reading is a coin flip

Ask the same question five times, back to back, same wording, and the brand list changes between runs. A single reading does not tell you whether you are visible. It tells you what the model said that one time.

An example from our own data, so we can publish the numbers. A photography-backdrop shop we monitor got its first ever non-zero result on Perplexity for one Spanish buy-intent query: mentioned in 1 of 3 samples. The next weekly run it was back to zero everywhere. If we had sampled once, we would have logged "first mention" one week and "lost visibility" the next. Neither happened. 1 of 3 is noise, and a tool that reports it as a yes/no is reporting noise as signal.

This cuts both ways. The same coin flip that invents a mention out of nothing also erases a real one. A brand that is genuinely recommended one time in three will read as absent on any given single check.

So we ran the same measurement three times

People comparing these tools keep asking a question nobody answers: do any two of them return the same number for the same prompt on the same day? You do not need two tools to find out. You need one tool, run twice.

On 3 September 2026 we ran three identical measurements of one of our own sites, six minutes apart. Same brand, same five questions, same four engines, three samples per cell, 60 samples per run. Nothing changed between them except that they ran three times.

On brand mentions, the three runs agreed perfectly — 0 mentions out of 180 samples, every cell identical. That is worth stating plainly, because it is not evidence that mention counting is stable. It is evidence that a zero is stable. There is no coin to flip when the true rate sits on the floor.

On citations, four in ten of the sources changed. Across the three runs the engines cited 132 distinct domains. Only 63 of them appeared in all three. 46 appeared in exactly one run out of three. Compare any two of the three runs and four in ten of the sources are new.

And the instability is not spread evenly. It is a property of the engine:

EngineDomains citedIn all 3 runsStable
Perplexity272696%
Gemini421741%
Claude39821%
ChatGPT6469%

Perplexity returned very nearly the same source list all three times. ChatGPT returned a nearly fresh one each time: of the 64 domains it cited in total, 40 showed up in a single run and only 6 survived all three. This is not the model choosing between memory and browsing, because we force the search on every call. It searched all three times. The churn is inside the search layer.

One cell caught it in the act. Asked which backdrops are best for product photography, Gemini cited our own domain in 3 of 3 samples at 10:35, 1 of 3 at 10:38, and 0 of 3 at 10:41. A report generated at 10:35 and one generated six minutes later would have said opposite things about whether that page is being used as a source.

Then we did it again where the brand was not at zero

The run above left the important hole open. A brand sitting at 0% is stable for a boring reason, so it cannot tell you anything about mention counting. Three hours later we repeated the whole design in another category, another language, on a brand that is genuinely in the answers: Dashlane, in password managers. Not ours. We used it only to get a mention rate in the middle of the range, where a coin actually shows.

It came back in 90 of 180 samples, almost exactly half. And there the mention axis does not hold still:

Against the first category, where nothing moved at all: the instability is not spread evenly, it is concentrated in the middle of the range, which is where every brand worth measuring lives. And the citation finding repeated in a second category and a second language, with the same ordering: Perplexity most reproducible at 92%, ChatGPT least at 24%.

Between two days, barely more than between two clicks

One more comparison, which cost nothing. A fourth run exists with the identical configuration, from 17 hours earlier. So we can ask whether a day of real-world drift moves the numbers more than pressing the button twice does.

Source-list overlapJaccard
Between three runs six minutes apart60.7%
Against a run 17 hours earlier52.2%
Difference8.4 points

A whole day of news, new content and reindexing moved the source list 8.4 points more than running the same measurement twice in a row did.

That is what makes weekly tracking so hard to read. When a report says your source list changed by four in ten since last week, most of that change would also have shown up between two runs six minutes apart. The week-over-week signal is real, but it sits under a noise floor that nearly matches it, and no tool that samples once can separate them.

We are not the first to find this, and it matters that we say so. Schulte, Bleeker and Kaufmann reported the same result in April 2026 on a much larger sample: four campaigns, 45 days, 3,409 pairwise comparisons across ChatGPT, Gemini, Google AI Mode and Perplexity, with same-day source overlap of 0.32 to 0.43 against day-to-day overlap of 0.34 to 0.42 — the same range. What our runs add is the split by engine, which their tables pool by campaign, plus Claude, which their four engines do not include. Don’t Measure Once, arXiv:2604.07585.

What this means for tool comparisons: before you can say two tools disagree, you have to know how far one tool disagrees with itself. That floor depends on which engine you ask, which axis you read, and how many samples you take. On Perplexity it is small. On ChatGPT it is large enough that two tools can report very different source lists with neither of them being wrong.

Scope, so this is not read as more than it is: two categories, two languages, three runs each, and the day-to-day comparison rests on a single pair of dates in a single category. Cell rates come from 9 samples, which is too few for fine movements. None of it measures other people's tools. It measures the surface all of them read.

"AI visibility" is already the wrong unit

Each engine builds its answer from its own set of sources. Being cited by Perplexity tells you almost nothing about Gemini. In one of our own comparisons, on the same question, Perplexity and Gemini shared 2 source domains out of 21. A single "AI visibility score" averages across engines that barely agree on reality, and an average of disagreeing measurements is not a measurement.

There is a second split inside each engine, and it is bigger than the first. An answer written from memory cites nothing live; an answer written while browsing cites the pages it just read. Those are two different behaviours, and one number that mixes them hides which one moved.

We measured that split on the same day and the same five questions, changing only whether the model was made to search. Left to itself the OpenAI API searched for one question in five and answered the rest from memory, and the two modes agreed on only 60% of the recommended brand set. It has its own page: memory, or the live web.

What honest measurement looks like

None of this means the numbers are useless. It means they need to be read as what they are. If you are measuring this yourself, this is the shape that survives contact with the noise:

Including ours

We run a paid tool that does exactly this, sampled and rate-based, and it undercounts for every reason above. We would rather say so than pretend a floor is a count.

It is the same discipline behind the rest of what we publish. On the census we probe every domain for a path that cannot exist, because most of what looks like an agent-ready web is a server answering 200 to anything; three quarters of the domains that appear to publish the machine-readable files are that. And our own usage numbers are published live, zeros included, at metrics.json. Right now most of them are zero. A number you can trust is one whose author shows you where it is weak.

If you want to see how your own site reads to an agent, that part is directly measurable, and free: the checker shows exactly what the discovery files say, with no sampling and no coin flip. What an engine then does with them is the part this page is about.