A common complaint from people doing this work: the tool reports a handful of citations a month, while their actual inbound from ChatGPT is several conversations a week. The tool is not broken. It is measuring the only thing it can measure, and that thing is a fraction of what happens. Here is why, and what to do about it.
When a person asks ChatGPT for a recommendation, the answer they get is shaped by their history, their memory settings, their location, and the exact words they used. A monitoring tool asks the same question from a clean session with none of that. So a tool measures the public, context-free answer, and a real user gets a private, context-loaded one.
Those are different questions with different answers, and the private one cannot be sampled from outside. This is not a gap a better crawler closes. It is structural: the measurable surface is smaller than the real one, always, so every honest number here is a floor, not a count.
Ask the same question five times, back to back, same wording, and the brand list changes between runs. A single reading does not tell you whether you are visible. It tells you what the model said that one time.
An example from our own data, so we can publish the numbers. A photography-backdrop shop we monitor got its first ever non-zero result on Perplexity for one Spanish buy-intent query: mentioned in 1 of 3 samples. The next weekly run it was back to zero everywhere. If we had sampled once, we would have logged "first mention" one week and "lost visibility" the next. Neither happened. 1 of 3 is noise, and a tool that reports it as a yes/no is reporting noise as signal.
Each engine builds its answer from its own set of sources. Being cited by Perplexity tells you almost nothing about Gemini. In one of our own comparisons, on the same question, Perplexity and Gemini shared 2 source domains out of 21. A single "AI visibility score" averages across engines that barely agree on reality, and an average of disagreeing measurements is not a measurement.
There is a second split inside each engine. An answer written from memory cites nothing live; an answer written while browsing cites the pages it just read. Those are two different behaviours, and lumping them into one number hides which one moved.
None of this means the numbers are useless. It means they need to be read as what they are. If you are measuring this yourself, this is the shape that survives contact with the noise:
We run a paid tool that does exactly this, sampled and rate-based, and it undercounts for every reason above. We would rather say so than pretend a floor is a count.
It is the same discipline behind the rest of what we publish. On the census we probe every domain for a path that cannot exist, because most of what looks like an agent-ready web is a server answering 200 to anything; three quarters of the domains that appear to publish the machine-readable files are that. And our own usage numbers are published live, zeros included, at metrics.json. Right now most of them are zero. A number you can trust is one whose author shows you where it is weak.