Reading the numbers

Memory, or the live web?
One setting decides.

Ask the OpenAI API a question with the search tool available but not required, and most of the time it will not search. It answers from what it remembers. If you are running a brand-visibility check that way, you are not measuring today's web. You are measuring a model's old catalogue, and the two do not recommend the same companies.

We shipped this bug ourselves

This is not a finding about other people's tools. It started as a defect in ours.

On 9 July 2026 our OpenAI column was returning zero sources. Not few. Zero, on every query, while the other three engines came back with citations as normal. The cause was one line: we offered the model the web_search tool and let it decide. It decided not to. Forcing the call with tool_choice: {type: 'web_search'} turned the same query that had returned 0 citations into 16.

We fixed it that day and moved on, which was the mistake. Every OpenAI reading we had taken until then was memory, not the live surface, so the baselines had to be thrown away and taken again. And the finding went into a changelog line, where it sat for two months as an implementation note rather than what it actually was: a fork in the road that every tool in this category walks past, and almost none of them tell you which way they went.

So we measured what it was worth

On 3 September 2026 we ran the same five questions three times against the same engine, changing one thing only: whether the model was made to search. Spanish buy-intent questions about photography backdrops, three samples each, fifteen samples per mode, thirteen brands watched, all three runs inside three minutes.

The citation side is stark:

Same five questionsForcedModel decidesNo tool
Distinct domains cited3060
Citations per answer7.61.20.0
Questions with any citation5 of 51 of 50 of 5

Left to itself the model searched for one question out of five. The other four it answered from memory. A tool that does not force the call is not measuring a bit less than one that does. Four times out of five it is measuring a different thing, and its citation report comes back nearly empty by construction, not because the brand is missing.

If a tool shows you a citation count near zero and no way to see whether the engine searched, that number cannot be read. Empty because nobody cites you and empty because nobody looked are the same picture from outside.

The part that actually costs money

Most tools in this category do not sell citation counts. They sell whether the model recommends you. So the question that matters is not how many links come back, it is whether the two modes name the same companies.

They do not. Mention rate per brand, averaged over the five questions:

BrandForcedModel decidesNo tool
Amazon0.200.200.47
Westcott0.200.070.40
Adorama0.200.200.33
B&H0.270.200.33
Lastolite0.070.130.33
Manfrotto0.000.130.27
Colorama0.000.000.13
Casanova Foto0.000.000.07
Kate Backdrop0.130.070.07
Savage Universal0.070.070.00

The two modes agreed on 60% of the recommended set. Four brands in ten were not in both. Three were named only from memory: Manfrotto, Colorama and Casanova Foto, all of them absent once the model actually looked at the web. One went the other way, Savage Universal, named only when it searched.

And the rates rise almost across the board when the search goes off. Manfrotto from 0.00 to 0.27. Lastolite from 0.07 to 0.33. Westcott from 0.20 to 0.40.

The shape of that is worth naming. Memory is an old catalogue. It reaches for established manufacturer names, and it reaches for them more often. The answer written while browsing looks more like the market as it is now, retailers included. Which of the two your tool reports is not a detail of implementation. It decides whether a legacy brand looks dominant or looks absent.

What this does to a tool comparison

People keep asking why two AI-visibility tools give different numbers for the same brand on the same day. Sampling explains part of it, and we have written about that separately. This is the other part, and it is cleaner: two tools can be reading two different universes.

There is no way to tell from the outside. The setting is invisible in the output. A report that says "visibility in ChatGPT" looks identical either way, and the vendor is rarely asked which one it is, because almost nobody knows the question exists.

So it is worth asking directly, of us as much as of anyone else: does this number come from an answer that searched? If the vendor cannot say, the number has no denominator. Our actor forces the call on every OpenAI check, and has since build 0.1.13 in July.

The other three engines expose this differently and none of them is a copy of OpenAI. Perplexity is a search product, so the question barely applies. Anthropic offers the tool optionally. Gemini has its own retrieval behaviour. What follows here is about OpenAI, which is also the engine most people mean when they say "ChatGPT visibility".

Run it yourself

The switch that produced these three runs is in our actor's source, and production does not set it, so the default is the behaviour we ship: forced. It exists so the comparison can be repeated rather than believed.

If you are building your own measurement instead, the check takes one run: ask five questions with the tool on auto and count how many answers came back with any citation at all. If that number is not five, your citation report is describing the training data.

Scope, so this is not read as more than it is. One engine, one model (gpt-5-mini), one category, one language, one day, fifteen samples per mode. At that size the per-brand rates are firmer than the sets: Savage Universal appearing in one mode and not another can be sampling noise, Manfrotto going from 0.00 to 0.27 rather less so. It does not measure what any other tool does. It measures a lever that every one of them has, and almost none of them documents.

Why we publish the ones that make us look worse

We sell a paid tool that measures this. It would be easier to publish the setting we got right and skip the two months our own OpenAI numbers were wrong. But a measurement business that only shows you its clean results is selling confidence, not measurement.

The same discipline runs through the rest of what we publish. On the census we probe every domain for a path that cannot exist, because a server that answers 200 to anything also answers 200 to /llms.txt, and three quarters of the apparently agent-ready web is that. Our own usage numbers are live at metrics.json, zeros included, and most of them are still zero.

If you want to see how your own site reads to an agent, that part is directly measurable and free: the checker shows exactly what your discovery files say, with no sampling involved. What an engine then does with them is what this page is about.