98.6% of visible citations in a recent OpenAI local-business study came from third-party websites.
It can also become misleading very quickly if the measurement contract is not visible.
The study was sent to me by James Tandy, founder of Empirank, after he read our AI Search Visibility Measurement Guide. Empirank analysed 2,403 visible citations across 300 local-business answers from what it describes as a search-enabled OpenAI configuration.
Only 1.4% of those citations were attributed to the recommended businesses’ own websites.
A business cannot depend on its website alone. It also needs accurate listings, reviews, local portals, professional associations, publishers and other independent sources.
But I would not place the 98.6% figure beside another AI citation percentage until I had compared the measurement designs.
The 300 answers contained an average of about eight citations each. But averages can hide an uneven distribution.
Imagine that 30 long answers produced most of the citations while the other 270 answers displayed only one or two sources. The pooled citation percentage would then describe the citation-heavy answers more strongly than the typical answer.
Of all recorded citation instances, what percentage pointed to owned and third-party sources?
What percentage of answers contained at least one owned source, at least one third-party source, both, or neither?
An API can return HTTP 200 while silently ignoring an unsupported field. A system can also return citations without exposing enough metadata to prove whether they came from live retrieval, a prebuilt index or another grounding layer.
But unless retrieval leaves independently observable evidence, “retrieval on” may only be a label created by the experimenter.
A credible retrieval-on study therefore needs to record how search execution was verified.
Depending on the platform, that evidence might include tool-call records, retrieved-source annotations, grounding metadata, provider documentation tied to the exact endpoint, or a controlled freshness test.
“Owned website” sounds like a simple classification until real businesses enter the dataset.
Empirank states that its business-source attribution was deliberately conservative and that exact root-domain matching may miss franchise relationships, verified subdomains and other owned properties.
A conservative matcher may classify some controlled or authorized properties as third-party. A permissive matcher may make the opposite mistake and treat partner-controlled pages as brand-owned.
Neither rule is automatically wrong. The classification just needs to be declared and versioned.
