Research · Dispatch #9 ·
A citation is not proof: three ways AI search cites you and still gets it wrong
The AI-search industry treats one question as the whole game: was I cited or not? But a valid citation has three independent properties: the source supports the claim, the citation points to the right source, and the version cited is current. Passing one says nothing about the other two. All three fail in the wild: an audit of AI-search verifiability finds about one in four citations does not support its sentence, and a Tow Center benchmark finds several engines emit fabricated or broken URLs in over half their citations; in our own probing an engine rendered our content under a domain we do not own, and another cited our live URL while reproducing figures we had already corrected. Our own two observations are rare; the point is that 'cited: yes' is not the same as citation integrity, which needs claim, destination, and version checks a binary counter never performs.
The ninth GEO Glossary dispatch. (The sixth failed its evidence gate and was retired unpublished; we number by our project ledger, so the gap is deliberate, and one of the three findings below is the thing that dispatch tried and failed to claim.)
The AI-search visibility industry runs on one number: were you cited, yes or no. Tools count it, dashboards chart it, and a citation is treated as the win. This dispatch is about what that number hides. "Cited" is treated as a single state, but a valid citation has three independent properties, and passing one says nothing about the other two. We treat citation integrity as three separate checks:
| integrity axis | the question it answers | the failure |
|---|---|---|
| support | does the source support the claim it is cited for? | unsupported or imprecise citation |
| destination | does the citation point to the right source? | fabricated, broken, or wrong address |
| version | is the cited version of the source current? | cited-version lag |
These are orthogonal: one citation can fail any of them without failing the others, and a citation that fails one still counts as "cited." We have seen all three fail, two from our own weekly frozen-panel probing and one from reading the field's papers alongside independent audits of AI search. Our own two observations are rare, and we say so below; the published audits show that support and destination failures can be far from rare in other engines and test settings, and we make no single cross-engine prevalence claim. What the three share is not a frequency but a blind spot: each is a way a citation that "happened" is still wrong, on an axis the yes/no metric does not have.
Support: the source does not back the claim
The support axis asks whether the cited source actually says the thing it is cited for. Measured directly on AI answer engines, it fails a meaningful share of the time. The best-known audit of generative-search verifiability found citation precision of about 74.5%: roughly one in four citations an engine attached did not fully support the sentence it was attached to. (The same study found citation recall of about 51.5%, a related but separate coverage problem: about half of generated sentences were not fully supported by any citation. Precision is the axis that fails here, not recall.) A reader cannot see either failure without opening the source and reading it against the claim. "Cited" tells you a link exists; it does not tell you the link holds.
The same structural failure runs through the field's human literature, which is worth naming because it is where practitioners meet it. We checked GEO's foundational numbers against the papers behind them. The single most-repeated figure, "GEO can boost visibility by up to 40%," is a real sentence from a real abstract, but the 40% is a position-adjusted word-share proxy measured on a 2023 GPT-3.5 testbed, and the same paper's real-search-engine headline quietly switches to a smaller metric. A separate multi-actor re-test found most content tactics ineffective or negative once everyone adopts them. The number survives the trip into marketing; its conditions do not. Same break, different actor: a real source attached to a claim materially broader than what the source measured.
Destination: the citation points to the wrong place
The destination axis asks whether the citation points to the intended source at all: the right site, the right article, at an address that resolves. This one is documented at scale, and it is far from rare in some engines. A Tow Center audit of eight AI search engines (Columbia Journalism Review, March 2025) gave each engine a quote and asked for the source; across 1,600 tests the engines returned the wrong source information more than 60% of the time, and for two engines more than half of the citations led to fabricated or broken URLs (154 of 200 tested from one). Fabricated links, syndicated copies cited in place of the original, and wrong-article attributions are all destination failures that a citation counter reads as success.
Our own instance is a narrower sub-case of the same axis, and it is the one we caught first and got wrong first, which is why it is worth telling carefully. On one round in July, Microsoft Copilot rendered five of its citations to our brand and our exact page paths under domains we do not own: four under a hyphenated address that does not resolve at all, and one under a parked lookalike domain. The content was ours, the brand was ours, the path was ours; the address was not. A citation like that is "cited" by every counter, and it delivers zero traffic, because the URL goes nowhere.
Here is the honest part. When we first saw it, we drafted a dispatch about it. Then we ran a held-out round to check whether it persisted, and it did not: on the re-probe, none of the five rendered the phantom domain again, and the next several rounds recorded zero phantom addresses. So we retired that dispatch unpublished rather than generalize a single round into "engines fabricate domains." It was not a rate; it was an instability. Six weeks later, on one term in a late-August round, a wrong-domain render appeared once more, which is why this failure mode earns a mention here but not a claim of frequency. The honest description is bounded: engines occasionally emit a wrong or unresolvable address for content that is genuinely yours, we have seen it twice, and it is intermittent, not systematic. We log the rendered domain of every citation precisely so this stays an observation and never becomes an inference.
Version: the cited version is stale
The version axis asks whether the version the engine reproduced is the current one, and this is the failure mode we found most recently and named: cited-version lag. An engine cites your page's live, current URL but reproduces a claim from an earlier version of that same page, one you have already corrected. The address resolves, the source is right, the content is genuinely something that page used to say. It just is not what the page says now.
We caught it on one of our own entries. We had corrected a set of figures on a page, and across two weekly rounds an engine cited that page's current URL while reproducing the pre-correction figures. We could rule out a simple misreading of the live page, because the specific numbers the engine returned appear nowhere on the current version; they only existed before the fix. The behavior is consistent with a stored, stale copy of the page, which is a documented retrieval reality: an independent investigation of one major engine's retrieval stack found a shared reading cache that serves stored page copies, some observed served more than ninety days after they were fetched, and that engine's own vendor documentation notes that in one workspace mode its indexed or cached pages can be older than the live version.
The uncomfortable property of this one is that it is nearly invisible. You can only detect it where the engine reproduces the exact fact you changed, and most of the time engines reproduce a page's framing, not its specific numbers, and the framing does not change when you fix a figure. So our evidence is a single term on a single engine across two rounds: enough to name the failure mode, not enough to say how often it happens. We report it as a bounded observation, not a rate, and the honest consequence is that the cases we cannot see are unknown, not assumed.
What this means if you measure AI citation
- Treat "cited" as the start of the check, not the end. A yes on the inclusion question leaves three things unverified: does the source support the claim, does the citation point to the right source, and is the version current. Each is a separate look, and each can fail while the citation counter reads success.
- Read the rendered URL, not just the brand. A citation that names you under a domain you do not own delivers nothing; a visibility tool that credits brand mentions will over-count it. Check that the address resolves to your site.
- Treat a correction as landed only after you verify it. After you fix a specific fact, re-query that fact and read whether the old or new value comes back. A clean-looking answer that reproduces only your framing is not evidence the correction propagated.
None of this requires distrusting AI citations wholesale; it requires not treating a citation as self-certifying. The three axes need checks a binary counter does not perform: claim-level source verification, destination validation, and version-aware re-probing. The first two you can run on a single answer, by reading the source against the claim and clicking the URL; the third needs your correction timestamps and a later re-probe, because a version mismatch is only visible over time. A dashboard that records only whether a citation appeared misses all three unless it separately validates the claim, the destination, and the source version.
Limits, honestly
The three axes do not rest on equal evidence, and we would rather say so than flatten them. The support and destination failures are documented at scale by others (a generative-search verifiability audit; the Tow Center's eight-engine news-citation benchmark), so "far from rare in some engines" is a claim those studies can carry, with the caveat that their numbers are their test settings, not a universal law. Our own two observations are the thin ones. The wrong-domain render is two intermittent sightings, and we killed an earlier dispatch for trying to make more of it. The stale-version case is one term on one engine across two rounds and is structurally hard to observe, so we make no frequency claim about it at all. What ties the three together is the shared blind spot, not a shared rate: each is a way a citation that "happened" is still wrong, on an axis the yes/no metric does not have. A name is what lets you go looking, and a field that measures only inclusion has not been looking.
More dispatches
- Dispatch #8We built this glossary on coining our own terms. Does AI search actually reward it?
- Dispatch #7We refuse to invent benchmarks. Does AI search punish us for it?
- Dispatch #5GEO's most-cited numbers, checked against the papers they come from
- Dispatch #4One panel, five engines, mostly separate citation sets: 'cited by AI' is not one thing
- Dispatch #3Cited more on Gemini, less on ChatGPT: a gradient, and what it is not
- Dispatch #2Google caught up: the AI-citation gap looks like a reporting lag
- Dispatch #1AI engines cited this page before Google indexed it