Research · Dispatch #7 ·
We refuse to invent benchmarks. Does AI search punish us for it?
Several of our glossary entries decline to give the number people search for: no target citation match rate, no standard attribution rate, no proven lift from authoritative tone. We tested whether that honesty costs us citations by asking five AI engines the number-seeking version of each question and comparing against the definitional control. In the expanded round, 16 of 25 pairs were eligible because the definitional prompt cited us; the number-seeking variant kept the citation in 11, and no loss went cleanly to a page that asserts a number. On the engines that carry editorial framing, the answers adopted our hedges and debunked the inflated figures. The honest reading is conditional: honesty held on the engines that transmit framing, where we hold the term.
The seventh GEO Glossary dispatch. (The sixth, a Copilot domain-misattribution finding, failed its held-out persistence gate and was retired unpublished. We number dispatches by our project ledger, not by what survives review; a numbering gap is what an evidence gate looks like from outside.)
This one is about a policy of ours that looked like a liability. Several of this glossary's metric entries refuse to give the number people actually search for. Citation match rate says the metric has no universally accepted target and that any threshold is a practitioner heuristic. Attribution rate says no vendor or academic literature defines a canonical formula, and reports per-engine ranges instead of a goal. Authoritative statement strength reports the honest, deflating number: in the original GEO benchmark (Aggarwal et al., arXiv:2311.09735), an authoritative-tone rewrite produced no statistically significant improvement, and the paper says so verbatim. Meanwhile, plenty of pages in this space will happily assert one. These are not a straw man: the competitor pages our number-seeking probes actually surfaced this round asserted exactly this kind of target, from per-response citation-rate figures to a 25 percent tone lift to stage-by-stage benchmark tables.
So the worry writes itself. When a user asks an AI engine "what is a good citation match rate to aim for," the engine wants to hand back a number. We decline to assert one. Does the citation go to whoever asserts one?
We had a concrete reason to take the worry seriously. On a sibling site we also run (an exam-prep tool in a crowded education niche), the page that honestly says "there is no official score threshold" lost exactly this way: it had been cited by three engines on the definitional question, then dropped to zero on the number-seeking question, which went to competitors who assert a pass mark. Honesty, on that site, was intent-conditional: it helped on definitional queries and lost the give-me-a-number ones.
The experiment
We ran the direct test on this glossary, twice. The design is a paired probe. For each hedged term, we ask the frozen definitional question our weekly panel always asks ("What is citation match rate in AI search?") as the control, and a number-seeking variant ("What is a good citation match rate to aim for in AI search?") as the treatment, in the same round, on the same five engines (ChatGPT, Perplexity, Claude, Copilot, Gemini), with the same source-request suffix. The interesting cell is the one the worry predicts: the definitional prompt cites us, the number-seeking prompt does not, and the winner is a page that asserts a number.
We ran it twice. Round one (late July) was a pilot on two terms. Round two (August 10) was the expanded test: five hedged metric terms (citation match rate, authoritative statement strength, attribution rate, citation share, citation velocity), each verified beforehand to genuinely decline a target number on the page. Five terms times five engines is 25 pairs. A pair can only test the hypothesis if the definitional control cites us in the first place (otherwise there is no citation to lose), and in the expanded round 16 of the 25 were eligible on that test. The numbers below are the expanded round unless noted; the pilot and expanded round combined gave 20 eligible pairs.
What happened
Of the 16 eligible pairs, the number-seeking variant kept citing us in 11, and in none of the five losses did the citation go cleanly to a page that asserts a number. Every eligible pair (competitor pages kept anonymous by category, as we name only engines):
| Term | Engine | Definitional | Number-seeking | Outcome | Number-asserter in the answer? |
|---|---|---|---|---|---|
| citation match rate | Claude | cited | cited | hold | yes (vendor figures co-cited, framed by our caveat) |
| attribution rate | Claude | cited | cited | hold | yes |
| citation share | Claude | cited | cited | hold | yes |
| citation velocity | Claude | cited | cited | hold | yes |
| authoritative tone | Claude | cited | cited | hold | not logged |
| authoritative tone | Perplexity | cited | cited | hold | yes |
| attribution rate | Perplexity | cited | cited | hold | yes |
| citation match rate | Perplexity | cited | cited | hold | not logged |
| citation match rate | ChatGPT | cited | cited | hold | not logged |
| authoritative tone | ChatGPT | cited | cited | hold | not logged |
| citation velocity | ChatGPT | cited | cited | hold | not logged |
| attribution rate | ChatGPT | cited | not cited | loss | no; went to a measurement-research page |
| citation match rate | Copilot | cited | not cited | loss | no; went to a how-to blog |
| attribution rate | Copilot | cited | not cited | loss | no; went to a how-to blog |
| citation match rate | Gemini | cited | not cited | loss | no; went to a GEO-metrics guide |
| authoritative tone | Gemini | cited | not cited | loss | yes; went to a page asserting a 25% lift |
(Two non-eligible pairs are worth naming: on Perplexity, citation share was not cited on the definitional prompt but was on the number-seeking one, a bonus in the wrong direction for the hypothesis; and Perplexity's citation velocity variant went elsewhere, but its definitional control had not cited us either, so it never entered the test.)
The holds concentrate where our authority does. Claude held all five of its pairs, though it also changed underlying models (to Sonnet 5) between the pilot and the expanded round, which flatters its round-two breadth, so read its five-for-five as the softest of the strong results. Perplexity held all three of its pairs. ChatGPT held three of four; its one loss (attribution rate) went to an independent measurement-research page, not to anyone asserting a benchmark.
We owe the holds the same scrutiny we gave the losses. A hold only refutes the worry if a number-asserting competitor was there to lose to; a hold with no competitor present is hollow, honesty "winning" an uncontested query, which is the same monopoly effect we have flagged before (Dispatch #3) now hiding on the winning side. So we checked competitor presence on every hold. Six of the eleven had a number-asserting vendor demonstrably in the same answer (the four Claude cells above plus two Perplexity ones), and the engine cited us anyway, often using our hedge to frame or debunk the vendor's figure. Those six carry the weight. For the other five we did not record competitor presence at capture time, so we cannot claim them as clean refutations either way; that is a recording gap this round exposed, and we are adding a competitor-present field to the probe schema so future rounds settle it. The honest count is six holds cleanly against an asserter, five holds unproven, five losses of which four went to non-asserting pages and one to an asserter sitting on a volatile control.
The five losses live on Gemini and Copilot, and they dissolve on inspection. Gemini's two look like the hypothesis until you check its definitional controls: one of the two, probed three times in the same session, went cited, not-cited, not-cited. An engine that flips two out of three times on the identical frozen prompt has an inclusion-volatility floor higher than any honesty effect we could measure through it. Copilot's two losses have a different texture: its definitional answers for both terms were built almost entirely on our pages (one was literally a single-reference answer, ours), and the variant then rotated to a how-to blog. That is retrieval churn on a references-list engine, not a penalty for hedging. The single cell all round where a number-asserting page did win a variant we lost, Gemini on authoritative tone against a page claiming a 25 percent lift, sits on that flipping Gemini control, so we do not count it as the hypothesis firing.
What the winning answers did with our content is the part we did not predict. On the engines that held, the number-seeking answers did not treat the hedge as a gap to route around. They adopted it as the frame. Perplexity's answer to "what is a good citation match rate" opens by stating there is no universal good rate yet, reproduces our warning that the denominator must be fixed before comparing tools, offers its own 80 and 90 percent working bands, and then explicitly labels them operational targets rather than industry standards. Its answer on authoritative tone goes further and debunks the going figures: "Do not plan on a fixed lift such as 25%, 30%, or 40%," on the grounds that those numbers come from marketing articles and are not supported by the primary benchmark, citing the original paper and our entry. ChatGPT's version opens with "Very little, based on the best available evidence" and prints the honest results table. Claude's answers repeatedly quote a sentence we added to the citation-match-rate entry one week earlier, cautioning against chasing one universal target. The number-seeking queries, on these engines, produced sourced debunks of the very numbers we refused to invent.
The reconciliation with the site that lost
Two sites, the same honest-hedge editorial policy, opposite outcomes on number-seeking queries. We did not measure authority, so treat what follows as our working explanation fitted after the fact, not a variable we controlled: the contrast is consistent with originating-source status, meaning who has standing to define the concept, rather than domain authority in the SEO sense. On this glossary, the tested terms are practitioner-coined measurement vocabulary where our entries are the closest thing the query has to an originating source; an engine that wants to answer carefully has to reckon with our framing. On the exam-prep site, the honest page is one voice in a crowded niche with a genuine canonical owner (the education department) and many competitors willing to assert a pass mark; there, the honest page is optional, and on the number-seeking query it lost.
So the working rule we take from this is conditional, not triumphant: an honest hedge survives number-seeking queries where you are the authority on the term, and is at risk where you are one contested voice among many. Honesty is not an independent citation moat. It is a benefit you can bank once the territory is yours, and a real risk where it is not.
Two refinements ride along. First, the effect is engine-architectural. The engines that held (Claude, Perplexity, ChatGPT) are the ones that transmit editorial framing into their answers; the engines where pairs failed (Copilot, Gemini) select references in ways that look more like retrieval rotation, and their definitional citations were unstable to begin with. Second, a number-seeking query can actively help an evidence-dense honest page: on authoritative statement strength, the plain definitional prompt lost to generic official documentation on two engines in round one, while the number-seeking variant surfaced our page, because ours is the one carrying the actual measured number, deflating as it is.
What this means if you run a content site
If your pages hedge honestly on questions where users want a number, the practical questions are where you can afford it and what to pair it with. On terms where you are the primary or originating source, our evidence says hold the line: the engines that carry framing will use your caveats as the skeleton of their answer, and the inflated numbers get debunked with your page as the source. On contested terms with a stronger canonical owner, an honest hedge alone is exposed; whatever real evidence you do hold (a measured range, a dated study, a named condition) is the part worth leading with, because specific empirical content is what pulled our pages into number-seeking answers. And read results per engine: a blended visibility score would have averaged Claude's five holds against Gemini's coin-flips and told you nothing.
This dispatch is also the other half of an earlier one. Dispatch #5 audited how the field's foundational numbers get quoted with their conditions stripped. The supply side of that problem is pages asserting figures the underlying research does not support; this dispatch is what the demand side looks like at the answer layer, and, at least where we hold the term, the engines are on the side of the conditions.
Limits, honestly
Twenty eligible pairs across the two rounds (sixteen of them in the wider second round) is a small sample, and the variant prompts are one wording each, not themselves frozen-tested. Of the eleven holds, only six had a number-asserting competitor we could confirm was present, so the clean evidence is thinner than eleven-of-sixteen makes it sound (we now log competitor presence to close that gap). The engine-architecture split (framing-transmitting versus references-list) is our post-hoc reading of five losses, not a pre-registered hypothesis, as is the originating-source explanation. The two-site contrast is suggestive, not causal: the sites differ in niche, competition, and age, not only in who originates the term, and we run both, so this is a replication across our own properties, not an independent one. And a citation held this week is not a citation held forever; the same instrument that produced this result exists because these things churn. We will keep running the pairs.
More dispatches
- Dispatch #9A citation is not proof: three ways AI search cites you and still gets it wrong
- Dispatch #8We built this glossary on coining our own terms. Does AI search actually reward it?
- Dispatch #5GEO's most-cited numbers, checked against the papers they come from
- Dispatch #4One panel, five engines, mostly separate citation sets: 'cited by AI' is not one thing
- Dispatch #3Cited more on Gemini, less on ChatGPT: a gradient, and what it is not
- Dispatch #2Google caught up: the AI-citation gap looks like a reporting lag
- Dispatch #1AI engines cited this page before Google indexed it