Your Brand’s AI Reputation Is Being Written by Strangers, Not You
Your Brand’s AI Reputation Is Being Written by Strangers, Not You
You’ve been obsessing over your website copy, your owned media, your press kit. Doesn’t matter. When an AI answers a question about your company, it’s pulling from sources you don’t control — 6 times more often than from anything you own. The rules of brand reputation just changed, and most founders haven’t noticed yet.
What happened
Researcher Dmitrij Zatuchin analyzed 167,551 URL-grounded citations across 128 brands in 12 markets and 13 languages, sourced from Rankfor.AI datasets, to ask a simple question: where does AI actually get its brand information? The answer is blunt. Third-party sources account for 85.7% of all citations; owned domains account for just 14.3%. The source distribution is brutally concentrated — 80% of citations come from roughly 18% of domains, following a Zipf law (α = 0.86, R² = 0.983), meaning a handful of platforms carry almost all the weight. Wikipedia is the single most-cited domain in 11 of 12 languages tested. The exception: Lithuanian, where business daily vz.lt edges it out with a marginal 4.38% share. Market-level anomalies exist — for 46 Polish national brands, YouTube tops the citation chart, and four HR and careers portals together supply 637 citations vs. 297 for Polish Wikipedia, roughly double. This is a study of source attribution in AI, not of answer quality, and that distinction matters enormously for how you interpret it.
Cold read
This is descriptive data, not a causal model — knowing which domains get cited tells you nothing proven about why those citations move AI outputs in one direction or another. The 128 brands are not a random sample; the paper doesn’t specify how they were selected or whether they skew toward large incumbents with deep Wikipedia coverage, which would inflate the Wikipedia dominance finding by design. The study is a snapshot, not a longitudinal track: we don’t know whether AI visibility patterns shift as models update, as RAG pipelines change, or as GEO tactics mature. The Zipf distribution finding is mathematically tidy but ecologically unsurprising — the web itself follows power-law distributions, so citations doing the same tells us less about AI behavior specifically than it might first appear. Finally, “citation” here means a URL retrieved and attributed; it does not establish that the cited content is faithfully represented in the AI’s actual answer.
What it means for you
- Signal maturity: 3/5 — Real data, real scale, but correlation-only; no intervention study
- Who gets hurt: Mid-market B2B brands with thin Wikipedia presence and no earned media strategy — your owned site is almost irrelevant to what AI says about you
- What breaks if this is true: The entire “content marketing = brand control” thesis collapses; your blog posts and landing pages are largely invisible to AI citation layers, making third-party review sites, HR portals, and trade press the actual moat
- Why it might not land: The finding may be model- and pipeline-specific; a retrieval-augmented generation system tuned to prefer owned or verified sources would flip these numbers, and enterprise deployments increasingly do exactly that
- Watch for: Any major AI platform (Perplexity, ChatGPT Search, Gemini) changing its default source-weighting policy — that single lever reshapes every number in this paper overnight
Forecast as of 2026-06-25
By Q2 2027, at least two enterprise brand-monitoring vendors will ship citation-source dashboards as a standard product SKU — not mention-tracking, but domain-level citation-share tracking — explicitly citing this class of research as the rationale. If that doesn’t happen, the market has decided this signal doesn’t convert to spend.
Source: How Large Language Models Source Brand Reputation Across Languages and Markets — Dmitrij Zatuchin. https://arxiv.org/abs/2606.25787v1
