Highlights
- The dataset: We analyzed over 220,000 AI source citations generated from prompts specific to the AI agents industry. The AI responses behind them come from four AI models and were run in seven languages over the course of Q2 2026.
- sanalabs.com climbed from #28 to #4 in the window (+99%): A vendor going from the periphery to a top-five source in eight weeks.
- 21% of Grok's citations are from user-generated domains: Reddit and forums, more than double Google AI Overview's 14%.
- 18% local sourcing for Grok vs ~45% for ChatGPT and Microsoft Copilot: Grok defaults to global English-origin domains even on non-English prompts.
- 61%–70% of citations are from commercial domains: Every model overwhelmingly cites vendor and business sites, not reference material.
- 90%+ of top sources differ between any two models: Any two of the four share only about 9% of their top-20 cited domains, a near-total citation split.
Which sources should you target to get cited as an AI agents brand?
Each model's shortlist has a distinctive character:
| Model | Top sources | Citation character |
|---|---|---|
| Grok | Reddit, youtube.com, linkedin.com, beam.ai, clutch.co | Social + vendor + agency directories |
| Google AI Overview | youtube.com, salesforce.com, lindy.ai, toptenaiagents.co.uk | Video + brand-direct + comparison lists |
| Microsoft Copilot | respond.io, clickup.com, sanalabs.com, superchat.de, fin.ai | Vendor sites + regional SaaS |
| ChatGPT | Reddit, arxiv.org, reuters.com, techradar.com | Forums + research + editorial press |
ChatGPT is the only model reaching for research (arxiv.org) and wire-service journalism (reuters.com) among its most-cited sources, a more editorial, evidence-seeking posture than the vendor-and-directory mix the other three favor.
How have AI source rankings changed over time in the AI agents industry?
We split the observation window at its midpoint (early May 2026) and compared each top domain's rank and citation volume in each half. Among the 30 most-cited domains, the movement was dramatic, this is a volatile, still-forming category, not a settled hierarchy.
The clearest signal is sanalabs.com, which climbed from #28 to #4 and nearly doubled its citation volume, a vendor going from the periphery to a top-five source in eight weeks. At the other end, several first-half staples (ensun.io, medium.com, clutch.co, datacamp.com) slid out of the top tier. Track it in the live AI visibility rankings.
What type of content do AI models cite for AI agents?
The models differ not just on which domains but on what type of content they trust. We classified every cited domain into six categories. Across all four, commercial domains dominate, but the second tier diverges sharply.
Grok is the user-generated-content engine: 21.0% of its citations go to sources like Reddit and forums, more than double Microsoft Copilot's and ChatGPT's 9–10%. ChatGPT and Microsoft Copilot are the editorial engines, sending 13.5%–15.7% of citations to news and trade press and cite reference works far more often than Grok and Google AI Overview. ChatGPT also cites institutional sources (3.1%) at nearly triple the rate of the others. For a vendor, this is a content-format map: earning Grok visibility means being talked about on Reddit; earning ChatGPT visibility means being covered by trade press and cited in research.
Do AI models cite local-language content for AI agents?
For the six non-English markets in scope (Spanish, French, German, Dutch, Swedish, Italian), we measured how often a cited source sits on a local country-code domain, a proxy for genuinely local-language sourcing. Across all non-English prompts, 28.3% of citations land on a local ccTLD; the remaining ~72% go to global domains, which in this category skew heavily English.
Smaller-language markets (Swedish, Dutch, German) cite local sources more than the larger Spanish and Italian markets, consistent with big global vendors publishing Spanish and Italian content directly on their .com properties. This is a ccTLD proxy, not language detection: a German-language page on a .com counts as non-local, so these figures are a floor.
Which AI model relies most on local sources for AI agents?
The localization gap between models is far larger than the gap between markets. Averaged across all non-English prompts, ChatGPT (45.7%) and Microsoft Copilot (44.6%) send nearly half their citations to local domains; Grok sends fewer than one in five (17.8%). The per-market breakdown shows the same ordering holding almost everywhere.
Microsoft Copilot relies most on local sources in five of six markets, reaching 72% local sourcing for Dutch prompts. Grok is the weakest everywhere, never exceeding 24%, and answers a French AI-agent prompt with largely the same global sources it uses for an English one. (Google AI Overview produced no French citations in scope, so that cell is n/a rather than zero.) A regional vendor that ranks well on its local .de or .nl footprint has a real shot at Microsoft Copilot and ChatGPT visibility, but will struggle to register with Grok regardless of language.
Should AI agents optimize for each AI model separately?
Yes, almost entirely. We built each model's top-20 most-cited domains and measured how many domains every pair shares. The average pairwise overlap is 9.4%, with no pair sharing even a quarter of its list.
ChatGPT and Microsoft Copilot are the most isolated pair, they agree on a single domain in their top-20. Grok and Google AI Overview overlap most, sharing recognizable anchors like Reddit, youtube.com, salesforce.com, and beam.ai. But even the friendliest pair of models disagrees on two-thirds of its sources.
How many sources does each AI model cite per answer for AI agents?
The models disagree sharply on how many sources to cite per answer. Grok is the outlier by a wide margin.
Grok cites roughly 3.8x as many sources per answer as Microsoft Copilot, the most economical model. That breadth is why Grok's citation share tilts so heavily toward Reddit and directories, a wider net pulls in more social and long-tail sources.
Context
This analysis draws from Temso's AI visibility monitoring platform, which tracks how brands appear in AI model responses across Grok (x.ai), Google AI Overview, Microsoft Copilot, and ChatGPT. The dataset covers AI-agent and AI-agent-platform prompts over roughly sixteen weeks (March–June 2026): 20,541 model responses and over 220,000 cited source links across seven languages (English, Spanish, Dutch, German, Italian, Swedish, French).
It measures what AI models cite, not whether a citation reflects a positive or negative recommendation. Citation frequency is a visibility signal, not an endorsement, and all findings are observational, they describe correlations in what the models cited, not causes.
Methodology
How we measured this
We collected every source link the four models cited when answering AI-agent prompts over the window, keeping only links the model actually cited (not merely retrieved), over 220,000 cited sources across 20,541 responses. We ranked each model's most-cited domains by citation count and built each model's top-20 list to measure overlap using Jaccard similarity on each pair. Domain categories (commercial, user-generated, editorial, reference, institutional) were assigned from a global domain registry. Cross-model overlap averaged 9.4% (range 5% to 35%); the most-vs-least-overlapping pair difference was validated with a two-proportion z-test using Wilson intervals.
Localization is measured by a country-code-domain proxy, a local-language page hosted on a .com is counted as non-local, so those figures are a conservative floor and exclude English-language prompts, which sit overwhelmingly on generic domains. The temporal analysis split the period at its midpoint and compared domain rankings in each half. Model coverage varied across the observation window, so temporal movements are reported as relative ranks within each half. Google AI Overview produced no French-language citations in scope. Every model exceeded recommended sample thresholds, with per-model cited-citation counts from 29,828 (ChatGPT) to 108,720 (Grok).
Frequently asked questions
Do AI models cite the same sources for AI-agent questions?
No, barely. Any two of the four models share on average about 9% of their top-20 cited domains, so more than 90% of the sources one model relies on are absent from another's shortlist. This is wider divergence than in more established industries, because the AI agents category is too new to have a settled set of authoritative sources.
Which AI model cites the most sources per answer for AI agents?
Grok, by a wide margin, 28.1 sources per response on average, versus 11.2 for Google AI Overview, 7.9 for ChatGPT, and 7.4 for Microsoft Copilot. That is roughly 3.8x as many sources per answer as the most economical model.
What type of content do these models cite for AI agents?
Overwhelmingly commercial. Vendor and business sites make up 61.1%–70.0% of every model's citations. The differences are in the supporting cast: Grok leans on user-generated content like Reddit (21.0%), while ChatGPT and Microsoft Copilot lean editorial (13.5%–15.7% news and trade press) and cite reference and research sources more often.
Does the language of the question change which sources get cited?
Yes, but it depends heavily on the model. Across non-English prompts, about 28% of citations go to local country-code domains. ChatGPT (45.7%) and Microsoft Copilot (44.6%) cite local sources strongly; Grok (17.8%) mostly stays on global English-origin sources regardless of the prompt language.
Which AI model is best for a regional, non-English AI-agent vendor?
Microsoft Copilot and ChatGPT. Microsoft Copilot reaches 72% local sourcing for Dutch prompts and 62% for German and Swedish; ChatGPT stays at 47%–56% local across most European markets. Grok rarely exceeds 24% local sourcing, so a regionally-focused vendor will struggle to register with it.
Did the source rankings change much over the period?
Substantially. Over sixteen weeks, sanalabs.com climbed from #28 to #4 among top domains while several first-half staples (ensun.io, medium.com, clutch.co) fell out of the top tier. The category's source hierarchy is still forming, not settled.
How should an AI-agent company use this?
Treat each model as a separate channel with its own content map. Reddit and YouTube presence drives Grok visibility; trade-press and research coverage drives ChatGPT; comparison sites and YouTube feed Google AI Overview; and a strong local-domain footprint matters most for Microsoft Copilot and ChatGPT. Optimizing for one model leaves you invisible to the rest.

