Home / Publications / Blog / AEO — Answer Engine Optimization / Your Google Rank Just Stopped Predicting Whether AI Cites You

Your Google Rank Just Stopped Predicting Whether AI Cites You

A University of Toronto audit ran 1,000 ranking queries through Google and four AI answer engines and found they cite almost entirely different sources: GPT-4o overlapped Google's top ten 0% of the time for the median query. Here is what that means for how you measure search visibility.

By Foragentis teamPublished 2026-06-23Updated 2026-07-066 min read

For twenty years, one number told a business whether the web could find it: its Google rank. An audit out of the University of Toronto just showed that number now tells you almost nothing about whether an AI answer engine will cite you.

The study, "Navigating the Shift: A Comparative Analysis of Web Search and Generative AI Response Generation" (Chen, Wang, Chen, and Koudas, Department of Computer Science, University of Toronto), appeared in the workshop proceedings of the EDBT/ICDT 2026 Joint Conference. Its domain-overlap experiment ran 1,000 ranking-style queries through five systems side by side: Google Search, GPT-4o, Claude 4.5 Sonnet, Perplexity Sonar Pro, and Gemini 2.5 Flash. Separate experiments in the same paper used their own, smaller query sets, so the numbers below belong to the experiment that produced them rather than to one pooled total. The PR trade outlet that summarized it for marketers called it the largest empirical audit of generative AI search published to date.

The number that should reset your dashboard

The researchers measured how often each AI engine's cited domains overlapped with Google's top ten results for the same query. The overlap is small, and it is different for every engine:

  • GPT-4o: 4.0% mean overlap, 0.0% median. A zero median means that for more than half of the queries tested, GPT-4o cited zero domains in common with Google's top ten.
  • Gemini 2.5 Flash: 11.1% mean, 8.5% median.
  • Claude 4.5 Sonnet: 12.6% mean, 8.7% median.
  • Perplexity Sonar Pro: 15.2% mean, 14.3% median.

Every pairwise difference between engines held up under paired bootstrap resampling over the same query set. This is not noise. It is a structural finding: the engines that increasingly answer your buyers' questions are reading from a different web than the one your SEO report tracks.

It gets sharper. The engines disagree with each other, not only with Google. The paper also measured overlap between each AI engine and Gemini and found it low throughout, so ranking well inside ChatGPT does not carry over to Perplexity, which does not carry over to Gemini. And the divergence is not only about which URLs get cited, it is about which kinds of sources. In a separate 300-query test on consumer electronics, Claude 4.5 Sonnet drew 65% of its sources from earned media and 1% from social content, against 57% and 8% for GPT-4o, while Google split 41% earned, 34% social and 26% brand. The authors attach a caveat to Claude's figures that is worth carrying: it returned no links at all for most informational and transactional queries unless it was explicitly prompted to search. Each engine has its own editorial center of gravity.

Why this breaks the standard SEO playbook

The entire discipline of search optimization assumes one ranked list that most people see. Optimize the page, earn the links, climb the list, get the traffic. That model quietly assumed the list was singular and stable.

The Toronto audit describes a world with five different lists, mostly non-overlapping, each weighting sources by its own logic, and several of them not exposed to you at all. A business can sit at position one on Google and be invisible in the AI answer that a buyer actually reads. Nothing in a rank-tracking dashboard would tell you that, because rank tracking measures the one surface that is now the minority case.

This is the gap that answer engine optimization (AEO) and generative engine optimization (GEO) exist to close. But you cannot optimize a surface you cannot measure, and most marketing teams have no instrument pointed at the AI-answer surface at all. They are flying the AI-search era on a Google-era altimeter.

What we built, and why this study is the proof point

ForaPost SEO Intelligence exists for exactly this measurement problem. It gives a business a real-time SERP landscape, content-gap detection, and AEO/GEO visibility signals, so you can see how you actually show up across answer engines, not just where you rank on Google.

We did not build it as a reaction to this paper. We built it because the decoupling the paper measures has been visible in the field for a while. What the Toronto audit adds is a hard number on a trend that was easy to wave away: a median overlap of zero is not a rounding error, it is a different web.

The practical takeaway is not "panic about AI search." It is narrower and more useful: stop treating one ranked list as the whole picture, and start measuring visibility per engine, because the engines no longer agree with each other or with Google. The businesses that adapt first will be the ones that can see the new surface clearly. That is a measurement problem before it is a content problem, and measurement is something you can start on this quarter.

FAQ

Does Google rank predict whether an AI engine will cite you?

No longer, according to a University of Toronto audit. Across 1,000 ranking-style queries it measured how often each AI engine's cited domains overlapped with Google's top ten results for the same query and found the overlap small and engine-specific. Mean overlap was 4.0% for GPT-4o, 11.1% for Gemini 2.5 Flash, 12.6% for Claude 4.5 Sonnet and 15.2% for Perplexity Sonar Pro. GPT-4o's median overlap was 0.0%, meaning that for more than half the queries it cited no domain in common with Google's top ten. Your Google rank now tells you almost nothing about whether an AI answer engine will cite you.

How much do AI answer engines overlap with Google's top ten results?

Very little, and it varies by engine. In the University of Toronto audit, mean overlap with Google's top ten ran from 4.0% for GPT-4o to 15.2% for Perplexity Sonar Pro, with Gemini and Claude in between. Medians were lower still: 0.0% for GPT-4o, 8.5% for Gemini, 8.7% for Claude and 14.3% for Perplexity. Every pairwise difference was significant under paired bootstrap resampling over the same query set, so they are structural rather than noise.

Do different AI engines cite the same sources as each other?

Mostly not. Every pairwise overlap the audit measured was low, and overlap between each AI engine and Gemini was low too, so ranking well inside ChatGPT does not carry over to Perplexity, which does not carry over to Gemini. The divergence is also about source type. In a separate 300-query consumer-electronics test, Claude 4.5 Sonnet drew 65% of its sources from earned media and 1% from social, against 57% and 8% for GPT-4o and a far more balanced split for Google. Read Claude's numbers with the authors' own caveat: it returned no links at all for most informational and transactional queries unless explicitly prompted to search. Each engine has its own editorial center of gravity, so visibility has to be measured per engine.

What is the difference between SEO and AEO/GEO?

SEO optimizes for one ranked list, Google's, on the assumption that one list is what most people see. Answer engine optimization (AEO) and generative engine optimization (GEO) optimize for the citations inside AI answers from engines like ChatGPT, Claude, Gemini, and Perplexity, which the Toronto audit shows draw from largely non-overlapping sources. Because those engines disagree with each other and with Google, AEO/GEO is a per-engine measurement problem before it is a content problem.


Sources: "Navigating the Shift: A Comparative Analysis of Web Search and Generative AI Response Generation," Chen, Wang, Chen, and Koudas, Department of Computer Science, University of Toronto. Published in the Proceedings of the Workshops of the EDBT/ICDT 2026 Joint Conference, CEUR-WS Vol-4192 (https://ceur-ws.org/Vol-4192/), and on arXiv as 2601.16858 (https://arxiv.org/abs/2601.16858). All figures above are read from the paper itself. Trade coverage: everything-pr.com, "Four Engines. One Thousand Queries. The Toronto AI Audit." (https://everything-pr.com/four-engines-one-thousand-queries-the-toronto-audit-of-how-ai-cites-the-web).