Which index do AI answer engines search? Only one of them says

Google documents that AI Overviews retrieve from the Search index. OpenAI, Perplexity and Anthropic document a crawler and stop. Checked 19 August 2026.

6 min readAdarsh Mishra

On this page

Ask which search index sits behind an AI answer and you get a confident answer from almost everyone except the companies running them. Of the four engines below, checked against their own live documentation on 19 August 2026, exactly one states where its answers are retrieved from.

That gap changes what you can reasonably plan. If you cannot know which index an engine reads, then "get into that index" is not a task anybody can hand you, and the work shifts to the parts that are documented.

Key Takeaways

  • Google states it plainly: AI Overviews use retrieval augmented generation against the Search index, through the core Search ranking systems. Its guide was last updated 10 July 2026.
  • OpenAI, Perplexity and Anthropic each document a crawler for their search feature and say nothing about the index behind it.
  • Every one of those three does document which agent governs whether you can appear at all, and that is the lever worth pulling.
  • OpenAI states that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, which is the clearest published cause and effect on this page.
  • Microsoft Copilot is missing from the table below because its own help page could not be read, and this post leaves the row empty rather than filling it from elsewhere.

What each vendor actually documents

Checked on 19 August 2026, against each vendor's own pages rather than against coverage of them.

Engine Agent it documents for search Index it says it retrieves from
Google AI Overviews and AI Mode Googlebot, with Google-Extended governing training rather than appearance The Google Search index, stated
ChatGPT search OAI-SearchBot Not stated
Perplexity PerplexityBot, plus Perplexity-User for a page fetched to answer one question Not stated
Claude Claude-SearchBot, plus Claude-User for a page fetched while answering Not stated
Microsoft Copilot Could not be read, see below Could not be read

One column is nearly full and the other is nearly empty, which is the finding.

Google is the exception, and says it in plain language

Google's guide to optimizing for generative AI features, last updated 10 July 2026, describes retrieval augmented generation as relying on its core Search ranking systems to retrieve relevant, up to date pages from the Search index, and then reviewing those pages to generate the response.

Two consequences follow, and both are more useful than most advice written about AI Overviews.

The first is that ordinary indexing work is the qualifying round. A page Google has not indexed cannot be retrieved by a system that retrieves from the Google index. There is no separate AI submission route to compensate.

The second is what Google says you do not need. The same page states you do not need to create machine readable files, AI text files, markup or Markdown to appear in Google Search including its generative AI capabilities, and adds that structured data is not required for generative AI search and there is no special schema markup to add. That is a vendor telling you to stop buying a category of work, which is worth more than most of what is sold in it. Our post on what is actually known about how AI Overviews pick sources goes further into what the ranking side does and does not reveal.

The other three document the door, not the room

OpenAI's crawler documentation names OAI-SearchBot as the agent that surfaces websites in ChatGPT's search features, separate from GPTBot for training. It also carries the sharpest published statement of consequence anywhere in this subject: sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers. What it does not say is what happens after that crawl, or which index the search step queries.

You will read that ChatGPT search runs on Bing. That may well be right. It is not on any OpenAI page that could be opened and quoted for this post, so it stays out of the table. The distance between "widely reported" and "documented by the vendor" is exactly the distance this blog tries to keep.

Perplexity's crawler guide names PerplexityBot, designed to surface and link websites in search results on Perplexity and explicitly not used to crawl for foundation model training, and Perplexity-User, which visits a page to answer a specific question. Nothing about the index.

Anthropic's crawler article, published 7 April 2026, names three: ClaudeBot for content that could contribute to training, Claude-User for a page fetched while answering somebody, and Claude-SearchBot, described as navigating the web to improve search result quality. Again nothing about where the search runs.

The pattern is consistent enough to be a decision rather than an oversight. Every vendor publishes the part a site owner controls and none publishes the part that would let anyone reverse engineer the ranking.

The empty row

Microsoft Copilot belongs in the table and is not in it.

The Microsoft Advertising help page for ads in Copilot returned a shell of navigation, a copyright line and cookie preferences, with the substance loaded afterwards by script. Nothing that reads the page without executing JavaScript gets anything, which is a small irony on a page about machine consumption of the web, and it means there is no sentence to quote.

Leaving the row empty is the honest version. A table where one cell says "could not be read" is worth more than a complete one where a reader has no way to tell which cells were checked.

What this changes about the work

If four out of five engines will not say where they retrieve from, then anybody promising to get you into the index behind an AI answer is describing a mechanism nobody has published.

What is documented is narrower and more useful. Each vendor names the agent that governs whether you can appear in its answers at all, and at least one states the consequence of blocking it outright. That makes crawler policy the one place where a published cause connects to a published effect, and it is a thirty minute job rather than a strategy. Our full list of AI crawler tokens has each one with its operator's own wording.

Everything after the crawl is unobservable from outside. You cannot see the index, the retrieval, the ranking, or why one page won. What you can see is the answer, and the answer names its sources.

That is why measuring the output is the only feedback loop available here, and why we built the product around reading answers rather than modelling the systems that produce them. Nobody outside these companies can model those systems honestly, and the ones claiming to are selling a diagram.

What to do this week

  1. Confirm the pages you care about are indexed by Google. For AI Overviews that is the documented qualifying round, not a proxy for it.
  2. Check your robots.txt against the agents each vendor names for its search feature, and separate them from the training agents. Blocking training is a choice. Blocking search is a removal.
  3. Stop buying work Google has said it does not use. Machine readable files and extra schema are not the route into generative results, by Google's own statement of 10 July 2026.
  4. Re-read this table before you rely on it. Four of its five rows are a vendor's current wording, and vendors change wording. The date at the top is there so you know how old this is.
  5. Measure the answer, since it is the only observable output. Run the free check, no account and no card, and see whether a real buyer question about your topic names you, and which sources it used instead.

Filed under

  • ai-overviews
  • chatgpt
  • perplexity
  • claude
  • retrieval
  • ai-crawlers

Last updated 19 August 2026

Questions

Do Google's AI Overviews use the normal Google Search index?
Yes, and Google says so directly. Its guide to optimizing for generative AI features, last updated 10 July 2026, describes retrieval augmented generation as relying on Google's core Search ranking systems to retrieve relevant, up to date web pages from the Search index. Google is the only major vendor that documents the index behind its answers.
Does ChatGPT search use Bing?
Not according to any OpenAI page that can be opened and quoted. OpenAI documents a crawler, OAI-SearchBot, whose job is surfacing sites in ChatGPT's search features, and states that sites opted out of it will not be shown in those answers. Which index sits behind that step is not stated in that documentation. Secondary reporting says Bing, and that is not the same as a vendor statement.
Does Perplexity have its own index?
Perplexity's crawler documentation, checked on 19 August 2026, names PerplexityBot as designed to surface and link websites in search results on Perplexity, and Perplexity-User for pages fetched to answer a specific question. It does not state whether the retrieval runs on an index Perplexity builds itself.
How does Claude find web pages?
Anthropic's crawler article, published 7 April 2026, names three agents: ClaudeBot for collecting content that could contribute to training, Claude-User for pages fetched while answering someone, and Claude-SearchBot, which it describes as navigating the web to improve search result quality. The article does not name a search provider behind it.
What can I actually do about this?
Allow the crawler each engine documents for its search feature, because that is the one lever every vendor publishes and the one that can disqualify you outright. Past that point nothing is observable from the outside, so the only feedback available is asking the question and reading the answer.

Related reading

Check the page, not the hunch

Is your page ready to be the source?

SEOBuilder asks 7 answer engines the questions your buyers ask and reports which answers cite you, which cite a competitor, and which cite nobody. Free to start, no card.

Or ask about one page right now: the free AI visibility check, no account and no card.

Run your first scan