AI search glossary: the terms that matter, with sources

Twenty-five AI search terms, each with a plain one-line definition and a link to the primary source. Built so you can check every entry rather than trust it.

7 min readAdarsh Mishra

On this page

Every entry here links to a primary source: the crawler operator's own documentation, Google's own docs, the RFC, or the paper. Where no primary source exists, the entry says so, because that is usually the most useful thing to know about a term.

Alphabetical. Each entry stands on its own.

AEO (answer engine optimization)

Optimising content to be selected as a direct answer rather than a link. No founding paper, no standards body and no dated first use that anyone citing an origin actually links to. Wikipedia's entry on the surrounding terminology records that no consensus definition separated AEO, GEO and AI SEO as of early 2026.

AI Mode

Google's conversational search surface, launched as a separate mode inside Search. Google's announcement is the primary source. Distinct from an AI Overview, which appears above ordinary results rather than replacing them.

AI Overview

The generated summary Google places above search results, with links to supporting pages. Rolled out broadly in May 2024. Gemini is a different product with different crawler controls, and the two are conflated constantly.

CCBot

Common Crawl's crawler. It builds a free public web archive that many organisations, including AI labs, then use as a training corpus. See Common Crawl's own page. Blocking CCBot affects a dataset, not any live search product.

ChatGPT-User

OpenAI's fetcher for user-initiated actions inside ChatGPT and custom GPTs, listed separately from GPTBot and OAI-SearchBot in OpenAI's bot documentation. It fires when a person asks for something, not on a crawl schedule.

Citation

A named or linked source inside a generated answer. Distinct from a ranking position: a page can be cited without ranking, and rank without being cited. Google describes AI Overviews and AI Mode as displaying supporting web pages in its AI features documentation. What tends to get picked is covered in what gets cited in AI answers.

ClaudeBot, Claude-SearchBot and Claude-User

Anthropic's three crawlers, all separately controllable. Per Anthropic's documentation: ClaudeBot collects content that may contribute to training, Claude-SearchBot improves search result quality, and Claude-User fetches pages when a person asks Claude a question.

E-E-A-T

Experience, expertise, authoritativeness and trustworthiness. Google's helpful content guidance states that "E-E-A-T itself isn't a specific ranking factor", and describes it as a way of naming what a mix of other signals is meant to identify. Anyone selling you an E-E-A-T score is selling their own invention.

A verbatim passage Google extracts from one page and shows at the top of results. You cannot opt in. Google's documentation answers the question of how to mark up a page as a featured snippet with two words: "You can't."

GEO (generative engine optimization)

Editing content to raise its share of a generated answer. Backed by one paper, Aggarwal et al., KDD 2024, which tested nine content edits and found keyword stuffing scored below making no changes at all. We read the whole thing in what the GEO paper actually measured.

GEO-bench

The benchmark from that paper: 10,000 queries from nine sources across 25 domains, each paired with the cleaned text of the top five Google results. Split 8,000 train, 1,000 validation, 1,000 test. Roughly 80% informational queries.

Google-Extended

A robots.txt token, not a crawler. It controls whether Google may use your content for training and grounding in Gemini Apps and Vertex AI. Google's crawler documentation states it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search". It has no user agent string of its own.

Googlebot

The crawler behind Google Search, documented in Google's common crawlers list. AI Overviews and AI Mode are Search surfaces and Google publishes no separate crawler for them. Blocking Googlebot removes you from Search entirely, which is a far larger decision than blocking a training bot.

GPTBot

OpenAI's training crawler. Its documentation says it "is used to crawl content that may be used in training our generative AI foundation models". It is not what powers ChatGPT search. See does blocking GPTBot remove you from ChatGPT.

Grounding

Attaching a model's answer to retrieved documents so the output can cite them. Google uses the term for Gemini Apps and Vertex AI, both covered by Google-Extended rather than by Search controls. Grounding is why a model can name a page it was never trained on.

llms.txt

A proposal by Jeremy Howard, first published September 2024 and revised August 2026, for a markdown file that helps agents use a website. Several AI labs publish one for their own docs. No major AI search product has documented reading one. The three files people confuse are compared in llms.txt vs robots.txt vs sitemap.

nosnippet, data-nosnippet and max-snippet

Google's robots meta rules for limiting what Search can display from a page, documented here. These, not Google-Extended, are the controls Google names when you want to limit what appears in AI Overviews. They reduce your visibility everywhere in Search at the same time.

OAI-SearchBot

OpenAI's search crawler. Its documentation says it "is used to surface websites in search results in ChatGPT's search features". A separate robots.txt token from GPTBot, and the one that matters if you want to appear in ChatGPT search.

PerplexityBot

Perplexity's search crawler, which its documentation describes as designed to surface and link websites in Perplexity results, explicitly "not used to crawl content for AI foundation models". Perplexity-User is the separate user-initiated fetcher, and Perplexity states that one generally ignores robots.txt.

Query fan-out

Google's term for issuing several related searches across subtopics before composing one answer. Described in its AI features documentation for both AI Overviews and AI Mode. It is why a cited page may rank for none of the terms you track.

Retrieval augmented generation (RAG)

Fetching documents at answer time and conditioning the model's output on them, rather than relying on training data. The originating paper is Lewis et al., 2020. Every AI search product is a RAG system with a product wrapped around it.

robots.txt

The exclusion protocol, standardised as RFC 9309 in 2022. It tells compliant crawlers what not to fetch. It does not remove content already fetched, does not stop non-compliant bots, and does not control what a model says about you. Whether to use it against AI bots is worked through in should you block AI crawlers.

Structured data

Machine-readable markup describing a page, usually schema.org JSON-LD. Google's search gallery lists the types that still earn a result. Google also states there is no special structured data needed for AI features. The HowTo type was removed in 2023 and FAQ stopped showing on 7 May 2026. What to do with the leftover markup is in HowTo schema after removal.

Training crawler vs search crawler

A training crawler collects text that may train a model. A search crawler fetches pages so they can be retrieved and cited in a live answer. GPTBot, ClaudeBot, CCBot and Google-Extended are the first kind. OAI-SearchBot, Claude-SearchBot and PerplexityBot are the second. They are separate robots.txt tokens, blocking one has no effect on the other, and confusing the two is how sites accidentally remove themselves from a search product they meant to stay in.

A search that ends without a click to any result. The best-sourced measurement is the Pew Research Center's March 2025 browsing study of 900 tracked US adults. Across 68,879 Google searches, users clicked a result on 8% of pages carrying an AI summary against 15% of pages without one. They clicked a link inside the summary on 1%. Pew also found browsing sessions ended on 26% of pages with a summary, against 16% without.

Terms we left out on purpose

LLMO, AIO, chat engine optimization. Coined labels for the same activity as AEO and GEO, with no distinct method behind any of them.

AI visibility score. Every vendor computes one differently and none publishes the formula. A score without a stated query set and a stated method is a vibe with a number on it.

Any citation-rate or AI-traffic-share statistic you have seen repeated without a link. We went looking for primary sources on several of the most quoted figures. What we found was vendor trackers citing each other, usually with the keyword sample undisclosed. Presence rates swing enormously with query mix. Pew measured about 18% of real searches producing an AI summary. Our own commercial buyer-intent questions came back at 100%. Both are honest numbers and neither one is the rate.

Our own measurements above come from asking Google live through Bright Data, listed at $0.0015 per question, which is list price rather than a figure diffed against our balance. A nine question scan of stripe.com cost $0.0135 and showed the domain cited in 44% of the questions Google answered, which at nine questions is a shape rather than a rate.

SEOBuilder is our product. It asks seven answer engines the same buyer questions, Google AI Overviews and Google AI Mode, ChatGPT, Perplexity, Gemini, Microsoft Copilot and Claude, and reports which of them cited you. If you are deciding whether to build this measurement yourself, we costed both routes in AI visibility, build vs buy, and the method by hand is in how to measure AI visibility. If you would rather see what an answer engine can already extract from one of your pages, the answer readiness tool runs client side and needs no account.

Filed under

  • glossary
  • ai-search
  • crawlers
  • definitions
  • strategy

Last updated 21 August 2026

Questions

What is the difference between a training crawler and a search crawler?
A training crawler collects text that may be used to train a model, such as GPTBot, ClaudeBot or Google-Extended. A search crawler fetches pages so they can be retrieved and cited in a live answer, such as OAI-SearchBot, Claude-SearchBot or PerplexityBot. They are separate robots.txt tokens and blocking one does nothing to the other.
Does blocking Google-Extended remove me from AI Overviews?
No. Google's crawler documentation states Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal. It governs training and grounding for Gemini Apps and Vertex AI.
Is llms.txt read by AI search engines?
Not by any that have documented it. The proposal is real and several AI labs publish an llms.txt for their own developer docs, but publishing a file is not the same as reading one. Google states you do not need to create machine readable or AI text files to appear in its AI features.

Related reading

Check the page, not the hunch

Is your page ready to be the source?

SEOBuilder asks 7 answer engines the questions your buyers ask and reports which answers cite you, which cite a competitor, and which cite nobody. Free to start, no card.

Or ask about one page right now: the free AI visibility check, no account and no card.

Run your first scan