L5 · Retrieval and AI confidence

Search and AI Crawlers: Googlebot, AI Bots, and What Each Fetches

Search crawlers, training crawlers, and user-triggered fetchers do different jobs, so each needs its own robots.txt decision and its own check in the server logs.

CRAWLERSROBOTSGOOGLEBOTAI SEARCHTRAININGALLOWED

Key takeaways

  • A crawler requests URLs automatically and passes the content to a search index, a model training corpus, or a live answer, and that purpose decides what blocking it costs.
  • Googlebot runs as Googlebot Smartphone and Googlebot Desktop under one token, and mobile-first indexing makes the smartphone crawler the one whose fetch is indexed.
  • Google-Extended is a robots.txt product token with no separate fetcher, and it affects neither inclusion nor ranking in Google Search.
  • OAI-SearchBot, Claude-SearchBot, PerplexityBot, and Bingbot decide visibility in AI answers, while GPTBot and ClaudeBot only collect training data.
  • A crawler is verified by reverse DNS and a matching forward DNS lookup, because a user-agent string can be copied by anyone.

A crawler is an automated client that requests URLs, follows links to find more, and hands what it fetches to a search index, a model training corpus, or a live answer. This article covers what a crawler is, Googlebot and its variants, Google-Extended and AI training controls, AI crawlers and answer bots, robots.txt rules per crawler, and reading server logs to verify crawlers. It then shows crawlers on holisticradar.com, connects crawlers to crawl budget and indexation, and closes with blocking AI bots, llms.txt, and crawl rate limits.

What a crawler is

A crawler is a program that requests URLs automatically, discovers further URLs from the links it finds, and passes the fetched content to the system that operates it. Every crawler identifies itself with a user-agent string and, for robots.txt purposes, with a shorter token that site owners write rules against.

Crawlers differ by what the fetched content is used for, and that purpose decides whether blocking them costs a site anything. Three purposes cover almost every bot a site sees today.

PurposeWhat happens to the fetched pageExamplesFollows robots.txt
Search indexingStored in an index and ranked for queries, including the queries behind AI answersGooglebot, Bingbot, OAI-SearchBot, Claude-SearchBot, PerplexityBotYes
Model trainingAdded to data used to train future AI modelsGPTBot, ClaudeBot; Google-Extended as a control tokenYes
User-triggered fetchingRetrieved once because a person asked an assistant or a tool to read the pageChatGPT-User, Claude-User, Perplexity-User, Google’s user-triggered fetchersVaries by operator

Google draws the same line in its own documentation. It separates common crawlers, which always respect robots.txt, from special-case crawlers and from user-triggered fetchers, which act on a user’s request and generally ignore robots.txt rules. The distinction matters for diagnosis: a fetcher hitting a disallowed URL is not a broken robots.txt file.

Googlebot and its variants

Googlebot is Google’s main crawler for Search, and it runs as Googlebot Smartphone and Googlebot Desktop under one robots.txt token, Googlebot. Google uses mobile-first indexing, so the smartphone crawler fetches the version of the page that Google indexes for almost every site, and a page that hides content from mobile visitors hides it from the index.

Crawlerrobots.txt tokenWhat it fetches
Googlebot SmartphoneGooglebotPages for Search, as a mobile browser renders them; the primary indexing crawler
Googlebot DesktopGooglebotPages for Search, as a desktop browser renders them
Googlebot ImageGooglebot-ImageImage files and the pages that embed them, for image results
Googlebot VideoGooglebot-VideoVideo files and the pages that embed them, for video results
Googlebot NewsGooglebot-NewsArticles for Google News
Google-InspectionToolGoogle-InspectionToolPages requested by Search testing tools such as URL Inspection and the Rich Results Test
GoogleOtherGoogleOtherPages for Google product teams’ research and development, outside Search
Storebot-GoogleStorebot-GoogleProduct and checkout pages for Google Shopping

The variants share Googlebot’s rules unless a site writes a group for their own token. Google’s crawlers obey the most specific user-agent group that matches them, so a rule written for Googlebot-Image overrides the Googlebot group for image crawling, and with no such group the image crawler follows the Googlebot rules.

Googlebot renders pages with an evergreen version of Chromium after fetching them. That makes it the exception among crawlers: most bots outside Google and Microsoft read only the HTML response.

Google-Extended and AI training controls

Google-Extended is a robots.txt product token that controls whether content Google crawls may be used to train future Gemini models and to ground answers in Gemini apps and Vertex AI, and it is not a separate fetcher. Google’s crawler documentation states that Google-Extended has no HTTP user-agent string of its own: the crawling is done by Google’s existing crawlers, and the token only tells Google how the content may be used afterwards.

Google also states that Google-Extended does not affect a site’s inclusion in Google Search and is not used as a ranking signal. Disallowing it removes content from Gemini training and grounding without changing anything in Search.

The token does not control AI features inside Google Search. AI Overviews and AI Mode are Search features, built from pages Googlebot crawled and indexed, so the controls that apply to them are the Search controls: noindex, nosnippet, data-nosnippet, and max-snippet. A site that wants to stay in Search but keep a passage out of AI Overviews uses snippet controls, not Google-Extended.

Other operators use the same pattern. Apple publishes Applebot-Extended, a token that controls whether content Applebot crawled is used to train Apple’s foundation models, while Applebot itself keeps crawling for Siri and Spotlight. Training permission and crawling permission are separate decisions, and the robots.txt file has to state both.

AI crawlers and answer bots

AI crawlers fall into three groups: training crawlers that collect data for future models, search crawlers that build the indexes AI answers retrieve from, and user-triggered fetchers that read a page during a conversation. Each major operator now runs a separate bot for each job, so a site can allow one and block another.

BotOperatorJobrobots.txt
GPTBotOpenAICollects content that may be used to train OpenAI’s generative modelsRespected
OAI-SearchBotOpenAIIndexes sites so they can appear in ChatGPT search resultsRespected
ChatGPT-UserOpenAIFetches pages for actions a user starts in ChatGPTOpenAI states the rules may not apply
ClaudeBotAnthropicCollects content for training Anthropic’s modelsRespected
Claude-SearchBotAnthropicCrawls to improve the quality of Claude’s search resultsRespected
Claude-UserAnthropicFetches pages when a Claude user asks a question that needs themRespected
PerplexityBotPerplexityIndexes sites to surface and link them in Perplexity answers; not used for trainingRespected
Perplexity-UserPerplexityFetches pages a user’s question requiresPerplexity states it generally ignores the rules
BingbotMicrosoftBuilds the Bing index, which grounds Microsoft Copilot answersRespected

The operational point is that the search crawlers decide visibility in AI answers. OpenAI states that a site which blocks OAI-SearchBot will not appear in ChatGPT search answers, though it may still appear as a navigational link. Blocking GPTBot, by contrast, only signals that content should not be used for training. A robots.txt file that blocks every AI user agent to stop training also removes the site from the answers those operators serve.

Copilot has no separate answer crawler. Microsoft grounds Copilot in the Bing index, so Bingbot access and Bing indexation decide whether a page can be cited there. Bing also supports the nocache and noarchive robots meta values to limit how page content is used in its chat answers.

Bot names change. OpenAI added OAI-AdsBot for checking ad landing pages, and Anthropic now documents three separate bots, each with its own token. The operator’s own documentation is the record to check before writing rules.

robots.txt rules per crawler

A robots.txt file assigns rules to crawlers by user-agent group, and each crawler obeys only the most specific group that matches its token. The table shows the rule a site writes for common decisions.

DecisionUser-agent tokenRuleEffect
Stay in Google Search, opt out of Gemini training and groundingGoogle-ExtendedDisallow: /Search inclusion and ranking unchanged
Opt out of OpenAI training, stay in ChatGPT searchGPTBot; OAI-SearchBotDisallow: / for GPTBot; Allow: / for OAI-SearchBotContent cited in ChatGPT search, not used for training
Opt out of Anthropic training, stay in Claude searchClaudeBot; Claude-SearchBotDisallow: / for ClaudeBot; Allow: / for Claude-SearchBotSame split for Anthropic
Stay in Perplexity answersPerplexityBotAllow: /Pages can be surfaced and linked in answers
Keep an image folder out of image resultsGooglebot-ImageDisallow: /private-images/Web Search crawling of pages unchanged
Slow Anthropic’s crawlersClaudeBotCrawl-delay: 1Honored by Anthropic; Google ignores Crawl-delay

Two limits apply to every rule in the table. robots.txt controls crawling, not indexing: Google can index a disallowed URL without its content when other pages link to it, and a noindex directive on a disallowed page is never seen because the page is never fetched. And robots.txt is a request, not an access control: well-behaved operators honor it, and anything that must stay private needs authentication.

Reading server logs to verify crawlers

Server logs verify a crawler by its source IP address, because a user-agent string is self-declared and any scraper can copy Googlebot’s. Verification runs in a fixed order:

  1. Filter the access log for requests whose user-agent claims to be the crawler in question.
  2. Run a reverse DNS lookup on each source IP. Google’s crawlers resolve to hostnames under googlebot.com, google.com, or googleusercontent.com; Bingbot resolves under search.msn.com.
  3. Run a forward DNS lookup on the returned hostname and confirm it resolves to the same IP. A match verifies the crawler; a mismatch marks an impersonator.
  4. For operators that publish IP ranges, such as Google, OpenAI, and Perplexity, match the IP against the published list instead of, or as well as, the DNS checks.

Verified requests then answer the questions that crawl reports in Search Console cannot answer for other engines: how often each bot visits, which status codes it receives, what share of its requests land on URLs in the sitemap rather than on parameters and archives, and how long a new page waits before its first request from each bot. A training crawler fetching thousands of URLs a day while the search crawler of the same operator fetches a few is a pattern only logs show.

Crawlers on holisticradar.com

Holistic Radar gives crawlers on holisticradar.com one XML sitemap, generated by Rank Math, that lists only URLs resolving with a 200 status. The theme filters the sitemap entries: any URL that appears in the site’s redirect map, and any old article slug that now redirects to a live permalink, is removed before the sitemap is served, so a crawler never spends a request on a URL that answers with a 301.

Crawler-facing surfaceRule on holisticradar.com
XML sitemapGenerated by Rank Math; redirected URLs and old slugs excluded; category, tag, and post format archives excluded
Thin viewsCategory, tag, date, search, and author views, paginated blog listings, and 404 pages carry noindex, follow
Author archives301 to the matching team profile, so every author resolves to one page
Attachment pages301 to the parent page
Old article slugs301 to the one live permalink under /blog/
llms.txtRank Math’s llms.txt module enabled, as a summary for clients that read it
Structured dataOne JSON-LD graph output by the theme; Rank Math’s own JSON-LD disabled so nothing is described twice

The noindex, follow rule is the one that matters most for crawlers. Thin views stay out of the index, but the links on them still pass, so a crawler that lands on an archive continues to the articles it lists. Every page a crawler can index on the site is a page in the topical map.

Crawlers and crawl budget and indexation

Crawlers connect to crawl budget and indexation because every request a crawler makes is spent either on a page in the topical map or on a URL that should not exist. Knowing which crawlers fetch a site answers who is reading it; it does not answer whether they spend their requests on the right URLs, or whether the pages they fetch end up in the index.

Those two questions are crawl budget and indexation. Every redirect, parameter URL, and thin archive a crawler requests is a request not spent on a page in the topical map, and every map page that is crawled but not indexed is coverage the site planned and does not have.

A reader who can now name the bots in a log file needs the rules that point those bots at the right URLs and the metric that shows whether the map is actually indexed.

→ Crawl budget and indexation: how Holistic Radar gets every map page crawled and indexed and keeps everything else out.

Crawlers: blocking AI bots, llms.txt, and crawl rate limits

Should a site block AI bots?

A site should block AI training crawlers only if it does not want its content used in future models, and it should not block AI search crawlers if it wants to be cited. Blocking GPTBot, ClaudeBot, or Google-Extended removes content from training without affecting search. Blocking OAI-SearchBot, Claude-SearchBot, PerplexityBot, or Bingbot removes the site from the answers those indexes feed.

Does llms.txt control what crawlers can access?

llms.txt does not control access; it is a proposed Markdown file that summarizes a site for language models and carries no allow or disallow rules. robots.txt remains the only standard control. Google states that no AI text files or special markup are needed to appear in its AI features, so llms.txt is a low-cost convenience for clients that choose to read it.

How do you slow a crawler that is overloading a server?

A crawler is slowed by the method its operator supports. Google ignores Crawl-delay and reduces its crawling when a site returns 500, 503, or 429 responses, which should be used only temporarily because persistent errors also slow indexing. Anthropic and Bing honor Crawl-delay in robots.txt, and Bing Webmaster Tools also offers crawl control settings.

Can a URL blocked in robots.txt still appear in search?

A URL blocked in robots.txt can still appear in Google’s results, usually without a description, when other pages link to it. Google indexes the URL from the links without fetching the page. Keeping a page out of the index requires allowing the crawl and serving noindex, or removing the page.

→ Technical SEO checklist: the ordered checks that start with crawl access.

→ AI search optimization: what happens after an AI search crawler fetches a page.

→ JavaScript SEO: why most crawlers outside Google read only the HTML response.

Questions readers ask

Frequently asked questions

What is the difference between a crawler and a fetcher?

A crawler discovers and requests URLs automatically, while a fetcher requests a single URL because a user or tool asked for it. Google's user-triggered fetchers generally ignore robots.txt for that reason.

Does blocking GPTBot remove a site from ChatGPT?

Blocking GPTBot does not remove a site from ChatGPT search. ChatGPT search relies on OAI-SearchBot, and GPTBot only collects content for model training.

Which crawler does Microsoft Copilot use?

Microsoft Copilot uses the Bing index, which Bingbot builds. Bingbot access and Bing indexation decide whether a page can be cited in Copilot answers.

How do you verify that a request really came from Googlebot?

A Googlebot request is verified with a reverse DNS lookup that returns a googlebot.com, google.com, or googleusercontent.com hostname, followed by a forward lookup that returns the same IP. Google also publishes its crawler IP ranges.

Does Google-Extended stop AI Overviews from using a page?

Google-Extended does not stop AI Overviews from using a page, because AI Overviews are a Search feature. The controls for them are noindex, nosnippet, data-nosnippet, and max-snippet.

Sourcing

Sources

Crawler names, jobs, and robots.txt behavior are taken from each operator's own documentation as published in 2026, and the worked example from Holistic Radar's own site configuration.

  1. Google Crawling Infrastructure, Overview of Google crawlers and fetchers
  2. Google Crawling Infrastructure, Google's common crawlers
  3. OpenAI, Overview of OpenAI crawlers
  4. Anthropic, Does Anthropic crawl data from the web, and how can site owners block the crawler? (help center)
  5. Perplexity, Perplexity crawlers

Author

Founder · Semantic SEO & Conversion Architecture

Bilal Sameer founded Holistic Radar and leads its topical map strategy: central entity, source context, and the query network behind every map. His work sits inside the Holistic Radar team, alongside the engineers, reviewers, and designers who deliver each engagement.

View the full profile

Get my free Radar ScanFree Radar Scan