A crawler is an automated client that requests URLs, follows links to find more, and hands what it fetches to a search index, a model training corpus, or a live answer. This article covers what a crawler is, Googlebot and its variants, Google-Extended and AI training controls, AI crawlers and answer bots, robots.txt rules per crawler, and reading server logs to verify crawlers. It then shows crawlers on holisticradar.com, connects crawlers to crawl budget and indexation, and closes with blocking AI bots, llms.txt, and crawl rate limits.
What a crawler is
A crawler is a program that requests URLs automatically, discovers further URLs from the links it finds, and passes the fetched content to the system that operates it. Every crawler identifies itself with a user-agent string and, for robots.txt purposes, with a shorter token that site owners write rules against.
Crawlers differ by what the fetched content is used for, and that purpose decides whether blocking them costs a site anything. Three purposes cover almost every bot a site sees today.
| Purpose | What happens to the fetched page | Examples | Follows robots.txt |
|---|---|---|---|
| Search indexing | Stored in an index and ranked for queries, including the queries behind AI answers | Googlebot, Bingbot, OAI-SearchBot, Claude-SearchBot, PerplexityBot | Yes |
| Model training | Added to data used to train future AI models | GPTBot, ClaudeBot; Google-Extended as a control token | Yes |
| User-triggered fetching | Retrieved once because a person asked an assistant or a tool to read the page | ChatGPT-User, Claude-User, Perplexity-User, Google’s user-triggered fetchers | Varies by operator |
Google draws the same line in its own documentation. It separates common crawlers, which always respect robots.txt, from special-case crawlers and from user-triggered fetchers, which act on a user’s request and generally ignore robots.txt rules. The distinction matters for diagnosis: a fetcher hitting a disallowed URL is not a broken robots.txt file.
Googlebot and its variants
Googlebot is Google’s main crawler for Search, and it runs as Googlebot Smartphone and Googlebot Desktop under one robots.txt token, Googlebot. Google uses mobile-first indexing, so the smartphone crawler fetches the version of the page that Google indexes for almost every site, and a page that hides content from mobile visitors hides it from the index.
| Crawler | robots.txt token | What it fetches |
|---|---|---|
| Googlebot Smartphone | Googlebot | Pages for Search, as a mobile browser renders them; the primary indexing crawler |
| Googlebot Desktop | Googlebot | Pages for Search, as a desktop browser renders them |
| Googlebot Image | Googlebot-Image | Image files and the pages that embed them, for image results |
| Googlebot Video | Googlebot-Video | Video files and the pages that embed them, for video results |
| Googlebot News | Googlebot-News | Articles for Google News |
| Google-InspectionTool | Google-InspectionTool | Pages requested by Search testing tools such as URL Inspection and the Rich Results Test |
| GoogleOther | GoogleOther | Pages for Google product teams’ research and development, outside Search |
| Storebot-Google | Storebot-Google | Product and checkout pages for Google Shopping |
The variants share Googlebot’s rules unless a site writes a group for their own token. Google’s crawlers obey the most specific user-agent group that matches them, so a rule written for Googlebot-Image overrides the Googlebot group for image crawling, and with no such group the image crawler follows the Googlebot rules.
Googlebot renders pages with an evergreen version of Chromium after fetching them. That makes it the exception among crawlers: most bots outside Google and Microsoft read only the HTML response.
Google-Extended and AI training controls
Google-Extended is a robots.txt product token that controls whether content Google crawls may be used to train future Gemini models and to ground answers in Gemini apps and Vertex AI, and it is not a separate fetcher. Google’s crawler documentation states that Google-Extended has no HTTP user-agent string of its own: the crawling is done by Google’s existing crawlers, and the token only tells Google how the content may be used afterwards.
Google also states that Google-Extended does not affect a site’s inclusion in Google Search and is not used as a ranking signal. Disallowing it removes content from Gemini training and grounding without changing anything in Search.
The token does not control AI features inside Google Search. AI Overviews and AI Mode are Search features, built from pages Googlebot crawled and indexed, so the controls that apply to them are the Search controls: noindex, nosnippet, data-nosnippet, and max-snippet. A site that wants to stay in Search but keep a passage out of AI Overviews uses snippet controls, not Google-Extended.
Other operators use the same pattern. Apple publishes Applebot-Extended, a token that controls whether content Applebot crawled is used to train Apple’s foundation models, while Applebot itself keeps crawling for Siri and Spotlight. Training permission and crawling permission are separate decisions, and the robots.txt file has to state both.
AI crawlers and answer bots
AI crawlers fall into three groups: training crawlers that collect data for future models, search crawlers that build the indexes AI answers retrieve from, and user-triggered fetchers that read a page during a conversation. Each major operator now runs a separate bot for each job, so a site can allow one and block another.
| Bot | Operator | Job | robots.txt |
|---|---|---|---|
| GPTBot | OpenAI | Collects content that may be used to train OpenAI’s generative models | Respected |
| OAI-SearchBot | OpenAI | Indexes sites so they can appear in ChatGPT search results | Respected |
| ChatGPT-User | OpenAI | Fetches pages for actions a user starts in ChatGPT | OpenAI states the rules may not apply |
| ClaudeBot | Anthropic | Collects content for training Anthropic’s models | Respected |
| Claude-SearchBot | Anthropic | Crawls to improve the quality of Claude’s search results | Respected |
| Claude-User | Anthropic | Fetches pages when a Claude user asks a question that needs them | Respected |
| PerplexityBot | Perplexity | Indexes sites to surface and link them in Perplexity answers; not used for training | Respected |
| Perplexity-User | Perplexity | Fetches pages a user’s question requires | Perplexity states it generally ignores the rules |
| Bingbot | Microsoft | Builds the Bing index, which grounds Microsoft Copilot answers | Respected |
The operational point is that the search crawlers decide visibility in AI answers. OpenAI states that a site which blocks OAI-SearchBot will not appear in ChatGPT search answers, though it may still appear as a navigational link. Blocking GPTBot, by contrast, only signals that content should not be used for training. A robots.txt file that blocks every AI user agent to stop training also removes the site from the answers those operators serve.
Copilot has no separate answer crawler. Microsoft grounds Copilot in the Bing index, so Bingbot access and Bing indexation decide whether a page can be cited there. Bing also supports the nocache and noarchive robots meta values to limit how page content is used in its chat answers.
Bot names change. OpenAI added OAI-AdsBot for checking ad landing pages, and Anthropic now documents three separate bots, each with its own token. The operator’s own documentation is the record to check before writing rules.
robots.txt rules per crawler
A robots.txt file assigns rules to crawlers by user-agent group, and each crawler obeys only the most specific group that matches its token. The table shows the rule a site writes for common decisions.
| Decision | User-agent token | Rule | Effect |
|---|---|---|---|
| Stay in Google Search, opt out of Gemini training and grounding | Google-Extended | Disallow: / | Search inclusion and ranking unchanged |
| Opt out of OpenAI training, stay in ChatGPT search | GPTBot; OAI-SearchBot | Disallow: / for GPTBot; Allow: / for OAI-SearchBot | Content cited in ChatGPT search, not used for training |
| Opt out of Anthropic training, stay in Claude search | ClaudeBot; Claude-SearchBot | Disallow: / for ClaudeBot; Allow: / for Claude-SearchBot | Same split for Anthropic |
| Stay in Perplexity answers | PerplexityBot | Allow: / | Pages can be surfaced and linked in answers |
| Keep an image folder out of image results | Googlebot-Image | Disallow: /private-images/ | Web Search crawling of pages unchanged |
| Slow Anthropic’s crawlers | ClaudeBot | Crawl-delay: 1 | Honored by Anthropic; Google ignores Crawl-delay |
Two limits apply to every rule in the table. robots.txt controls crawling, not indexing: Google can index a disallowed URL without its content when other pages link to it, and a noindex directive on a disallowed page is never seen because the page is never fetched. And robots.txt is a request, not an access control: well-behaved operators honor it, and anything that must stay private needs authentication.
Reading server logs to verify crawlers
Server logs verify a crawler by its source IP address, because a user-agent string is self-declared and any scraper can copy Googlebot’s. Verification runs in a fixed order:
- Filter the access log for requests whose user-agent claims to be the crawler in question.
- Run a reverse DNS lookup on each source IP. Google’s crawlers resolve to hostnames under googlebot.com, google.com, or googleusercontent.com; Bingbot resolves under search.msn.com.
- Run a forward DNS lookup on the returned hostname and confirm it resolves to the same IP. A match verifies the crawler; a mismatch marks an impersonator.
- For operators that publish IP ranges, such as Google, OpenAI, and Perplexity, match the IP against the published list instead of, or as well as, the DNS checks.
Verified requests then answer the questions that crawl reports in Search Console cannot answer for other engines: how often each bot visits, which status codes it receives, what share of its requests land on URLs in the sitemap rather than on parameters and archives, and how long a new page waits before its first request from each bot. A training crawler fetching thousands of URLs a day while the search crawler of the same operator fetches a few is a pattern only logs show.
Crawlers on holisticradar.com
Holistic Radar gives crawlers on holisticradar.com one XML sitemap, generated by Rank Math, that lists only URLs resolving with a 200 status. The theme filters the sitemap entries: any URL that appears in the site’s redirect map, and any old article slug that now redirects to a live permalink, is removed before the sitemap is served, so a crawler never spends a request on a URL that answers with a 301.
| Crawler-facing surface | Rule on holisticradar.com |
|---|---|
| XML sitemap | Generated by Rank Math; redirected URLs and old slugs excluded; category, tag, and post format archives excluded |
| Thin views | Category, tag, date, search, and author views, paginated blog listings, and 404 pages carry noindex, follow |
| Author archives | 301 to the matching team profile, so every author resolves to one page |
| Attachment pages | 301 to the parent page |
| Old article slugs | 301 to the one live permalink under /blog/ |
| llms.txt | Rank Math’s llms.txt module enabled, as a summary for clients that read it |
| Structured data | One JSON-LD graph output by the theme; Rank Math’s own JSON-LD disabled so nothing is described twice |
The noindex, follow rule is the one that matters most for crawlers. Thin views stay out of the index, but the links on them still pass, so a crawler that lands on an archive continues to the articles it lists. Every page a crawler can index on the site is a page in the topical map.
Crawlers and crawl budget and indexation
Crawlers connect to crawl budget and indexation because every request a crawler makes is spent either on a page in the topical map or on a URL that should not exist. Knowing which crawlers fetch a site answers who is reading it; it does not answer whether they spend their requests on the right URLs, or whether the pages they fetch end up in the index.
Those two questions are crawl budget and indexation. Every redirect, parameter URL, and thin archive a crawler requests is a request not spent on a page in the topical map, and every map page that is crawled but not indexed is coverage the site planned and does not have.
A reader who can now name the bots in a log file needs the rules that point those bots at the right URLs and the metric that shows whether the map is actually indexed.
→ Crawl budget and indexation: how Holistic Radar gets every map page crawled and indexed and keeps everything else out.
Crawlers: blocking AI bots, llms.txt, and crawl rate limits
Should a site block AI bots?
A site should block AI training crawlers only if it does not want its content used in future models, and it should not block AI search crawlers if it wants to be cited. Blocking GPTBot, ClaudeBot, or Google-Extended removes content from training without affecting search. Blocking OAI-SearchBot, Claude-SearchBot, PerplexityBot, or Bingbot removes the site from the answers those indexes feed.
Does llms.txt control what crawlers can access?
llms.txt does not control access; it is a proposed Markdown file that summarizes a site for language models and carries no allow or disallow rules. robots.txt remains the only standard control. Google states that no AI text files or special markup are needed to appear in its AI features, so llms.txt is a low-cost convenience for clients that choose to read it.
How do you slow a crawler that is overloading a server?
A crawler is slowed by the method its operator supports. Google ignores Crawl-delay and reduces its crawling when a site returns 500, 503, or 429 responses, which should be used only temporarily because persistent errors also slow indexing. Anthropic and Bing honor Crawl-delay in robots.txt, and Bing Webmaster Tools also offers crawl control settings.
Can a URL blocked in robots.txt still appear in search?
A URL blocked in robots.txt can still appear in Google’s results, usually without a description, when other pages link to it. Google indexes the URL from the links without fetching the page. Keeping a page out of the index requires allowing the crawl and serving noindex, or removing the page.