Rankwise logoRankwise
Free audit
PricingCompareResources
Sign inGet started

Footer

Rankwise

The infrastructure layer for AI search: measure what AI says, publish what wins, and pinpoint the fixes that get you cited.

Built for the answer engines

Product

  • Pricing
  • Templates

Solutions

  • For agencies
  • For in-house SEO
  • For founders
  • Use Cases
  • Integrations
  • Compare
  • Alternatives

Resources

  • Resources
  • Articles
  • Guides
  • Case Studies
  • Learn
  • Glossary
  • Topics
  • API

Company

  • About
  • Contact
ImpressumAGBDatenschutz

© 2026 Rankwise. All rights reserved. Built by Founder Ventures.

Tracking ChatGPT · Perplexity · Claude · Google AI Overviews
  1. Rankwise
  2. /AI crawlers
Sources checked 2026-09-24

AI crawler directory

Who runs each AI crawler, what it does with your pages, whether it follows robots.txt, and the exact lines to allow or block it. Every fact links to the vendor's own documentation.
Check your robots.txt

24

Crawlers documented

11

Operators

7

Fetch on a user's request

Search index

Builds the index an AI answer searches. Block it and you drop out of that product's results.

User-triggered fetch

Visits a page because a person asked a question that needs it. This is the visit that can end in a citation.

Model training

Collects pages for training future models. Slower to pay off, and separate from search and citations.

By operator

Every crawler, grouped by who runs it

One company often runs three bots with three different jobs. Blocking the training crawler does not have to cost you the search one.
OpenAI
Search index
OAI-SearchBotOpenAI's search crawler. It is used to surface websites in the search features of ChatGPT.
Follows robots.txt
User-triggered fetch
ChatGPT-UserVisits a web page when a user asks ChatGPT or a custom GPT a question that needs it.
May ignore robots.txt on user requests
Model training
GPTBotCrawls content that may be used to train OpenAI's generative AI foundation models.
Follows robots.txt
Anthropic
Search index
Claude-SearchBotNavigates the web to improve the quality of search results for Claude users.
Follows robots.txt
User-triggered fetch
Claude-UserVisits a website when a person asks Claude a question that needs it.
Follows robots.txt
Model training
ClaudeBotCollects web content that could contribute to training Anthropic's models.
Follows robots.txt
Google
Model training
Google-ExtendedLets a site choose whether content Google crawls may be used to train future Gemini models, and for grounding answers in Gemini Apps and in Grounding with Google Search on Vertex AI.
A robots.txt token, not a crawler
Search index
GooglebotGoogle's main crawler. It feeds Google Search, including all Search features, and other Google products.
Follows robots.txt
Perplexity
Search index
PerplexityBotSurfaces and links websites in Perplexity's search results. Perplexity says it is not used to crawl content for AI foundation models.
Follows robots.txt
User-triggered fetch
Perplexity-UserVisits a web page when a user asks Perplexity a question, and may link that page in the answer. Perplexity says it is not used for crawling or training.
May ignore robots.txt on user requests
Microsoft
Search index
bingbotBing's web crawler. Microsoft says Copilot and Copilot Chat may fetch information from the Bing search service to ground a response.
Follows robots.txt
Apple
Search index
ApplebotPowers search features across Apple's products, including Spotlight, Siri and Safari. Apple says the data may also be used to help train its foundation models.
Follows robots.txt
Model training
Applebot-ExtendedLets a publisher opt out of having its content used to train Apple's generative foundation models.
A robots.txt token, not a crawler
Meta
Model training
meta-externalagentCrawls the web for uses such as training foundation AI models or improving products by indexing content directly.
Follows robots.txt
Search index
meta-webindexerNavigates the web to improve Meta AI search results. Meta says allowing it helps Meta AI cite and link your content in its responses.
Follows robots.txt
User-triggered fetch
meta-externalfetcherFetches individual links at a user's request, supporting product functions such as agentic AI features.
May ignore robots.txt on user requests
Amazon
Model training
AmazonbotUsed to improve Amazon's products and services. Amazon says the data may be used to train Amazon AI models.
Follows robots.txt
Search index
Amzn-SearchBotUsed to improve search experiences in Amazon's products and services. Amazon says it does not crawl content for generative AI model training.
Follows robots.txt
User-triggered fetch
Amzn-UserSupports user actions, such as answering Alexa questions that need up-to-date information. Amazon says it does not crawl content for generative AI model training.
May ignore robots.txt on user requests
DuckDuckGo
User-triggered fetch
DuckAssistBotCrawls pages in real time for DuckDuckGo's AI-assisted answers, which cite their sources. DuckDuckGo says the data is not used to train AI models.
Follows robots.txt
Mistral
User-triggered fetch
MistralAI-UserVisits a web page when a user asks Mistral's assistant a question, and may link the source in the answer. Mistral says it is not used for automatic crawling or for training.
Follows robots.txt
Search index
MistralAI-IndexCrawls the web automatically for indexing only. Mistral says it is not used for generative AI training.
Follows robots.txt
Model training
MistralAI-TrainingCrawls web content to build datasets for training Mistral's generative AI models.
Follows robots.txt
Allen Institute for AI
Model training
AI2BotCrawls web content that the Allen Institute for AI uses to train open language models.
Vendor does not say

Named in our robots.txt, undocumented

Our own robots.txt also names Claude-Web, anthropic-ai, Bytespider, YouBot, cohere-ai, GrokBot and xAI-Grok. We could not find current documentation for these tokens from the companies they are attributed to, so they have no page here: a page would have to guess what they do.

Check your own file

See which of these your robots.txt lets in.

The free AI crawler checker reads your robots.txt and reports which AI crawlers it allows or blocks. No signup.
Check AI crawler accessOr start free →