Preview environment
CrawlPact

AI crawler directory

Reference pages for AI crawlers CrawlPact's registry tracks, each with its purpose, official source, and last-verified date. This is a growing, versioned registry — see methodology for how entries are verified.

22 crawlers from 8 operators · most recently verified 7/30/2026

How entries are verified

Each entry requires a reliable official source, a verified user-agent token, a purpose classification, and a verification date before publication. Entries are re-checked against each operator's own current documentation on a rolling basis, not left to go stale indefinitely — see the last-verified date on each crawler's page. Where an official source could not be reliably fetched for a verified crawler, that is disclosed on its own registry record rather than published as though automatically confirmed.

Search

  • Amzn-SearchBot

    Improves search experiences in Amazon products and services; Amazon's own documentation states it is not used for generative AI model training.

  • Claude-SearchBot

    Navigates the web to improve the relevance and accuracy of Claude's search results.

  • Googlebot

    Google's primary web crawler for Search indexing — not an AI-training-specific crawler.

  • Meta-WebIndexer

    Navigates the web to improve Meta AI search result quality.

  • OAI-SearchBot

    Used to discover and surface links to websites in ChatGPT search results.

  • PerplexityBot

    Indexes web content to power Perplexity's AI-generated search answers.

Training

  • Applebot-Extended

    Controls use of website content for training Apple Intelligence and other Apple generative AI models.

  • ClaudeBot

    Used by Anthropic to crawl publicly accessible content for model training.

  • Google-Extended

    Controls use of website content for training Gemini and Vertex AI generative models, independent of Search indexing.

  • GPTBot

    Used by OpenAI to crawl publicly accessible web content that may be used to train future models.

  • Meta-ExternalAgent

    Used by Meta to crawl content for training AI models and improving AI products.

User-triggered

  • Amzn-User

    Fetches a page on behalf of an end user or Amazon application, such as responding to an Alexa query that needs up-to-date information; not used for generative AI model training.

  • ChatGPT-User

    Fetches a web page in direct response to a user's question inside ChatGPT or a Custom GPT.

  • Claude-User

    Fetches a web page when a person directs Claude to access it as part of a query.

  • Perplexity-User

    Fetches a page in direct response to a user's question inside Perplexity.

Agent / action

  • Google-CloudVertexBot

    Crawls requested by site owners for building Vertex AI Agents.

  • Meta-ExternalFetcher

    Fetches individual links at a user's request to support agentic AI capabilities in Meta products.

Advertising / validation

  • Meta-ExternalAds

    Crawls the web for use cases such as improving advertising and other business-related products.

  • OAI-AdsBot

    Validates the safety and relevance of web pages submitted as ads on ChatGPT — not used for AI training.

Research

  • CCBot

    Builds the open Common Crawl web corpus, which is reused by many third-party model trainers.

Mixed

  • Amazonbot

    Used by Amazon to improve its products and services, including Alexa answers, and may be used to train Amazon AI models.

Unknown

  • GoogleOther

    A generic Google crawler various internal product teams may use — Google's own documentation does not specify which teams or purposes.