Preview environment
CrawlPact

Audit and monitor your website's AI crawler policy.

CrawlPact independently audits robots.txt and related public signals, separates search from training and agent crawlers, explains conflicts, and monitors changes across any hosting provider.

No installation, server-log access, or AI API required. Try example.com.

  • Vendor-neutral
  • Evidence-based findings
  • Works with any hosting provider
  • Deterministic recommendations

One crawler-policy mistake can create the wrong outcome.

Search crawlers may be blocked unintentionally

A broad AI-crawler rule may also affect crawlers used for search or user-requested retrieval — not only model-training crawlers.

Training access may remain unspecified

A website may intend to restrict model-training crawlers while publishing no explicit rule for several verified crawler tokens. Unspecified is a neutral state, not a guarantee either way.

Deployments and CDN settings can change the public policy

The policy a team intends may differ from the public response after a deployment, framework update, or CDN-generated rewrite.

Not every website needs the same policy, and not all AI crawling is harmful — the point is knowing, with evidence, what your website currently declares.

See what a report looks like

Illustrative example using a synthetic demonstration domain — not a real scan result.

sample-domain.example (demonstration data)

AI crawler policy report

Scanned 01/08/2026 · Registry v2026.07.3

68/ 100Needs attention
  • Resource availability90
  • Syntax & evaluation85
  • Objective alignment45
  • Cross-signal consistency70
How this score is calculated

OAI-SearchBot

OpenAI · Search

Allowed

GPTBot

OpenAI · Training

No explicit rule

ClaudeBot

Anthropic · Training

Blocked

Google-Extended

Google · Training

Allowed

Example finding

GPTBot has no explicit rule

robots.txt, fetched 2026-08-01 — no User-agent group names GPTBot.

Recommended action: Add an explicit Disallow rule naming GPTBot's user-agent token.

How CrawlPact works

  1. 1

    Audit public policy signals

    CrawlPact requests publicly accessible resources such as robots.txt and related policy signals.

  2. 2

    Evaluate by crawler purpose

    Verified crawler tokens are grouped by search, training, user-triggered retrieval, or agent use.

  3. 3

    Explain findings with evidence

    The report identifies explicit rules, unspecified policies, conflicts, unavailable resources, and recommended next steps.

  4. 4

    Monitor changes

    Saved domains can be compared against future website responses and crawler-registry updates on paid plans.

Public policy auditing does not require server-log access or installation. Actual crawler traffic is outside the audit's proof boundary — see methodologyand crawler directory.

Not every AI crawler serves the same purpose.

Crawler operators increasingly separate purposes into distinct tokens. Blocking one purpose does not automatically affect another, and classifications are reviewed against official documentation rather than assumed.

Search

Used to support search or discovery experiences, indexing content to help answer queries.

Training

Used to collect content for model development or training. A separate decision from search, often using a different crawler token.

User-triggered retrieval

Used when a person explicitly asks a product to retrieve a specific page in real time.

Agents

Used by automated or semi-automated agent workflows performing a task, distinct from a one-shot retrieval.

Registry classifications may change when an operator updates its own documentation — see the AI crawler directory for every verified entry and its source.

Built for how your team actually works

The same audit and monitoring engine, with guidance shaped around what your role actually needs to know.

  • Agencies

    Agencies must understand, explain, and monitor AI crawler policy state across many client websites at once — one-off manual checks don't scale, and clients ask about it in language that doesn't map cleanly to a robots.txt file.

  • Publishers

    Search and AI-training crawlers often warrant different decisions, but a site's public signals frequently don't distinguish them clearly — and CDN or deployment changes can alter the actual response without an editorial decision being made.

  • SaaS and documentation teams

    Documentation discoverability decisions, the search-versus-training distinction, and deployment-platform-driven policy drift all intersect on the domains SaaS teams run — often across multiple subdomains with different owners.

  • Web developers

    The crawler policy a developer intends to ship and the crawler policy a framework, hosting platform, or CDN actually serves can diverge — and that gap is invisible until someone checks the live, deployed response.

AI crawler directory

View all crawlers

Govern crawler policy across every client website.

CrawlPact is built primarily for agencies and multi-site teams — save every client domain in one place, apply objective-based recommendations, and monitor for website or registry changes without exposing internal tooling to clients.

  1. 1Audit client domains
  2. 2Establish policy baselines
  3. 3Identify conflicts and unspecified rules
  4. 4Apply client-approved recommendations
  5. 5Monitor website and registry changes
  6. 6Share or export evidence

Portfolio monitoring

Group domains, batch-import a client list, and schedule recurring rechecks — plan limits apply, see pricing.

Client-ready sharing

Share a report with your own agency name and logo attached — CrawlPact's methodology and limitations always remain visible on a shared report, and branding never replaces them.

Independent evidence, not another hosting-provider control panel.

CrawlPact works across publicly accessible hosting setups, does not require a website to use one CDN, preserves the evidence behind every finding, uses a source-backed crawler registry, distinguishes website-policy changes from registry-driven changes, and publishes its own methodology and limitations.

CrawlPact audits the public policy signals a website publishes. It does not control external crawlers or guarantee that they will comply.

Supported public signals

Primary signals (robots.txt, HTTP headers, HTML meta directives) are checked on every scan. Optional or emerging signals are checked when publicly available and evaluated according to the current published methodology — none is required for AI visibility.

  • robots.txt
  • AI crawler user-agent groups
  • llms.txt
  • llms-full.txt
  • RSL
  • Content Signals
  • Robots meta
  • X-Robots-Tag
  • Sitemap declarations
  • Relevant HTTP headers

A correct policy today can change tomorrow.

Website-policy changes

  • robots.txt, header, or meta directive changed
  • A resource became unavailable
  • A CDN rewrote the public response
  • A deployment changed the effective policy

Registry-driven changes

  • A new crawler token was added
  • An operator changed its documented purpose
  • Source verification changed
  • Classification uncertainty changed

Monitoring frequency depends on your plan — see pricing. Changes surface through an in-app notification centre and a private Atom feed.

Monthly and yearly pricing. Billing is handled by Paddle; applicable taxes may be calculated during checkout.

Free

For a one-time check on a single site.

Scan a website

Solo

For individual site owners tracking one project.

Start monitoring

Pro

For growing portfolios that need weekly monitoring.

Choose Pro

Agency

For agencies managing many client domains.

Choose Agency

Frequently asked questions

See what your website currently tells AI crawlers.

Run a free public-policy audit with no installation or server-log access.

Results describe published policy signals and do not guarantee crawler behaviour.

View a sample report