BotWitness, a mechanical mite with an orange lens BotWitness

For publishers · Crawler Compliance Audit

Did AI crawlers respect your robots.txt?

Your bot-detection tool can tell you a crawler visited. Only an independent, timestamped archive can tell you what your robots.txt said at that exact moment. Paste your server log; we evaluate every AI-crawler hit against the rules in effect at each hit's time and deliver the finding as a citable, branded PDF: "between date X and date Y, crawler A made N requests to paths that were disallowed at the time of access."

Why this needs a witness

Your logs prove the visit but never the rule

robots.txt shows only its current state. By the time a dispute matters, the file has changed and the past is gone. BotWitness has been independently capturing, SHA-256-hashing and RFC-3161-timestamping robots.txt across the web on a recurring schedule. The audit joins the two halves: your log supplies the hits; our archive supplies what the rules were, provably, at each moment. Neither half alone makes the case.

How it works

How an audit runs

Step 1

Paste your server log

Apache/nginx common or combined format, straight from your access log, or explicit {time, path, agent} rows via the API. We recognize the major AI crawlers (GPTBot, ClaudeBot, CCBot, PerplexityBot, Bytespider and more) by user-agent.

Step 2

Each hit meets the archived rule

For every hit we find the governing capture (the most recent archived robots.txt at or before that moment) and evaluate the path under RFC 9309 semantics, exactly as the major crawlers document: longest match wins, Allow wins ties, wildcards honored.

Step 3

A citable finding

The PDF reports the headline counts, the per-crawler breakdown, every disallowed hit as an exhibit row, and the exact snapshots (SHA-256 + timestamp status) that governed each verdict. Any snapshot can be escalated to a signed Certificate of Capture or a full litigation evidence package.

Custody

Your logs are never stored

Hits are evaluated and returned in the same request: nothing is retained on our side. You keep custody of your evidence; we only lend the archive. That separation is what keeps the audit neutral.

Who it's for

When you have to show that a crawler ignored you

Publishers & site owners

Turn a suspicion in your access logs into a dated, quantified finding: the difference between a complaint and a case file.

Counsel

The audit cites the governing snapshot for every verdict, each verifiable by hash and timestamp, and each escalatable to a certificate or evidence package for filing.

Licensing teams

Documented disallowed crawling is leverage in a content-licensing negotiation: here is what you took, here is what the rules said, dated.

Read this first: what the audit does and doesn't say

  • The audit evaluates the log you supply against the rules we independently archived. We did not observe the traffic ourselves and express no opinion on the log's authenticity.
  • A user-agent string identifies who a request claimed to be; operators can be impersonated. Corroborate with IP verification where it matters.
  • robots.txt is a voluntary protocol. "Disallowed at the time of access" is a factual finding about the published rules, not legal advice or a legal conclusion.
  • Hits before our first capture of a domain are reported honestly as "no archive coverage".

Get started

One audit, one PDF, back tonight

Self-serve audits are $199: pay, then run from your dashboard against up to 1,000 hits. On a Monitoring plan, audits are included and unlimited. For recurring audits across a portfolio, or expert-witness support, talk to us.

Suspect the crawling ended up in a model? The Exposure Report probes the models directly, and the Opt-Out Efficacy Report joins your dated opt-out record with those probes.