For publishers · Crawler Compliance Audit
Your bot-detection tool can tell you a crawler visited. Only an independent, timestamped archive can tell you what your robots.txt said at that exact moment. Paste your server log; we evaluate every AI-crawler hit against the rules in effect at each hit's time and deliver the finding as a citable, branded PDF: "between date X and date Y, crawler A made N requests to paths that were disallowed at the time of access."
Why this needs a witness
robots.txt shows only its current state. By the time a dispute matters, the file has changed and the past is gone. BotWitness has been independently capturing, SHA-256-hashing and RFC-3161-timestamping robots.txt across the web on a recurring schedule. The audit joins the two halves: your log supplies the hits; our archive supplies what the rules were, provably, at each moment. Neither half alone makes the case.
How it works
Step 1
Apache/nginx common or combined format, straight from your access log, or explicit {time, path, agent} rows via the API. We recognize the major AI crawlers (GPTBot, ClaudeBot, CCBot, PerplexityBot, Bytespider and more) by user-agent.
Step 2
For every hit we find the governing capture (the most recent archived robots.txt at or before that moment) and evaluate the path under RFC 9309 semantics, exactly as the major crawlers document: longest match wins, Allow wins ties, wildcards honored.
Step 3
The PDF reports the headline counts, the per-crawler breakdown, every disallowed hit as an exhibit row, and the exact snapshots (SHA-256 + timestamp status) that governed each verdict. Any snapshot can be escalated to a signed Certificate of Capture or a full litigation evidence package.
Custody
Hits are evaluated and returned in the same request: nothing is retained on our side. You keep custody of your evidence; we only lend the archive. That separation is what keeps the audit neutral.
Who it's for
Turn a suspicion in your access logs into a dated, quantified finding: the difference between a complaint and a case file.
The audit cites the governing snapshot for every verdict, each verifiable by hash and timestamp, and each escalatable to a certificate or evidence package for filing.
Documented disallowed crawling is leverage in a content-licensing negotiation: here is what you took, here is what the rules said, dated.
Read this first: what the audit does and doesn't say
Get started
Self-serve audits are $199: pay, then run from your dashboard against up to 1,000 hits. On a Monitoring plan, audits are included and unlimited. For recurring audits across a portfolio, or expert-witness support, talk to us.
Suspect the crawling ended up in a model? The Exposure Report probes the models directly, and the Opt-Out Efficacy Report joins your dated opt-out record with those probes.