BotWitness, a mechanical mite with an orange lens BotWitness

FAQ

Questions we get asked

Can you prove my work was trained on?

No one honestly can, and we say so. What we report is tiered: verbatim reproduction and recognition of your exact wording are facts about a model's behavior; membership statistics are a clearly-labelled relative signal, never a probability; and a timestamped copyright canary that a future model reproduces is the one prospectively sound method. Tools that print "87% trained on your book" are selling a number no expert will defend.

What's an opt-out efficacy report?

The two-sided document only this archive can produce: our timestamped record proves when your domain blocked each AI crawler, and our probes test what models released after that date still reproduce or recognize of your text. It measures outcomes honestly: signal in a newer model doesn't by itself prove post-opt-out crawling, and the report says so on the page.

How can AI companies use BotWitness as a safeguard?

If you operate an AI crawler, it's your independent proof of good-faith compliance. We record, daily, with trusted timestamps, across 1,000,000+ domains, exactly what each site's robots.txt and llms.txt told your bot. If a publisher later claims you ignored their rules, you can show what the site actually permitted on the date you crawled it, from a neutral third party, not your own logs.

We already log crawl permissions ourselves. Why a third party?

Self-collected logs are easy to dispute: you could have edited them. BotWitness is a neutral outside party with independent RFC-3161 timestamps and a public, consistent methodology. Corroboration you didn't produce yourself carries far more weight in an audit or a dispute.

What if a site changes its robots.txt after we've crawled it?

That's exactly what we guard against. Permissions change, and history can't be reconstructed after the fact. Because we capture daily, we hold the version that was live on the date you accessed the site, so a later tightening can't be used to claim you ignored rules that didn't exist yet.

Can it also prove a site blocked a crawler?

Yes. The same record cuts both ways. A publisher can prove a site told a specific AI crawler "no" on a given date; an AI company can prove it was permitted. It's neutral evidence of what the file said, whichever side you're on.

How is this admissible?

Each capture stores the exact bytes, a SHA-256, and an independent RFC-3161 trusted timestamp, with a documented, consistent methodology. It attests to what a public file returned at a stated time; it is supporting evidence, not legal advice.

Why can't I just use the Wayback Machine?

It captures most sites only occasionally, skips the long tail, and keeps no chain of custody: weak as evidence. BotWitness records daily with tamper-evident timestamps.

Still have a question?

Billing and purchase questions are answered on the pricing page. For anything else, email [email protected], or [email protected] for evidence packages, defensive audits and Enterprise.