Type one in below. You will get what its robots.txt says about each AI crawler today, the SHA-256 of the capture, and the date we first recorded it. Nothing to sign up for.
On 14 August 2026 our crawler asked nytimes.com for its robots.txt and kept every byte that came back. Line 214 of the 344 it returned said this about OpenAI's crawler.
nytimes.com/robots.txt, lines 213 to 216
User-agent: GPTBot Disallow: / User-agent: ImagesiftBot
The record of that capture
In the last thirty days, 98,721 tracked domains rewrote their robots.txt. Each of those edits carries the same hash, the same timestamp and the same nightly digest as the capture above. That is the part nobody can go back and collect afterwards.
A Certificate of Capture turns one archived fetch into a dated document you can hand to somebody else. This is the certificate for the record above, produced by the same code that runs when a customer buys one, from the same row in the archive.
It names the domain and the crawler, states what the file said on that date, and prints the SHA-256, the capture time and the timestamp status underneath. The captured content follows on page two.
Download it and check the hash yourself against the record above. If the two disagree, we have a problem and you should not buy anything from us.
The archive proves what a site said. The probes test what the models kept. Holding both is the only reason we can answer the question anybody actually asks.
Daily captures of robots.txt and llms.txt across more than a million domains, plus EU disclosure pages, MCP manifests, agent key directories and model licences. None of it can be reconstructed later. Somebody has to write it down on the day, and that somebody should have nothing riding on the answer.
Four methods, each labelled by what it can and cannot show. Verbatim reproduction and recognition of your exact wording are facts about how a model behaves. A membership statistic is a relative signal, and we will not dress it up as a probability, because no honest one exists. When a model refuses to answer, the refusal goes in the report too.
Dated, hash-verifiable answers about what the rules said at a given moment, and honest tests of what the models retained. Whichever side of the argument you are on.
A signed exhibit stating what a site told a crawler on a date, with the hash, the timestamp status and the captured content.
Watched domains, signed change alerts, unlimited history, and every report on this page unmetered.
Send your server log. We judge every AI-crawler hit against the robots.txt that was live at that exact moment. We never keep the log.
For AI companies. What robots.txt permitted your crawler, across your crawl list, over any range. Drawn from a record you do not control.
You opted out on a date. We prove when, then test the models released since and show what they still reproduce of your text.
Reproduction and recognition probes across the flagship models, plus a membership signal we label as relative and leave that way.
For litigators. Raw captures, timestamp tokens, Merkle proofs and a verifier, so the other side can check it without trusting us.
For AI labs. Find out what your model reproduces before opposing counsel does it for you.
The same recorder archives what the rest of the AI industry publishes and then quietly edits. Looking is free.
| Stream | What we keep | Go |
|---|---|---|
| EU AI Act | What each general-purpose AI provider disclosed about its training content under Article 53(1)(d), revision by revision, including the days it published nothing at all. Enforcement started on 2 August 2026. | Disclosures |
| MCP | Daily sweeps of the registry and live tool manifests from remote servers. What a tool claimed to be on the day an agent trusted it. | Manifests |
| Web Bot Auth | Daily snapshots of the signing-key directories run by the major agent operators. Keys rotate, and yesterday disappears unless somebody kept it. | Keys |
| Hugging Face | Whether a model was public, under which licence, on which date. Deletions, gatings and licence edits become dated events. | Models |
| RSL | Really Simple Licensing terms declared in robots.txt across the tracked web, captured as served from the day a site first declares one. | Adoption |
Lookups, the live report and the streams cost nothing. Everything else sits on one page with the price next to it.
One-time reports run from $49 to $499. You pay once and keep the PDF. Monitoring starts at $49 a month and unmeters every report on this page. Evidence packages, defensive audits and enterprise access are quoted per matter, because the scope is different every time.