For rights holders · Training-Data Exposure Report
Submit your text; we probe the major AI models and report what they reveal, how much of your work a model reproduces verbatim, and where it sits against known-trained and known-untrained reference text. Honest, tiered evidence you can act on: never a made-up percentage.
What this is, and isn't
The peer-reviewed consensus is blunt: for a closed production model you cannot soundly compute the probability that specific text was in its training set, you'd have to know the whole training corpus or retrain the model. Any tool that prints a confident percentage is selling you a number it can't stand behind. We built the opposite: we report facts where facts exist, a clearly-labelled relative signal where only a signal exists, and (going forward) the one method that is genuinely sound.
How it works
Tier A: fact
We give each model a short prefix of a distinctive passage and measure how much of the true continuation it reproduces word-for-word. "The model output a 240-word verbatim run of your Chapter 3 from a 15-word prompt" is a fact about the model, and the soundest signal there is. Guarded flagship models often refuse this (recorded as a fact), which is why we pair it with recognition.
Tier A: fact
We show the model the verbatim passage beside three close paraphrases and ask which is the original. A model that memorized your text picks the exact wording well above the 25% chance line. It's a comprehension question, not a reproduction request, so the guarded flagships that refuse the test above answer this one, and it needs no special access.
Tier B: relative signal
For open-weight models we compute a membership statistic over token probabilities and report where your work falls between known-trained and known-untrained reference text, a position, never an absolute probability. Closed models don't expose the needed signal, and we say so rather than fake it.
Tier C: sound, prospective
We generate a unique random canary, publish and RFC-3161 timestamp it on BotWitness's infrastructure, then re-probe models over time. Because the canary is random and provably published on a date, a later reproduction is genuinely sound evidence of training. The one method with a defensible false-positive rate.
See what you get
This is an actual BotWitness Exposure Report, run on a public-domain work, the opening of Jane Austen's Pride and Prejudice, so you can see the exact format. Your report runs on your text, across the same models.
Exposure indicator
Strong verbatim reproduction on at least one model
An investigative signal, not proof that the work was used in training.
Tier A: verbatim reproduction, by model
| Model | Strong | Partial | Refused |
|---|---|---|---|
| GPT-4o | 0 | 0 | 3 |
| Claude Sonnet 4.5 | 0 | 2 | 1 |
| Gemini Pro | 0 | 1 | 0 |
| Llama 3.3 70B (open-weight) | 1 | 1 | 0 |
3 passages probed per model. A refusal is recorded as a fact, not a reproduction: GPT-4o declined every reproduction probe (its guardrails), while Gemini Pro and Llama reproduced the text. But watch what happens when we stop asking them to reproduce it…
Tier A: recognition (DE-COP), by model
Each model was shown the verbatim passage beside three close paraphrases and asked which is the original, a comprehension question, not a reproduction request, so the guarded models answer instead of refusing:
| Model | Correct | Accuracy | vs 25% chance |
|---|---|---|---|
| GPT-4o (refused above) | 3/3 | 100% | +75 pts |
| Claude Sonnet 4.5 | 3/3 | 100% | +75 pts |
| Gemini Pro | 3/3 | 100% | +75 pts |
| Llama 3.3 70B (open-weight) | 3/3 | 100% | +75 pts |
Every model (including the two that refused to reproduce the text) identified the verbatim original 3 of 3 times, versus the 25% you would expect from guessing. That gap is a strong familiarity signal (an investigative signal, not proof). Paraphrases generated by GPT-4o.
Reproduced-text exhibit: Llama 3.3 70B
A 55-word verbatim run of the true continuation, produced from a 15-word prompt:
Tier B: membership statistics
On an open-weight model that exposes token probabilities (Qwen 3.5 9B), 2 of 3 passages sat closer to the known-trained pole than to freshly-random text (passage surprise −1.22 vs −6.35 for text no model has seen), reported only as a relative position, never an absolute probability.
Every model transcript is SHA-256 hashed and retained, so any result is reproducible and auditable. The full report also carries a plain-language explanation of each method and its limits. This report is an investigative signal, not proof, and not legal advice.
Who it's for
See whether frontier models can recite your catalogue, with the reproduced passages as exhibits, the starting point for a licensing conversation or a claim.
A neutral, reproducible, hash-audited report with the science stated plainly (including its limits) is more useful in a filing than a black-box score no expert will defend.
Run the same probes across a body of works and enroll canary traps now, so next year's models are testable against a timestamped baseline.
Read this first: limitations
New · Exposure Monitoring · $99/mo
A one-shot report answers "what do today's models know?" Monitoring keeps answering it: enroll up to three works and we re-probe automatically when the model roster changes: a new flagship ships, or monthly, and email you the day results change. Your enrolled copyright canaries are re-probed on the same trigger, so a future reproduction (the one genuinely sound signal) reaches you the day it appears, with the timestamped publication proof attached.
Get started
Self-serve reports are $299 each: pay, then run from your dashboard. Ongoing coverage is $99/mo with Exposure Monitoring. For a body of works, expert context, or a filing-ready package, we deliver per matter.
Already blocked the crawlers? The Opt-Out Efficacy Report joins this probe suite with your dated opt-out record, and a Compliance Audit checks whether the crawlers respected your rules at all.