BotWitness, a mechanical mite with an orange lens BotWitness

For rights holders · Training-Data Exposure Report

Was your work used to train AI?

Submit your text; we probe the major AI models and report what they reveal, how much of your work a model reproduces verbatim, and where it sits against known-trained and known-untrained reference text. Honest, tiered evidence you can act on: never a made-up percentage.

What this is, and isn't

No one can honestly hand you "87% trained on your book."

The peer-reviewed consensus is blunt: for a closed production model you cannot soundly compute the probability that specific text was in its training set, you'd have to know the whole training corpus or retrain the model. Any tool that prints a confident percentage is selling you a number it can't stand behind. We built the opposite: we report facts where facts exist, a clearly-labelled relative signal where only a signal exists, and (going forward) the one method that is genuinely sound.

How it works

Four kinds of evidence, each labelled

Tier A: fact

Verbatim reproduction

We give each model a short prefix of a distinctive passage and measure how much of the true continuation it reproduces word-for-word. "The model output a 240-word verbatim run of your Chapter 3 from a 15-word prompt" is a fact about the model, and the soundest signal there is. Guarded flagship models often refuse this (recorded as a fact), which is why we pair it with recognition.

Tier A: fact

Recognition (DE-COP)

We show the model the verbatim passage beside three close paraphrases and ask which is the original. A model that memorized your text picks the exact wording well above the 25% chance line. It's a comprehension question, not a reproduction request, so the guarded flagships that refuse the test above answer this one, and it needs no special access.

Tier B: relative signal

Membership statistics

For open-weight models we compute a membership statistic over token probabilities and report where your work falls between known-trained and known-untrained reference text, a position, never an absolute probability. Closed models don't expose the needed signal, and we say so rather than fake it.

Tier C: sound, prospective

Copyright traps

We generate a unique random canary, publish and RFC-3161 timestamp it on BotWitness's infrastructure, then re-probe models over time. Because the canary is random and provably published on a date, a later reproduction is genuinely sound evidence of training. The one method with a defensible false-positive rate.

See what you get

What a finished report contains

This is an actual BotWitness Exposure Report, run on a public-domain work, the opening of Jane Austen's Pride and Prejudice, so you can see the exact format. Your report runs on your text, across the same models.

Exposure indicator

Strong verbatim reproduction on at least one model

An investigative signal, not proof that the work was used in training.

Tier A: verbatim reproduction, by model

Model Strong Partial Refused
GPT-4o 0 0 3
Claude Sonnet 4.5 0 2 1
Gemini Pro 0 1 0
Llama 3.3 70B (open-weight) 1 1 0

3 passages probed per model. A refusal is recorded as a fact, not a reproduction: GPT-4o declined every reproduction probe (its guardrails), while Gemini Pro and Llama reproduced the text. But watch what happens when we stop asking them to reproduce it…

Tier A: recognition (DE-COP), by model

Each model was shown the verbatim passage beside three close paraphrases and asked which is the original, a comprehension question, not a reproduction request, so the guarded models answer instead of refusing:

Model Correct Accuracy vs 25% chance
GPT-4o (refused above) 3/3 100% +75 pts
Claude Sonnet 4.5 3/3 100% +75 pts
Gemini Pro 3/3 100% +75 pts
Llama 3.3 70B (open-weight) 3/3 100% +75 pts

Every model (including the two that refused to reproduce the text) identified the verbatim original 3 of 3 times, versus the 25% you would expect from guessing. That gap is a strong familiarity signal (an investigative signal, not proof). Paraphrases generated by GPT-4o.

Reproduced-text exhibit: Llama 3.3 70B

A 55-word verbatim run of the true continuation, produced from a 15-word prompt:

fortune, must be in want of a wife. However little known the feelings or views of such a man may be on his first entering a neighbourhood, this truth is so well fixed in the minds of the surrounding families, that he is considered the rightful property of some one or other of their daughters.

Tier B: membership statistics

On an open-weight model that exposes token probabilities (Qwen 3.5 9B), 2 of 3 passages sat closer to the known-trained pole than to freshly-random text (passage surprise −1.22 vs −6.35 for text no model has seen), reported only as a relative position, never an absolute probability.

Every model transcript is SHA-256 hashed and retained, so any result is reproducible and auditable. The full report also carries a plain-language explanation of each method and its limits. This report is an investigative signal, not proof, and not legal advice.

Download the full sample report (PDF) Run one on your work · $299

Who it's for

When you have to show your working

Authors & publishers

See whether frontier models can recite your catalogue, with the reproduced passages as exhibits, the starting point for a licensing conversation or a claim.

Counsel

A neutral, reproducible, hash-audited report with the science stated plainly (including its limits) is more useful in a filing than a black-box score no expert will defend.

Rights organizations

Run the same probes across a body of works and enroll canary traps now, so next year's models are testable against a timestamped baseline.

Read this first: limitations

  • This report is an investigative signal, not proof, and not legal advice.
  • Verbatim reproduction shows a model can output your text, a fact about the model. It does not by itself establish how your work entered the model, or that any right was infringed.
  • Membership statistics are a relative position against reference corpora, not a probability that your work was trained on. No such probability is soundly computable for a production model.
  • A negative or null result does not prove your work was excluded.
  • You may only submit text you own or are authorized to submit; self-serve runs require you to acknowledge this.

New · Exposure Monitoring · $99/mo

New models ship monthly, so the report re-runs itself

A one-shot report answers "what do today's models know?" Monitoring keeps answering it: enroll up to three works and we re-probe automatically when the model roster changes: a new flagship ships, or monthly, and email you the day results change. Your enrolled copyright canaries are re-probed on the same trigger, so a future reproduction (the one genuinely sound signal) reaches you the day it appears, with the timestamped publication proof attached.

Get started

Run one now, or plant a canary for next year

Self-serve reports are $299 each: pay, then run from your dashboard. Ongoing coverage is $99/mo with Exposure Monitoring. For a body of works, expert context, or a filing-ready package, we deliver per matter.

Already blocked the crawlers? The Opt-Out Efficacy Report joins this probe suite with your dated opt-out record, and a Compliance Audit checks whether the crawlers respected your rules at all.