BotWitness, a mechanical mite with an orange lens BotWitness

For eval vendors & researchers · Benchmark Contamination Report

Is your benchmark in the training data?

A model that saw your eval items during training scores on memory, not ability, and every leaderboard built on those numbers quietly overstates it. We probe the major models for familiarity with your items: can they continue one verbatim? Do they pick the exact original wording over close paraphrases at rates far above chance? Honest, tiered evidence: never an accusation dressed as a number.

The problem

Contamination is common and usually nobody's fault

Benchmarks leak into training corpora through arXiv papers, GitHub mirrors, and scraped discussion of the questions: no intent required. That's exactly why the finding has to be framed honestly: familiarity is a fact you can measure; how the items got there usually isn't. Our report gives you the measurement with the caveats printed on the same page, so you can act on it (rotate items, hold out a private split, annotate results) without overclaiming.

How it works

The same probes, pointed at your eval set

Fact

Verbatim reproduction

Each model gets a short prefix of an eval item and is asked to continue it. A model that completes your held-out question word-for-word has seen it. Refusals are recorded as facts.

Fact

Recognition (DE-COP)

The model sees your verbatim item beside three machine paraphrases and picks the original. Above-chance accuracy is a familiarity signal that works on guarded, closed models: no logprobs needed.

Relative signal

Membership statistics

On open-weight models with token probabilities, a Min-K%++ statistic positions your items between known-trained and known-untrained reference text: a position, never a probability.

Deliverable: machine-readable JSON plus a branded, hash-audited PDF with per-model results and the methods explained in plain language. Runs self-serve from your dashboard against the configured flagship roster.

Who it's for

Who this is for

Eval vendors

Show customers your benchmark is monitored for contamination, and know when it's time to rotate items or cut a private split.

Researchers

A familiarity check across models before you publish results, with a reproducible, hash-audited method you can cite instead of a home-rolled script.

Enterprise buyers

Vendor claims 95% on a public benchmark? Probe whether the model has simply seen the test. Due diligence, honestly framed.

Read this first: limitations

  • Familiarity signal is an investigative finding, not an accusation. Benchmark text leaks into corpora through public channels without intent.
  • Above-chance recognition evidences familiarity with the exact wording; it does not by itself establish when or how the items entered training.
  • A negative result does not certify a clean benchmark: items can be memorized in forms our probes don't reach.
  • Submit only items you own or are authorized to submit; results are not legal advice.

Get started

Probe the models before you trust the leaderboard

Self-serve reports are $299: pay, then run from your dashboard (tick "contamination mode" on the report card). For recurring sweeps across releases, pair with Exposure Monitoring.