For eval vendors & researchers · Benchmark Contamination Report
A model that saw your eval items during training scores on memory, not ability, and every leaderboard built on those numbers quietly overstates it. We probe the major models for familiarity with your items: can they continue one verbatim? Do they pick the exact original wording over close paraphrases at rates far above chance? Honest, tiered evidence: never an accusation dressed as a number.
The problem
Benchmarks leak into training corpora through arXiv papers, GitHub mirrors, and scraped discussion of the questions: no intent required. That's exactly why the finding has to be framed honestly: familiarity is a fact you can measure; how the items got there usually isn't. Our report gives you the measurement with the caveats printed on the same page, so you can act on it (rotate items, hold out a private split, annotate results) without overclaiming.
How it works
Fact
Each model gets a short prefix of an eval item and is asked to continue it. A model that completes your held-out question word-for-word has seen it. Refusals are recorded as facts.
Fact
The model sees your verbatim item beside three machine paraphrases and picks the original. Above-chance accuracy is a familiarity signal that works on guarded, closed models: no logprobs needed.
Relative signal
On open-weight models with token probabilities, a Min-K%++ statistic positions your items between known-trained and known-untrained reference text: a position, never a probability.
Deliverable: machine-readable JSON plus a branded, hash-audited PDF with per-model results and the methods explained in plain language. Runs self-serve from your dashboard against the configured flagship roster.
Who it's for
Show customers your benchmark is monitored for contamination, and know when it's time to rotate items or cut a private split.
A familiarity check across models before you publish results, with a reproducible, hash-audited method you can cite instead of a home-rolled script.
Vendor claims 95% on a public benchmark? Probe whether the model has simply seen the test. Due diligence, honestly framed.
Read this first: limitations
Get started
Self-serve reports are $299: pay, then run from your dashboard (tick "contamination mode" on the report card). For recurring sweeps across releases, pair with Exposure Monitoring.