Eval contamination
Plain English. Benchmark questions — or their answers — leaking into a model's training data, so the model has effectively seen the exam. Contamination is rarely deliberate; the test set was on the public web, and the web is the training corpus.
Why it moves money. A contaminated benchmark overstates capability precisely where capability is being sold. The canonical case: OpenAI retired SWE-bench Verified, the standard coding comparison, after its own contamination pipeline found cases of models reproducing near-verbatim solutions. Benchmark-led model rankings now carry a contamination caveat by default, which means every capability-based valuation argument inherits the caveat too.
What to watch. Whether labs publish contamination analyses alongside scores, and a corrective from the research: the best formal study of benchmark decay found contamination is not the main killer — most benchmarks die because their tasks were easy to exhaust, and expert curation predicts survival better than data secrecy. Contamination is real; it is not the whole measurement crisis.
From the signals. OpenAI retires SWE-bench Verified — the instrument crossed the contamination threshold. The saturation study: contamination is not the main driver. Mark's extension: human assessment has been contaminated since ChatGPT shipped.