Held-out set
Plain English. Test questions deliberately withheld from a model's training data, so a score measures generalisation rather than memory. Variants matter: a public set is visible to everyone, a private set only to the benchmark operator, a semi-private set sits in between — released to labs under conditions.
Why it moves money. A test the model has seen converts a capability claim into a memory claim, and capital is priced on the former. As public sets decay, operators retreat toward privacy: Artificial Analysis took 40 per cent of its Intelligence Index private in September 2026, and OpenAI retired SWE-bench Verified after its own pipeline found contamination. The held-out set is the thin wall between a benchmark and an advertisement.
What to watch. Whether privacy actually saves the instrument. The best formal study of saturation found that expert curation, not keeping test data private, is what predicts a benchmark's resilience — a leaked easy question was always going to saturate; a genuinely hard one survives exposure longer.
From the signals. Artificial Analysis retires a saturated benchmark and takes 40% of the index private. OpenAI retires SWE-bench Verified on measured contamination. The saturation study: curation beats secrecy.
Further reading. ARC-AGI — the public / semi-private / private set structure, explained by its operator.