Legal / Retractions

2026-07-10 GPT-3 verdict correction

A production audit found that a GPT-2 checkpoint measurement had been attached to the GPT-3 paper Language Models are Few-Shot Learners. The resulting REPRODUCED label was not a valid finding about that paper and has been retracted.

Current state: arXiv:2005.14165 has no current POST verdict. The GPT-2 measurement is excluded from every public verdict page, API, badge, feed, leaderboard, statistic, and metadata surface.

What happened

The row used agent version v0.1.0-gpt2-perplexity-microslice, but its paper identifier was arXiv:2005.14165, the GPT-3 paper. The measurement and paper identity therefore described different models. A successful GPT-2 run cannot substantiate any sentence or result in the GPT-3 paper.

Correction

The reproduction measured the 124M-parameter GPT-2 checkpoint on WikiText-103 but attached that result to the GPT-3 paper. The GPT-3 paper does not report that checkpoint result as its claim, so the stored row is excluded from every public verdict surface.

The underlying record is retained for audit. A central public exposure policy now rejects this exact paper and agent pairing, and also rejects every verdict still inside the 72-hour pre-publication-notice window.

Prevention

Public verdict readers now share one visibility rule instead of implementing legal holds and known-invalid identities independently. A source-level regression test covers paper pages, APIs, badges, feeds, claims, anchors, timelines, leaderboards, lab and author pages, validator counts, and homepage statistics.

Records