Taste-Bench: Best Frontier Model Picks the Better Direction at Only 59.7% of Long-Horizon Agent Decision Forks
arXiv·medium signal
Taste-Bench mines decision forks automatically from parallel agent attempts and from detours inside single engineering and research trajectories. At each fork one direction led to a better outcome, and the model must choose without seeing what came after. The best frontier model answered only 59.7% correctly, which isolates judgment on which hypothesis or implementation to pursue as a weak spot that end-to-end success metrics hide.