Research
Reflexion-Style Verbal Memory Sometimes Lowers Success Versus Plain Retry, and Replay Experiments Show Why
VRL-Bench evaluates trial-and-error learning under finite trial budgets across three models on MiniWoB and WebShop, and every prominent verbal-memory method from Reflexion onward improves observed success over memory-free retry in some settings while reducing it in others. Replay experiments show reflection itself can lower success rates, exposing a trade-off between exploiting written-down experience and continuing to explore. Their VEX-squared scheduler uses a language model to jointly pick policies and allocate the remaining trial budget, and is the only evaluated update with positive gains over retry in all six settings.
↳ Follow the thread