HERO'S JOURNEY: Benchmark for Complex Rule Induction in Goal-Directed Agent Tasks
arXiv·medium signal
Introduces HERO'S JOURNEY, a benchmark where agents must infer hidden rules from demonstrations and execute multi-step plans in text games. Covers eight tasks across attribute and relational reasoning with episodic structure requiring both rule induction and execution. Provides a concrete evaluation framework for measuring whether LLM agents can learn rules from context rather than being explicitly programmed — a key capability gap in current agent systems.