Research
HarnessEval-W Replaces Fixed World-Model Rubrics With a Sub-Agent Tree That Shows Its Reasoning, Across 18 Models and 330 Cases
The most-upvoted paper on HuggingFace Daily Papers today (104 upvotes) argues a benchmark score without a reasoning chain is untrustworthy for world models, where judging a rollout means checking whether physics, causality and world state evolve correctly. HarnessEval-W interprets each evaluation case, decomposes it into measurable subproblems, and spawns specialized sub-agents with tailored diagnostic tools; the parent agent validates the evidence and produces a verdict backed by a full evidence tree. Applied to 18 world models over 330 cases, judgments align closely with human preference, and the pipeline is open-sourced as a live benchmark others can extend with new skills.
↳ Follow the thread