Research
TART Raises a State-of-the-Art Multi-Agent System's Multilingual GAIA Accuracy 5.6 Points Across Eleven Languages
Multilingual multi-agent systems degrade sharply outside English, and this work locates the loss at the planner — the request-to-action interface — deriving a taxonomy of planning-grounding failures from real failed executions. LLM-based analysis shows these failures make up an increasing share of unsuccessful runs as language-resource availability drops, worst in low-resource languages. TART makes the taxonomy's key aspects explicit to the planner and downstream sub-agents, improving results across three LLM backbones, two datasets, and two agentic configurations, with a 5.6-point average gain on multilingual GAIA.
↳ Follow the thread