ToMAS converts diagnosed multi-agent coordination failures into training items, then reports honestly that its own training run did nothing
ToMAS applies four explicit convertibility criteria to diagnosed execution traces, converting inter-agent misalignment cases (labelled FC2 in MAST-Data) into partner-state reasoning items; a full pass over 242 eligible non-AG2 traces produced 39 clean items, with an 18-trace reliability pilot reaching 94.4% raw agreement and Cohen's kappa 0.92. The GRPO feasibility experiment on Qwen2.5-1.5B is a documented null: every condition cleared the ROUGE-L threshold on the same 2 of 28 held-out Magentic GAIA items, and post-hoc adapter checks showed the LoRA update stayed numerically negligible at max abs delta W around 7e-6, so all conditions decode identically to the untrained checkpoint. The authors name the two blockers any conclusive study must fix — a provenance gap between training and evaluation items, and lexical-overlap scoring.
Source
↳ Follow the thread