Round-trip test finds the natural-language channel between models is lossy and asymmetric
Submitted 2026-09-18, 2609.21509 measures how much tree-structured content survives serialization into natural language: a generator turns a procedurally generated arithmetic expression into a word problem, a separate extractor recovers the expression from the word problem alone, and symbolic equivalence gives an exact oracle with no judge in the loop. Evaluating all pairwise combinations of sixteen models produces a communication matrix whose marginals separate generation quality from extraction quality, and the headline finding is that the channel is lossy and asymmetric, with results changing when you swap which model generates and which extracts. Anyone passing free-text intermediates between agents now has a cheap oracle-backed way to measure what their handoff is dropping.
Source
↳ Follow the thread