Fetching from the wire…
Public story · 2026-09-02 · high
Matching the offer format and choice rule across arms wiped out most of the gain, and reruns across model generations erased more of it.
Why now: The audit entered coverage on September 2, before the original 87-point figure gets cited elsewhere as a guardrail result.
Guardrails for buyer-seller agents earned an 87.4-point welfare gain in a multi-turn negotiation test, an audit of the study found. Anyone using that number to justify a guardrail policy was reading results from a testbed that gave guarded and unguarded agents different offer schemas and different rules for how the buyer picked an offer, so the two arms were never measuring the same game.
The study also reported gains of +35.0 and +28.8 across a Qwen2.5 model ladder from 1.5 billion to 14 billion parameters. The audit reran the comparison with the offer schema and the buyer's choice procedure held identical for both arms. The three welfare gains moved to +7.2, -13.9 and +23.8. One of the three headline results flipped sign once the setup stopped favoring one side.
A second problem showed up at the largest model size. The four biggest single-generation effects at 14 billion parameters averaged +229, a number that reads like a guardrail win on its own. Averaged across three separate generations of the identical setup, that average fell to +37.6, with a 95% bootstrap interval spanning -34.2 to 109.3. Differences between generations of the same run accounted for 49.9% of the variation.
The paper's fix is a construct-validity check that runs before any welfare number gets used to argue a policy point. It confirms the schema and chooser match across arms and that the effect holds across multiple generations, returning INVALID or INCONCLUSIVE if either check fails. That check is laid out in the paper.
Each link below shares sources, entities, or timing with this story.
A 4,181-problem study found confidence-based escalation beats a frozen router by 4.2 points while using 37% fewer tokens.
A preregistered test of 18,000 multi-agent missions shows failures cluster instead of scattering, which breaks the math teams use to size redundancy.
Evolved harnesses lost to a matched-budget sampling baseline once tested on a benchmark the search never saw.
It splits agent composition from runtime adaptation, and its GitHub repos are still active, not archived research code.
Across 46 model endpoints, block rates on the same forged-command test swing up to 47 points between configurations.
Four model tiers spanning a 15x price gap failed at the same rate: no model buys its way out of a stale-data problem.
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.