Gate on metadata before the agent acts — stale-but-well-formed data converts to wrong actions ~60% of the time, and no model tier is better at catching it
SARC-DQ (July 28, 2026) targets defects that never reach the agent's context at all: a stale price or superseded record that is perfectly well-formed in the payload and wrong only in its freshness, lineage, or provenance. On a pricing-replenishment benchmark, competent agents silently turned these into costly actions about 60% of the time, and both data-quality flags and the agents' own hedging language detected them at chance (AUC ≤ 0.50). The finding that should change architecture decisions: the conversion rate was flat across four model tiers spanning a 15x inference-price range — you cannot buy your way out of this with a better model, because evidence integrity is a separate axis from model capability. The proposed control is a metadata-aware pre-action gate rather than downstream remediation.
↳ Follow the thread