Fetching from the wire…
Research2026-08-25 · source-backed
The paper formalizes a two-stage failure in the pattern everyone uses: sample N outputs, filter with a learned safety model, pick the highest-reward survivor (arXiv 2608.22915). An imperfect proxy admits unsafe outputs into the feasible set, then reward maximization amplifies that contamination. If unsafe-but-feasible outputs carry the heavier upper reward tail, selecting one becomes asymptotically certain as N grows, even when average proxy error is arbitrarily small. Bounding policies within a chi-squared divergence of the reference gives an N-independent bound, but coverage control limits amplification without repairing an already contaminated set.
Each link below shares sources, entities, or timing with this story.
Shared entity: Constrained / Same source domain / Earlier coverage
Both cover Constrained; reported by the same outlet (arxiv.org); earlier Constrained coverage from 2026-06-14.
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (already, bound); pushes against this story (against).
Same source domain / Shared topic
Reported by the same outlet (arxiv.org); overlapping topics (amplification, average).
Reported by the same outlet (arxiv.org); overlapping topics (admit, already).
Reported by the same outlet (arxiv.org); overlapping topics (already, average).
Reported by the same outlet (arxiv.org); overlapping topics (already, average).
Reported by the same outlet (arxiv.org); overlapping topics (average, bound).
Reported by the same outlet (arxiv.org); overlapping topics (amplify, output).