Safety Hacking: Constrained Best-of-N Becomes Asymptotically Certain to Return an Unsafe Output as N Grows
The paper formalizes a two-stage failure in the common pattern of sampling N outputs, filtering with a learned safety model, then picking the highest-reward survivor. An imperfect safety proxy first admits unsafe outputs into the feasible set, and reward maximization then amplifies that contamination; the derived finite-N bounds show that if unsafe-but-feasible outputs have the heavier upper reward tail, selecting an unsafe output becomes asymptotically certain as N grows, even when average proxy errors and false-positive mass are arbitrarily small. Bounding policies within a chi-squared divergence of the reference distribution gives an N-independent bound, but coverage control limits amplification without repairing an already contaminated feasible set.
↳ Follow the thread