Stacked LLM Defenses Fail on the Same Inputs: All 15 Measurable Layer Pairs Show Positive Correlation, phi 0.30 to 0.75
arXiv 2608.28327 (2026-08-28, cs.AI/cs.CL/cs.CR) measures the independence assumption that stacked LLM defenses quietly rely on. Running one adaptive adversary against a seven-layer stack, failure correlation is positive in all fifteen measurable pairs (phi 0.30 to 0.75) and the joint residual exceeds the multiplicative prediction by up to 0.172, while the same stack refuses four in five benign prompts and stays statistically indistinguishable from its strongest single layer. The dependence is architectural rather than sampling-based because members correlate through the shared wrapped model, so adding more diverse layers does not weaken it; the paper also gives an Adversary Access-Tier Model (A0 to A4) and a five-class inference-cost taxonomy for sizing a stack.
↳ Follow the thread