Capability Gating Inside One Set of Weights Composes for Security but Not for Utility
Deployed safeguards vary by principal only outside the weights: filters get reconfigured and tiers multiplied, but inside one set of weights every request meets the same configuration. The authors define capability-gated deployment, per-principal access control inside one weight set whose configurations form a lattice where meets accumulate restrictions and joins pool a coalition's reach, instantiated by sparse rank gating over a nested factorisation with results read once from a pre-registered held-out split. Security composes provably at meets under a monotone-elicitation assumption, but utility does not: individually harmless profiles can compose into retention and fluency damage, and no compositional bound exists.
↳ Follow the thread