A Safety Trilemma: Useful Capability, Reliable Safety, and Open Access Cannot Coexist Under Copyable Context
This paper (arXiv 2607.27951, July 30) separates the capability a model releases from the evidence it has about downstream use, and shows that when that evidence is copyable — a request, a persona, an interaction history an attacker can imitate — there is an exact worst-case floor on attacker assistance for any safeguard that still preserves useful answers on dual-use tasks. The result is a trilemma: Useful Capability, Reliable Safety, and Open Access cannot all hold. The authors argue hard-to-copy trusted credentials that predict actual downstream use are the complement that lowers the floor, and identify the stronger condition needed to eliminate it, citing dual-use evaluations, adaptive attacks, and deployed trusted-access programs as supporting evidence.
Source
↳ Follow the thread