Fetching from the wire…
Public story · 2026-07-31 · high
The math only breaks if access is gated by credentials an attacker can't fake, not by safety training, per a new arXiv paper.
Why now: The paper posted to arXiv in July 2026, adding a formal argument to the live debate over open-weight safety.
A new paper puts a hard number on why open-weight AI can't be safe, useful, and open at the same time, per arXiv 2607.27951.
That's a problem for anyone counting on safety training to close the risk gap on open releases. The floor holds regardless of how good that training gets, as long as the model still has to answer real dual-use questions.
Their argument splits the capability a model releases from the evidence a safeguard has about who's asking. Copyable evidence, a stated purpose, a persona, a chat history an attacker can imitate, is what breaks it.
Given that, the authors work out an exact worst-case floor. It caps how much help any safeguard can deny an attacker while still answering dual-use questions usefully.
Per the authors, the way out is hard-to-copy trusted credentials that predict downstream use. That's something an attacker can't fake just by wording a prompt right. The paper doesn't say what such a credential looks like in practice, or who would issue it.
Each link below shares sources, entities, or timing with this story.
Same source domain / Shared topic
Reported by the same outlet (arxiv.org); overlapping topics (access, against, answer, arguing, argument).
Same source domain / Shared topic / Downstream implication
Reported by the same outlet (arxiv.org); overlapping topics (access, against, attacker); traces where this leads (which means).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (author, capability); pushes against this story (but).
Reported by the same outlet (arxiv.org); overlapping topics (against, attacker); pushes against this story (against).
Same source domain / Shared topic / Downstream implication
Reported by the same outlet (arxiv.org); overlapping topics (capability, downstream); traces where this leads (downstream).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (access, answer); pushes against this story (but).
Reported by the same outlet (arxiv.org); overlapping topics (against, safety); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (against, safety); pushes against this story (against).