Research
Semantic Overlays Add a Non-Text Channel to Mark Untrusted Spans, Dropping All Four PIArena Attack Families to 0% Compliance
The premise is that the serving stack knows which span is user input, tool output or instruction, but the model only sees tokens and has to infer span identity from text an attacker controls. Semantic Overlays are small learned adapters applied at chosen prefill positions to a frozen model's residual stream, creating an out-of-band annotation channel that tokens cannot forge. Marking a span 'non-executable' raised SEP separation from 24.3% to 96.5% with utility unchanged, cut TensorTrust attack success from 34.8% to 6.6%, and dropped all four PIArena attack families to 0% compliance, while marked spans stayed readable at a 92.5% exact copy rate.
↳ Follow the thread