TrustShiftProbe: a compromised MCP server that behaves for N calls then defects hits 69.5% attack success, and the best defense only halves it
A paper posted 24 August (arXiv 2608.23763) names a server-side MCP threat class the authors call TrustShift: an MCP server behaves benignly during a conditioning phase, then switches to an adversarial payload once an interaction threshold is reached, which makes the defection invisible to pre-deployment static analysis that only ever sees the honest phase. The adversary is the trusted server endpoint itself, not the prompt channel or the transport, so it sits outside both indirect-prompt-injection and MITM threat models. Across frontier proprietary and open-weight models the nine attack variants reach a 69.5% mean success rate, and SHIELD, their zero-oracle runtime defense that baselines server payloads during clean trust windows, only brings that down to 42.7%.
Source
↳ Follow the thread