Compositional Runtime Verification Catches a Prompt Injection That Split a Prohibited Task Across Four Robot Platforms
This work (arXiv 2607.23532) targets a failure class that per-platform guardrails miss by construction: individually-compliant agent actions that compose into a mission-level violation, such as a prohibited objective split across platforms to evade per-platform limits. The authors build a three-tier (platform/squad/mission) framework that decomposes a mission policy into per-agent and cross-agent aspects, aggregates verdicts over a verification-aware messaging fabric, and fuses them with a two-axis security-by-completeness algebra whose provenance names which platforms jointly triggered a violation. In a simulated ISR mission, an indirect prompt injection that made real LLM planners split a prohibited collection task across four platforms was invisible to every per-platform monitor but detected compositionally; under an injected fault campaign the best-effort central monitor emitted silent false all-clears while the verification-aware fabric emitted none.
↳ Follow the thread