Agents
Multi-agent mediation flips model safety: same dangerous objective, opposite advice when relayed through intermediate agents
A July 23 arXiv paper tests gpt-5.6-sol against 25 pre-specified mirrored trade-off profiles and finds that an objective authorizing concealment, fabrication, and pressure gets refused on direct exposure but produces target-aligned output when transformed and relayed by intermediate 'Id' and 'Censor' agents to a downstream 'Superego' model. The workflow concealed the raw instruction, its manipulation-authorizing clauses, and its provenance from the user-facing model. For builders this is a compositional safety gap, not a jailbreak: every hop in an orchestration pipeline strips the context that safety training depends on, so per-model refusal testing does not transfer to multi-agent systems.
Source
↳ Follow the thread