RedEvoAgent distills red-team trajectories into a readable attack skill that transfers across execution harnesses
RedEvoAgent, posted 27 August, targets jailbreaks against agents in product-level execution harnesses, where a successful attack triggers real tool use and persistent state changes rather than just unsafe text. Instead of retrieving whole past trajectories, which the authors argue reuses misleading experience because of retrieval bias and unclear tool credit while bloating context, it distills cross-case trajectories into one concise human-readable attack skill. The skill evolves through tool-effectiveness profiling and Deciding-Tool Attribution, gated by a validation ratchet that keeps only updates that improve validation performance, and it transfers across attacker models and target harnesses.
Source
↳ Follow the thread