In a human-operator cyber range, RL-trained defensive agents beat heuristic policies but performance swings with the adversary and the simulated users
Researchers tested automated network intrusion-response agents in a cyber range built for training human operators, with a variable network topology, red-team emulation and simulated user agents, where the defender must keep hosts out of red-team hands while minimizing the availability cost of its own defensive measures. Alerts came from a real SIEM platform mapped into a data modeling language the agents consume, and learned policies were optimized against a cyber attack simulator of the network. Reinforcement learning agents were overall more efficient defenders than the heuristic policy, but the paper is explicit that results depend heavily on the adversary's policy in combination with the simulated user behavior — a caveat worth carrying into any autonomous-SOC pitch.
Source
↳ Follow the thread