Agents
Pentest-agent study: an autonomous PentestGPT on Claude Opus 4.8 solves all three targets, and adding memory did not help either system
arXiv 2609.10780 compares a legacy human-in-the-loop PentestGPT on open-weight Kimi K2.5 with a newer autonomous version on Claude Opus 4.8. Across three public targets, the autonomous one solved everything. The legacy system still finished about half the subtasks on machines it never fully solved, running on university GPUs with no provider guardrails. Adding a coverage-memory layer improved neither system. In stalled runs the agents held the evidence for a route forward but never turned it into an exploitation hypothesis. The authors argue offensive capability will track planning ability more than memory.
Source
↳ Follow the thread