News
NVIDIA Researchers Show a Fine-Tuned 30B Open Model Matches Frontier Models at Attacking AI Agents for 70–125x Less
At Black Hat USA 2026, NVIDIA researchers demonstrated a fine-tuned 30B open-source model achieving a 56% exploit success rate against AI agents — matching frontier models including GPT-4o, Claude and Gemini — at 70 to 125 times lower cost and with full local privacy. The result collapses the cost floor for automated agent exploitation, which had been implicitly protected by frontier API pricing and provider abuse filters. For defenders, it means the economics of red-teaming your own agents just improved dramatically, and so did the economics of attacking them.
Source
↳ Follow the thread