From Benchmarks to Breaches: Scaling Offensive Security

No ratings

Presented at OAIC 2025 by

Offensive AI is past its demo phase. Models now identify vulnerabilities, generate exploits, and conduct penetration tests. Yet as these systems mature and scale, developers confront a paradox: the scaffolding and abstractions designed to make models capable increasingly become their primary bottleneck. How should domain experts navigate this transition?