None
EN
BSides 2026: How AI Agents Really Perform in Offensive Security
['Ken Underhill']
eSecurity Planet
However, new research presented at the BSides 2026 conference suggests those scores reveal little about how AI agents actually behave during offensive security tasks.
Rather than focusing solely on benchmark solve rates, researcher Tarun Koyalwar analyzed the decision-making processes of AI agents during offensive security tasks.
Key takeaways of the LLM benchmark researchTraditional benchmark scores reveal little about how AI agents actually perform offensive security tasks.
For security teams evaluating AI-powered offensive security tools, the findings suggest that behavioral telemetry may provide more actionable insight than benchmark percentages alone.
As organizations adopt autonomous AI for offensive security, understanding how agents reason and where they fail will become increasingly important.