Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing

Agent Agent Planning Tool Learning Multi-agent Collaboration Agent Eval Benchmarks
我们首次在真实的企业环境中,对人工智能代理与人类网络安全专业人员进行了全面对比评估。我们在一个包含约8000台主机、横跨12个子网的大型大学网络中,对十名网络安全专业人员、六个现有AI代理以及我们新开发的代理框架ARTEMIS进行了测试。ARTEMIS是一个多代理系统,具备动态提示生成、可任意扩展子代理和自动漏洞分级功能。在本次对比研究中,ARTEMIS总体排名第二,共发现9个有效漏洞,有效提交率达到82%,表现优于10名人类参与者中的9位。尽管现有的Codex和CyAgent等代理框架的表现不及大多数人类参与者,但ARTEMIS展现出与最强人类参与者相当的技术深度和报告质量。我们观察到,AI代理在系统性枚举、并行化漏洞利用以及成本控制方面具有优势——某些ARTEMIS变体的运行成本仅为每小时18美元,而专业渗透测试人员的成本则为每小时60美元。同时,我们也发现了AI代理的关键能力短板:其误报率较高,在涉及图形用户界面(GUI)的任务中表现不佳。
We present the first comprehensive evaluation of AI agents against human cybersecurity professionals in a live enterprise environment. We evaluate ten cybersecurity professionals alongside six existing AI agents and ARTEMIS, our new agent scaffold, on a large university network consisting of ~8,000 hosts across 12 subnets. ARTEMIS is a multi-agent framework featuring dynamic prompt generation, arbitrary sub-agents, and automatic vulnerability triaging. In our comparative study, ARTEMIS placed second overall, discovering 9 valid vulnerabilities with an 82% valid submission rate and outperforming 9 of 10 human participants. While existing scaffolds such as Codex and CyAgent underperformed relative to most human participants, ARTEMIS demonstrated technical sophistication and submission quality comparable to the strongest participants. We observe that AI agents offer advantages in systematic enumeration, parallel exploitation, and cost -- certain ARTEMIS variants cost $18/hour versus $60/hour for professional penetration testers. We also identify key capability gaps: AI agents exhibit higher false-positive rates and struggle with GUI-based tasks.
许愿