Introduces the first open multi-host cyber range benchmark AgentCyberRange with Cage toolchain and evaluates six frontier AI systems on web exploitation and post-exploitation tasks across 110 vulnerabilities.
On the feasibility of using LLMs to autonomously execute multi-host network attacks
10 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CR 10representative citing papers
ExploitBench decomposes LLM exploitation into 16 oracle-verified capability flags and finds public frontier models trigger crashes but rarely reach arbitrary code execution on 41 V8 bugs.
A decoupled evaluation framework shows LLM penetration agents reach 90% exploitation success with ground-truth context but only 50% reconnaissance recall due to telemetry parsing failures across 50 vulnerabilities.
CyberEvolver introduces a four-layer self-evolving agent architecture with trace-to-diagnosis and population beam search that raises seed agent success rates by 13.6% on CTF, exploitation, and penetration tasks across four LLMs.
SETC framework provides the first systematic comparison of CIM, OCSF, and ECS logging standards by running 50 RCE exploits and measuring how well each captures attack indicators.
Agentic offensive security tools share design flaws enabling API key exfiltration, persistence, and sandbox escape, addressed via a new cyber kill chain and robust architecture principles.
PocketAgents introduces a manifest-driven library for LLM-based autonomous defense agents, evaluated in 18 closed-loop trials against a DarkSide-inspired attack where 13 trials produced validated blocking actions.
Structured CTI standards like ATT&CK describe adversary actions but lack the ordering, preconditions, and environmental details needed for direct multi-stage emulation, and a translation method can bridge this gap when assumptions are recorded.
Expert-defined action plans for LLM agents achieve higher task completion in lateral-movement scenarios than fully autonomous or self-scaffolded modes, but failures remain common due to brittle commands and state handling.
ATT&CK-style structured CTI narrows candidate SUT environments but cannot derive replay-ready configurations, as 97.6% of Enterprise software objects lack version or CPE information.
citing papers explorer
-
AgentCyberRange: Benchmarking Frontier AI Systems in Realistic Cyber Ranges
Introduces the first open multi-host cyber range benchmark AgentCyberRange with Cage toolchain and evaluates six frontier AI systems on web exploitation and post-exploitation tasks across 110 vulnerabilities.
-
ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents
ExploitBench decomposes LLM exploitation into 16 oracle-verified capability flags and finds public frontier models trigger crashes but rarely reach arbitrary code execution on 41 V8 bugs.
-
Decoupling Reconnaissance and Exploitation: Measuring the Capability Boundaries of LLM-Based Web Penetration Testing
A decoupled evaluation framework shows LLM penetration agents reach 90% exploitation success with ground-truth context but only 50% reconnaissance recall due to telemetry parsing failures across 50 vulnerabilities.
-
CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly
CyberEvolver introduces a four-layer self-evolving agent architecture with trace-to-diagnosis and population beam search that raises seed agent success rates by 13.6% on CTF, exploitation, and penetration tasks across four LLMs.
-
Beyond Collection: Measuring the Detection Efficacy of Modern Security Logging Standards
SETC framework provides the first systematic comparison of CIM, OCSF, and ECS logging standards by running 50 RCE exploits and measuring how well each captures attack indicators.
-
Red-Teaming the Agentic Red-Team
Agentic offensive security tools share design flaws enabling API key exfiltration, persistence, and sandbox escape, addressed via a new cyber kill chain and robust architecture principles.
-
PocketAgents: A Manifest-Driven Library of Autonomous Defense Agents
PocketAgents introduces a manifest-driven library for LLM-based autonomous defense agents, evaluated in 18 closed-loop trials against a DarkSide-inspired attack where 13 trials produced validated blocking actions.
-
The Procedural Semantics Gap in Structured CTI: A Measurement-Driven STIX Analysis for APT Emulation
Structured CTI standards like ATT&CK describe adversary actions but lack the ordering, preconditions, and environmental details needed for direct multi-stage emulation, and a translation method can bridge this gap when assumptions are recorded.
-
Autonomous Adversary: Red-Teaming in the age of LLM
Expert-defined action plans for LLM agents achieve higher task completion in lateral-movement scenarios than fully autonomous or self-scaffolded modes, but failures remain common due to brittle commands and state handling.
-
AutoSUT: The Environment Semantics Gap in Structured CTI for Adversary Emulation
ATT&CK-style structured CTI narrows candidate SUT environments but cannot derive replay-ready configurations, as 97.6% of Enterprise software objects lack version or CPE information.