Pith. sign in

CyberGym-E2E: Scalable Real-World Benchmark for AI Agents' End-to-End Cybersecurity Capabilities

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

AI has the potential to transform cybersecurity by enabling systems that can autonomously detect, analyze, and remediate software vulnerabilities. However, existing cybersecurity evaluations of AI systems are limited in scale or scope, and fail to capture the end-to-end lifecycle of real-world software vulnerability discovery and remediation. To address this gap, we propose CyberGym-E2E, a large-scale and realistic end-to-end cybersecurity benchmark that comprehensively evaluates AI agents' abilities across the full lifecycle of vulnerability discovery, PoC generation, and patch generation. CyberGym-E2E is comprehensive and scalable, as we build an automated, agent-enhanced pipeline for transforming open-source vulnerability data into realistic evaluation environments. Currently, the benchmark consists of 920 real-world vulnerabilities across 139 different open-source projects.

fields

cs.CR 1

years

2026 1

verdicts

UNVERDICTED 1

representative citing papers

Red-Teaming the Agentic Red-Team

cs.CR · 2026-06-23 · unverdicted · novelty 6.0

Agentic offensive security tools share design flaws enabling API key exfiltration, persistence, and sandbox escape, addressed via a new cyber kill chain and robust architecture principles.

citing papers explorer

Showing 1 of 1 citing paper.

  • Red-Teaming the Agentic Red-Team cs.CR · 2026-06-23 · unverdicted · none · ref 41 · internal anchor

    Agentic offensive security tools share design flaws enabling API key exfiltration, persistence, and sandbox escape, addressed via a new cyber kill chain and robust architecture principles.