PenTest2.0 demonstrates that an LLM-driven agent can autonomously suggest and run privilege escalation commands on a purposely vulnerable Linux VM, reaching root in every tested configuration but achieving automatic root detection in only four of seven.
Technical report, Royal Holloway, University of London (2024),https://pure.royalholloway.ac.uk/files/58692091/ TechReport_UnleashingAIinEthicalHacking.pdf
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
PenTest2.0: Towards Autonomous Privilege Escalation Using GenAI
PenTest2.0 demonstrates that an LLM-driven agent can autonomously suggest and run privilege escalation commands on a purposely vulnerable Linux VM, reaching root in every tested configuration but achieving automatic root detection in only four of seven.