REVIEW 3 cited by
Depending on yourself when you should: Mentoring LLM with RL agents to become the master in cybersecurity games
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Integrating LLM and reinforcement learning (RL) agent effectively to achieve complementary performance is critical in high stake tasks like cybersecurity operations. In this study, we introduce SecurityBot, a LLM agent mentored by pre-trained RL agents, to support cybersecurity operations. In particularly, the LLM agent is supported with a profile module to generated behavior guidelines, a memory module to accumulate local experiences, a reflection module to re-evaluate choices, and an action module to reduce action space. Additionally, it adopts the collaboration mechanism to take suggestions from pre-trained RL agents, including a cursor for dynamic suggestion taken, an aggregator for multiple mentors' suggestions ranking and a caller for proactive suggestion asking. Building on the CybORG experiment framework, our experiences show that SecurityBot demonstrates significant performance improvement compared with LLM or RL standalone, achieving the complementary performance in the cybersecurity games.
Forward citations
Cited by 3 Pith papers
-
Large Language Models are Autonomous Cyber Defenders
LLM agents can be integrated as blue-team defenders in the CybORG CAGE 4 multi-agent environment, but they are about 100x slower and earn lower rewards than a GNN-based RL team.
-
VulnBot: Autonomous Penetration Testing for A Multi-Agent Collaborative Framework
VulnBot, a three-role LLM agent team with a penetration task graph and summarizer, raises penetration testing completion rates over raw GPT-4o and Llama3.1 on public benchmarks, with one end-to-end real-machine succes...
-
ARCeR: an Agentic RAG for the Automated Definition of Cyber Ranges
ARCeR, an agentic RAG system, generates syntactically valid CyRIS cyber range configuration files from natural language descriptions, outperforming a plain LLM and basic RAG in a small evaluation.
Discussion (0). Continue with ORCID to comment.