LLM cyber-defense policies, obtained by prompt engineering alone, can be behavior-cloned into a 64,910-parameter RL agent that matches a heavily trained PPO baseline in the CybORG simulator.
Algorithms17(2), 60 (Feb 2024)
1 Pith paper cite this work, alongside 10 external citations. Polarity classification is still indexing.
1
Pith paper citing it
10
external citations · OpenAlex
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Distilling Knowledge from Large Language Models into Lightweight Reinforcement Learning Agents for Autonomous Cyber Operations
LLM cyber-defense policies, obtained by prompt engineering alone, can be behavior-cloned into a 64,910-parameter RL agent that matches a heavily trained PPO baseline in the CybORG simulator.