LLM cyber-defense policies, obtained by prompt engineering alone, can be behavior-cloned into a 64,910-parameter RL agent that matches a heavily trained PPO baseline in the CybORG simulator.
IJRDO -JOURNAL OF MATHEMATICS9, 1–5 (Sep 2023)
1 Pith paper cite this work, alongside 4 external citations. Polarity classification is still indexing.
1
Pith paper citing it
4
external citations · OpenAlex
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Distilling Knowledge from Large Language Models into Lightweight Reinforcement Learning Agents for Autonomous Cyber Operations
LLM cyber-defense policies, obtained by prompt engineering alone, can be behavior-cloned into a 64,910-parameter RL agent that matches a heavily trained PPO baseline in the CybORG simulator.