Environmental perturbations to the initial state sharply reduce the rewards of PPO-trained Overcooked agents, and the proposed BAT defense, supervised kickstarting followed by adversarial fine-tuning, restores robustness and often improves clean-environment scores.
Robust deep reinforcement learning through adversarial loss,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Towards Robust Deep Reinforcement Learning against Environmental State Perturbation
Environmental perturbations to the initial state sharply reduce the rewards of PPO-trained Overcooked agents, and the proposed BAT defense, supervised kickstarting followed by adversarial fine-tuning, restores robustness and often improves clean-environment scores.