PPO agents warm-started with behavioral cloning on GA-optimized trajectories outperform standard PPO in a simulated sorting task, while GA demonstrations added to DQN replay do not help.
In: 2019 IEEE 15th International Conference on Automation Science and Engineering (CASE)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Leveraging Genetic Algorithms for Efficient Demonstration Generation in Real-World Reinforcement Learning Environments
PPO agents warm-started with behavioral cloning on GA-optimized trajectories outperform standard PPO in a simulated sorting task, while GA demonstrations added to DQN replay do not help.