A hybrid offline-online PPO with TWTL reward shaping is claimed to accelerate delayed-reward learning, yet the theoretical guarantees do not hold as written.
Overcoming exploration: Deep reinforcement learning for continuous control in cluttered environments from temporal logic specifications
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2024 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards
A hybrid offline-online PPO with TWTL reward shaping is claimed to accelerate delayed-reward learning, yet the theoretical guarantees do not hold as written.