A TD3 variant that evaluates multiple perturbed actions via short Monte Carlo rollouts reports faster learning and higher returns on HalfCheetah, Walker2d, and Swimmer.
”Continuous control actions learning and adaptation for robotic manipulation through reinforcement learning.” Autonomous Robots 46.3 (2022): 483-498
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control
A TD3 variant that evaluates multiple perturbed actions via short Monte Carlo rollouts reports faster learning and higher returns on HalfCheetah, Walker2d, and Swimmer.