SHAC-ASAM, which wraps the SHAC policy gradient with ASAM, shows improved tolerance to action noise and friction variations in simulated Ant and Humanoid walking tasks, but does not measure whether the learned minima are actually flatter.
Rethinking Optimization with Differentiable Simulation from a Global Perspective
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Differentiable simulation is a promising toolkit for fast gradient-based policy optimization and system identification. However, existing approaches to differentiable simulation have largely tackled scenarios where obtaining smooth gradients has been relatively easy, such as systems with mostly smooth dynamics. In this work, we study the challenges that differentiable simulation presents when it is not feasible to expect that a single descent reaches a global optimum, which is often a problem in contact-rich scenarios. We analyze the optimization landscapes of diverse scenarios that contain both rigid bodies and deformable objects. In dynamic environments with highly deformable objects and fluids, differentiable simulators produce rugged landscapes with nonetheless useful gradients in some parts of the space. We propose a method that combines Bayesian optimization with semi-local 'leaps' to obtain a global search method that can use gradients effectively, while also maintaining robust performance in regions with noisy gradients. We show that our approach outperforms several gradient-based and gradient-free baselines on an extensive set of experiments in simulation, and also validate the method using experiments with a real robot and deformables. Videos and supplementary materials are available at https://tinyurl.com/globdiff
citation-role summary
citation-polarity summary
fields
cs.RO 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
support 1representative citing papers
citing papers explorer
-
Improving generalization of robot locomotion policies via Sharpness-Aware Reinforcement Learning
SHAC-ASAM, which wraps the SHAC policy gradient with ASAM, shows improved tolerance to action noise and friction variations in simulated Ant and Humanoid walking tasks, but does not measure whether the learned minima are actually flatter.