Pith. sign in

REVIEW 5 cited by

Accelerated Policy Learning with Parallel Differentiable Simulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.07137 v1 pith:4CDN42IQ submitted 2022-04-14 cs.LG cs.AIcs.GRcs.RO

classification cs.LGcs.AIcs.GRcs.RO
keywords learningdifferentiablealgorithmcontrolgradientsworkclassicalcomplex
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep reinforcement learning can generate complex control policies, but requires large amounts of training data to work effectively. Recent work has attempted to address this issue by leveraging differentiable simulators. However, inherent problems such as local minima and exploding/vanishing numerical gradients prevent these methods from being generally applied to control tasks with complex contact-rich dynamics, such as humanoid locomotion in classical RL benchmarks. In this work we present a high-performance differentiable simulator and a new policy learning algorithm (SHAC) that can effectively leverage simulation gradients, even in the presence of non-smoothness. Our learning algorithm alleviates problems with local minima through a smooth critic function, avoids vanishing/exploding gradients through a truncated learning window, and allows many physical environments to be run in parallel. We evaluate our method on classical RL control tasks, and show substantial improvements in sample efficiency and wall-clock time over state-of-the-art RL and differentiable simulation-based algorithms. In addition, we demonstrate the scalability of our method by applying it to the challenging high-dimensional problem of muscle-actuated locomotion with a large action space, achieving a greater than 17x reduction in training time over the best-performing established RL algorithm.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Transformer Transformer: A Unified Model for Motion-Conditioned Robot Co-design

    cs.RO 2026-07 conditional novelty 7.0 of 10

    A single diffusion transformer trains on tokenized robot bodies and motions to generate and optimize robot designs for unseen rewards and trajectories, outpacing evolutionary search in speed and often in reward.

  2. Physics-enhanced reinforcement learning for real-time optimal control of dynamical systems

    cs.LG 2026-07 conditional novelty 6.0 of 10

    PEARL trains a policy by differentiating a short horizon of the simulator and using a learned adjoint network to account for long-term return gradients, beating PPO, TD3, BPTT, and SHAC on two double-gyre navigation tasks.

  3. The HydroGym Reinforcement Learning Platform for Fluid Dynamics

    physics.flu-dyn 2025-12 reject novelty 6.0 of 10

    HydroGym provides 42+ (abstract claims 61+) standardized RL environments for flow control, and reports policies that transfer across Reynolds numbers and geometries, including a 38% drag-reduction transfer claim not s...

  4. Physics-Grounded Differentiable Simulation for Soft Growing Robots

    cs.RO 2025-01 conditional novelty 6.0 of 10

    A differentiable simulator for soft growing robots with a new wrinkling-based bending stiffness model, fitted and validated against real robot trajectories.

  5. ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards

    cs.RO 2025-01 conditional novelty 5.0 of 10

    ABPT averages a zero-step value gradient with an N-step backpropagation gradient so that non-differentiable reward components do not fully block policy learning.

Pith tools