Pith. sign in

REVIEW 2 cited by

Do Differentiable Simulators Give Better Policy Gradients?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.00817 v2 pith:HA5BQLJW submitted 2022-02-02 cs.LG cs.AIcs.RO

classification cs.LGcs.AIcs.RO
keywords alphadifferentiableestimatorfirst-ordergradientssimulatorsestimatesestimators
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Differentiable simulators promise faster computation time for reinforcement learning by replacing zeroth-order gradient estimates of a stochastic objective with an estimate based on first-order gradients. However, it is yet unclear what factors decide the performance of the two estimators on complex landscapes that involve long-horizon planning and control on physical systems, despite the crucial relevance of this question for the utility of differentiable simulators. We show that characteristics of certain physical systems, such as stiffness or discontinuities, may compromise the efficacy of the first-order estimator, and analyze this phenomenon through the lens of bias and variance. We additionally propose an $\alpha$-order gradient estimator, with $\alpha \in [0,1]$, which correctly utilizes exact gradients to combine the efficiency of first-order estimates with the robustness of zero-order methods. We demonstrate the pitfalls of traditional estimators and the advantages of the $\alpha$-order estimator on some numerical examples.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Robust and Efficient MuJoCo-based Model Predictive Control via Web of Affine Spaces Derivatives

    cs.RO 2025-12 conditional novelty 6.0 of 10

    WASP derivative reuse speeds up MuJoCo MPC by 1.26–2.08x versus finite differences on selected locomotion tasks, with comparable task costs.

  2. Enhancing Robotic System Robustness via Lyapunov Exponent-Based Optimization

    cs.RO 2024-12 conditional novelty 5.0 of 10

    A differentiable 'sum of Lyapunov exponents' score is used as a robustness objective to co-optimize robot hardware and control policies in simulation.

Pith tools