Pith. sign in

REVIEW 3 major objections 4 minor 42 references

Differentiating the converged contact residual, not the solver trace, keeps gradient memory almost flat as contact solves tighten, and that scale lets full-horizon plans be distilled into short-horizon residual control with much higher succ

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 04:55 UTC pith:VMFFHA4Q

load-bearing objection Solid MJX systems paper: residual IFT kills solver-trace memory, and residual-MPC distillation actually moves short-horizon success; gradient checks stay inside fixed contact sets. the 3 major comments →

arxiv 2607.24959 v1 pith:VMFFHA4Q submitted 2026-07-27 cs.RO cs.DCcs.LGcs.SYeess.SYmath.OC

Amortising Trajectory Optimisation for Residual MPC via Implicit Contact Differentiation

classification cs.RO cs.DCcs.LGcs.SYeess.SYmath.OC
keywords differentiable simulationimplicit function theoremcontact dynamicstrajectory optimisationiLQRresidual MPCoptimiser distillationmodel predictive control
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Contact-rich robot planning needs derivatives of how controls change future outcomes, but iterative contact solvers make those derivatives expensive: unrolling them stores a growing computation trace, while hand-built sensitivity systems are brittle. This paper shows you can get the same local sensitivities by differentiating only the stationarity residual at the tolerance-converged contact solution, via the Implicit Function Theorem, without changing the forward solver. Compiled temporary memory then barely moves as solver iterations rise, and grows far more slowly with active contacts and degrees of freedom than unrolled automatic differentiation. With that capacity, batched full-horizon trajectory optimisation can act as an online teacher whose actions are distilled into a policy; at runtime the policy supplies the long-horizon nominal action and short-horizon residual iLQR adds local contact-aware corrections. On three contact-rich tasks the hybrid controller raises six-step success by tens of percentage points over plain short-horizon iLQR, so amortised planning becomes practical where pure short look-ahead fails.

Core claim

An AD-assisted implicit backward pass on the stationarity residual of a regularised smooth contact solve matches converged whole-rollout finite-difference and unrolled-AD gradients at high iteration counts, while keeping compiled temporary memory nearly constant in solver effort and substantially lower as contacts and model dimension grow. That scalability enables optimiser distillation: batched full-horizon iLQR teachers feed a policy that guides short-horizon residual iLQR, raising six-step success by 28–98 percentage points over standard iLQR on Finger, Franka, and Unitree.

What carries the argument

AD-assisted implicit differentiation of the contact stationarity residual F(a, θ)=0 at the tolerance-converged acceleration: by the Implicit Function Theorem the local sensitivity is da*/dθ = −F_a⁻¹ F_θ, implemented as one transposed linear solve and a vector–Jacobian product without unrolling solver iterations or assembling a hand-derived KKT system. That derivative then powers residual MPC in which a distilled policy supplies the nominal action and short-horizon iLQR optimises only the residual.

Load-bearing premise

The derivative is trusted only inside a fixed contact set and smooth friction-cone region; if contacts appear, disappear, or switch friction mode, the local map can jump and the same guarantee does not automatically hold.

What would settle it

Re-run the whole-rollout central finite-difference check on trajectories that deliberately cross contact creation, removal, or friction-mode changes (instead of preserving the nominal contact schedule) and test whether implicit and finite-difference control gradients still agree through those events; large systematic disagreement would falsify the claim that the residual derivative is reliable for closed-loop contact-rich control.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Tighter contact solves no longer force a hard trade-off against batch size in GPU trajectory optimisation.
  • Full-horizon iLQR can be kept online as a teacher without storing per-iteration contact traces.
  • Short-horizon residual MPC guided by a distilled policy can succeed on contact tasks where plain short-horizon iLQR fails.
  • The same residual interface can be reused for other regularised contact models that expose a differentiable stationarity map.
  • Open-source release of the backward pass lets other planners swap unrolled contact AD for constant-in-K memory derivatives.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Memory headroom from dropping the solver trace is most valuable when many parallel full-horizon teachers are needed for continual distillation, not only for single-trajectory MPC.
  • Tasks whose success hinges on discovering a long contact sequence (e.g. multi-push manipulation) are the natural stress test of whether policy guidance plus short residual horizons generalises beyond the three evaluated domains.
  • Pairing the residual derivative with terminal-value or latent-model priors could further shrink the online horizon while keeping contact-aware feedback gains.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper introduces an AD-assisted implicit backward pass for the regularised contact solve in Mujoco MJX: instead of unrolling the K-iteration Newton/CG solver trace, it differentiates the stationarity residual F(a,θ)=0 (Eq. 6) at the converged solution via the Implicit Function Theorem (Eqs. 16–17), implemented with jax.lax.custom_root and a dense QR adjoint solve. The authors validate whole-rollout control gradients against central finite differences and unrolled AD on a sphere bounce-to-target task (Fig. 2), measure compiled temporary memory versus solver effort, active contacts, and model dimension (Fig. 3), and then use IFT-enabled batched full-horizon iLQR as a teacher for "optimiser distillation": a policy supplies long-horizon nominal actions while short-horizon residual iLQR supplies contact-aware corrections (Eq. 19). Closed-loop evaluation on Finger, Franka, and Unitree over 160 randomised trials reports 28–98 percentage-point gains in six-step success over standard short-horizon iLQR (Fig. 5).

Significance. If the results hold, this is a practically useful contribution at the intersection of differentiable simulation and contact-rich MPC. Specific strengths worth crediting: (i) the derivation is clean and correctly notes that the Newton metric and line search drop out of the local solution derivative (Eqs. 12–14), so no hand-assembled KKT system is needed; (ii) gradient correctness is checked against an independent central-FD reference and against unrolled AD at matched K, not just against the method's own assumptions; (iii) the memory claims are direct systems measurements across three scaling axes; (iv) the closed-loop comparison isolates policy guidance by holding the differentiable model fixed between hybrid and standard iLQR; and (v) the implementation is released open source. The honest disclosure of the locality limitation in §III is also to the authors' credit. The distillation application is a natural and apparently effective use of the memory savings, and the Franka result (standard iLQR near zero success at short horizons vs. a working hybrid controller) is a meaningful demonstration if reproducible.

major comments (3)
  1. [§V-C / Tables I–III vs. §III-C, Eqs. (13)–(17)] Tables I–III (evaluation rows) set solver iterations = 1 for all three deployed tasks, but the IFT derivative in §III-C (Eqs. 15–17) is derived at the tolerance-converged root F(a*,θ)=0, and Eq. (13)–(14) explicitly relies on F vanishing at the evaluation point. With a single Newton iteration the forward solve has generally not converged, so the deployed adjoint differentiates the residual at a non-stationary point — the same under-solved regime the paper criticises for K=1 unrolled AD (Fig. 2, §III-D-a). This directly bears on the mechanism claimed to enable the closed-loop gains. Please report ||F|| residuals at K=1 on the evaluation tasks, justify if one iteration is near-converged under the chosen regularisation R, and/or show sensitivity of the Fig. 5 results to the evaluation-time solver budget.
  2. [§III-D-a, Fig. 2] The gradient-accuracy validation holds the contact schedule fixed for all 480 FD perturbations (Fig. 2 caption). Consequently FD, unrolled AD, and IFT are only compared within a single smooth region: unrolled AD differentiates the executed (fixed-branch) trace, schedule-preserving central FD converges to the same local one-sided derivative, and IFT provably returns the local-root derivative under the C¹ / Fa≻0 assumptions stated after Eq. (9). Their agreement is fully consistent with the global contact map being nonsmooth at contact creation/removal, yet the deployed use — receding-horizon iLQR with batched line searches — perturbs trajectories across impact timing at every iteration. The paper discloses this locality (§III intro), which is good, but no experiment bounds the error when the assumption fails. Please add a complementary check, e.g., FD with schedule-crossing perturbations n
  3. [§III-D-a, Fig. 2] Gradient validation is restricted to a single free-sphere rollout (one body, nV=6, zero nominal controls, Fig. 2), while the memory and trajectory-optimisation claims cover articulated systems up to 96 DoF with sustained multi-contact phases (Hand, Unitree). The mechanism is generic, but the headline accuracy claim ('IFT matches AD and FD in accuracy at high iteration counts', Abstract/Contribution 2) currently rests on one low-dimensional scenario. Please add at least one FD/AD/IFT agreement check on an articulated contact-rich model used later in the paper (e.g., Finger or a Unitree contact phase), or scope the accuracy claim explicitly to the validated setting.
minor comments (4)
  1. [Abstract / §V-C / Table III] Reconcile the abstract/§V-C claim of 'six-step success' with Table III, which lists the Franka evaluation iLQR horizon as 4 (Finger and Unitree use 6). State which horizon the 28–98 pp figure corresponds to per task.
  2. [§III-D-b, Fig. 3] No wall-clock comparison is reported. Since the motivation is a convergence–parallelism trade-off, please report forward/backward runtime for IFT vs. unrolled AD at matched K, at least for the Fig. 3 settings; memory-only evidence leaves open whether IFT trades time for memory (e.g., the dense QR solve, §III-C).
  3. [§V-B, Algorithm 1] §V-B and Algorithm 1 (line 5) retain only 'successful' teacher trajectories, but the retention criterion is not defined per task; please state it, and comment on potential distribution shift from imitating a filtered subset of teacher behaviour.
  4. [Throughout] Typos/consistency: 'continuos' (§II-A); mixed British/American spelling throughout ('amortising' vs. 'optimized', 'optimiser'); heading spacing in §III and §V; 'Admissible contact forces.' formatting in §III-A. Eq. (8): please define f_smooth and f_constraint explicitly relative to Eq. (6). Fig. 2 top panel: state which method the control-gradient magnitude curve corresponds to.

Circularity Check

0 steps flagged

No circularity: IFT contact derivatives and residual-MPC distillation are independently validated, not forced by construction.

full rationale

The paper’s load-bearing claims do not reduce to their inputs. The contact backward pass is the standard Implicit Function Theorem applied to MJX’s stationarity residual F(a⋆,θ)=0 (Eqs. 15–17); uniqueness and Fa≻0 follow from strong convexity of the regularised objective (Eqs. 5, 9), not from a self-defined target. Gradient agreement is checked against whole-rollout central finite differences and against unrolled AD at matched K (Fig. 2)—external numerical references, not fitted quantities renamed as predictions. Memory scaling is a direct systems measurement (Fig. 3). Trajectory optimisation and optimiser distillation imitate full-horizon iLQR actions into a policy and evaluate closed-loop success on randomised held-out trials (Fig. 5, success criteria in §V-C); success is not defined by the distillation loss. Citations (Blondel modular IFT, Todorov/MuJoCo contact, classical iLQR) are external prior art, not author uniqueness theorems that force the result. Locality of IFT across contact-mode switches is a correctness/scope limitation, not circularity. No self-definitional step, fitted-input-as-prediction, or load-bearing self-citation chain appears.

Axiom & Free-Parameter Ledger

3 free parameters · 5 axioms · 2 invented entities

The central claims rest on standard IFT/convex-contact mathematics, MuJoCo’s existing regularised cone model, and ordinary supervised distillation plus iLQR assumptions. No new physical entities. Free parameters are the usual task costs, horizons, network sizes, and solver tolerances chosen for experiments—not constants fitted to invent the IFT identity.

free parameters (3)
  • Contact regularisation R and solver tolerance / max iterations K = K up to 10–20 in experiments; R from MJX defaults/task setup
    Define the smooth contact map being differentiated and the forward convergence level; chosen as MJX configuration, not derived.
  • iLQR horizons, residual penalty R, and task cost weights = e.g. Finger H_train=150, H_eval=6; Unitree height/velocity/rotation weights as in Table II
    Appendix tables set per-task state/control/terminal weights and residual penalties that shape both teacher trajectories and reported success.
  • Policy architecture and distillation hyperparameters = [64,64] or [32,32,32]; lr 1e-3 decaying; epochs 100–150
    MLP widths, AdamW schedules, EMA τ, batch sizes, and success filters determine the amortised nominal controller.
axioms (5)
  • standard math Implicit Function Theorem: if F(a*,θ)=0 and Fa is nonsingular and C¹, then a*(θ) is locally C¹ with da*/dθ = -Fa^{-1} Fθ.
    Invoked in Section III-C, Eqs. 15–17, as the sole justification for the contact backward pass.
  • domain assumption MuJoCo regularised cone contact yields a unique minimiser of the strongly convex reduced objective L within a fixed contact set and smooth projection region (M≻0, R≻0).
    Section III-A, Eqs. 4–5; taken from Todorov/MJX contact model rather than re-proved.
  • standard math Newton/CG forward iterations that reach the same local root do not change the IFT derivative of that root (P terms vanish when F=0).
    Section III-B, Eqs. 12–14; standard fixed-point/IFT argument.
  • domain assumption Local linear–quadratic iLQR models with IFT Jacobians A_t, B_t are adequate for the reported contact-rich tasks when paired with line search.
    Section IV-C and V; standard shooting assumption, not proved globally for discontinuous contact schedules.
  • ad hoc to paper Successful full-horizon hybrid iLQR trajectories are sufficient supervision for a policy that remains useful as a short-horizon nominal under residual correction.
    Algorithm 1 and Section V; empirical working hypothesis of the distillation pipeline.
invented entities (2)
  • AD-assisted residual IFT backward for MJX contact via jax.lax.custom_root on stationarity residual F independent evidence
    purpose: Obtain da*/dθ without unrolling solver iterations or hand-derived KKT blocks while leaving the forward MJX solver unchanged.
    Engineering construct, not a new physical object; purpose is modular differentiation of an existing residual (Eqs. 6–8, 17).
  • Hybrid residual MPC with stop-gradient policy nominal and iLQR residual independent evidence
    purpose: Amortise full-horizon teachers into short-horizon closed-loop control.
    Composition of known policy distillation and residual correction ideas; evaluated empirically in Section V.

pith-pipeline@v1.2.0-grok45-kimik3 · 18042 in / 3899 out tokens · 77062 ms · 2026-07-31T04:55:03.520774+00:00 · methodology

0 comments
read the original abstract

Differentiable simulation can accelerate contact-rich trajectory optimisation by exposing local sensitivities of task outcomes to controls. Existing approaches either use finite differences, which are expensive and step-size sensitive; differentiate iterative contact solvers by unrolling automatic differentiation (AD), which stores a growing computation trace; or require intricate, solver-specific KKT sensitivity derivations. We introduce an AD-assisted implicit derivative for regularised smooth contacts and apply it to Mujoco MJX, based on the Implicit Function Theorem (IFT). The method differentiates the stationarity residual at the tolerance-converged solution, avoiding both solver unrolling and hand-assembled KKT systems. IFT keeps compiled temporary memory nearly constant with solver effort, changing by less than 4$\%$ from one to ten iterations versus 10.6$\times$ growth for unrolled AD. IFT memory grows slower with active contacts and model dimension, using 20$\times$ less memory at 256 contacts and 6$\times$ less at 16 contacts and 96 DoF. We further introduce optimiser distillation for residual MPC, amortising batched full-horizon iLQR into a policy that guides short-horizon residual iLQR. Across Finger, Franka, and Unitree, this raises six-step success by 28-98 percentage points over standard iLQR.

Figures

Figures reproduced from arXiv: 2607.24959 by Aditya Kamireddypalli, Calum Arnott, Daniel Layeghi, Hashim Al-Obaidi, Michael Mistry, Steve Tonneau, Thomas Corb\`{e}res.

Figure 1
Figure 1. Figure 1: Examples of trajectories solved across different scenarios with our [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Open-loop sensitivity of a ball bounce-to-target rollout. Top: control [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Compiled temporary memory for batched differentiation through primitive-contact steps at [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Running-cost profiles of the final trajectories produced by Adam and iLQR on four contact-rich tasks. Curves show the mean over 100 trajectories [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Closed-loop performance versus planning horizon [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

42 extracted references · 6 linked inside Pith

  1. [1]

    Optnet: Differentiable optimization as a layer in neural networks,

    B. Amos and J. Z. Kolter, “Optnet: Differentiable optimization as a layer in neural networks,” inProceedings of the International Conference on Machine Learning, 2017

  2. [2]

    Differentiable convex optimization layers,

    A. Agrawal, B. Amos, S. Barratt, S. Boyd, S. Diamond, and J. Z. Kolter, “Differentiable convex optimization layers,” inAdvances in Neural Information Processing Systems, 2019

  3. [3]

    Efficient and modular implicit differentiation,

    M. Blondel, Q. Berthet, M. Cuturi, R. Frostig, S. Hoyer, F. Llinares- López, F. Pedregosa, and J.-P. Vert, “Efficient and modular implicit differentiation,” inAdvances in Neural Information Processing Systems, 2022

  4. [4]

    Differentiable Implicit Soft-Body Physics,

    J. Rojas, E. Sifakis, and L. Kavan, “Differentiable Implicit Soft-Body Physics,” 2021, arXiv:2102.05791 [cs.LG]

  5. [5]

    An implicit time-stepping scheme for rigid body dynamics with inelastic collisions and coulomb friction,

    D. E. Stewart and J. C. Trinkle, “An implicit time-stepping scheme for rigid body dynamics with inelastic collisions and coulomb friction,” International Journal for Numerical Methods in Engineering, vol. 39, no. 15, pp. 2673–2691, 1996

  6. [6]

    Optimization-based simulation of nonsmooth rigid multibody dynamics,

    M. Anitescu, “Optimization-based simulation of nonsmooth rigid multibody dynamics,”Mathematical Programming, vol. 105, pp. 113– 143, 2006

  7. [7]

    Convex and analytically-invertible dynamics with contacts and constraints: Theory and implementation in mujoco,

    E. Todorov, “Convex and analytically-invertible dynamics with contacts and constraints: Theory and implementation in mujoco,” inProceedings of the IEEE International Conference on Robotics and Automation, 2014

  8. [8]

    End-to-end differentiable physics for learning and control,

    F. d. A. Belbute-Peres, K. A. Smith, K. R. Allen, J. B. Tenenbaum, and J. Z. Kolter, “End-to-end differentiable physics for learning and control,” inAdvances in Neural Information Processing Systems, 2018

  9. [9]

    Fast and feature-complete differentiable physics for articulated rigid bodies with contact,

    K. Werling, D. Omens, J. Lee, I. Exarchos, and C. K. Liu, “Fast and feature-complete differentiable physics for articulated rigid bodies with contact,” inRobotics: Science and Systems, 2021

  10. [10]

    Dojo: A differentiable simulator for robotics,

    T. A. Howell, S. Le Cleac’h, J. Brüdigam, J. Z. Kolter, M. Schwager, and Z. Manchester, “Dojo: A differentiable simulator for robotics,” in Proceedings of the Conference on Robot Learning, 2022

  11. [11]

    Scalable differentiable physics for learning and control,

    Y .-L. Qiao, J. Liang, V . Koltun, and M. C. Lin, “Scalable differentiable physics for learning and control,” inProceedings of the International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, vol. 119, 2020

  12. [12]

    End-to-end and highly-efficient differentiable simulation for robotics,

    Q. Le Lidec, L. Montaut, Y . de Mont-Marin, F. Schramm, and J. Carpentier, “End-to-end and highly-efficient differentiable simulation for robotics,” 2024, arXiv:2409.07107 [cs.RO]

  13. [13]

    Difftaichi: Differentiable programming for physical simulation,

    Y . Hu, L. Anderson, T.-M. Li, Q. Sun, N. Carr, J. Ragan-Kelley, and F. Durand, “Difftaichi: Differentiable programming for physical simulation,” inProceedings of the International Conference on Learning Representations, 2020

  14. [14]

    Brax: A differentiable physics engine for large scale rigid body simulation,

    C. D. Freeman, E. Frey, A. Raichuk, S. Girgin, I. Mordatch, and O. Bachem, “Brax: A differentiable physics engine for large scale rigid body simulation,” inAdvances in Neural Information Processing Systems Datasets and Benchmarks Track, 2021

  15. [15]

    Mujoco xla (mjx) documentation,

    Google DeepMind, “Mujoco xla (mjx) documentation,” 2026, authori- tative software documentation, accessed May 2026

  16. [16]

    Differentiable physics simulations with contacts: Do they have correct gradients with respect to position, velocity and control?

    Z. Zhong, X. Han, and E. B. Brikis, “Differentiable physics simulations with contacts: Do they have correct gradients with respect to position, velocity and control?” 2022, arXiv:2207.05060 [cs.LG]

  17. [17]

    Improving gradient computation for differentiable physics simulation with contacts,

    Z. Zhong, X. Han, S. Dey, and E. B. Brikis, “Improving gradient computation for differentiable physics simulation with contacts,” in Proceedings of the Learning for Dynamics and Control Conference, ser. Proceedings of Machine Learning Research, vol. 211, 2023, pp. 128–141

  18. [18]

    Differentiable Simulation of Hard Contacts with Soft Gradients for Learning and Control,

    A. Paulus, A. R. Geist, P. Schumacher, V . Musil, S. Rappenecker, and G. Martius, “Differentiable Simulation of Hard Contacts with Soft Gradients for Learning and Control,” inProceedings of the International Conference on Learning Representations, Oct. 2025

  19. [19]

    A direct method for trajectory optimization of rigid bodies through contact,

    M. Posa, C. Cantu, and R. Tedrake, “A direct method for trajectory optimization of rigid bodies through contact,”The International Journal of Robotics Research, vol. 33, no. 1, pp. 69–81, 2014

  20. [20]

    Contact-Implicit Trajectory Optimization with Hydroelastic Contact and iLQR,

    V . Kurtz and H. Lin, “Contact-Implicit Trajectory Optimization with Hydroelastic Contact and iLQR,” in2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2022, pp. 8829–8834

  21. [21]

    Synthesis and stabilization of complex behaviors through online trajectory optimization,

    Y . Tassa, T. Erez, and E. Todorov, “Synthesis and stabilization of complex behaviors through online trajectory optimization,” in Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems, 2012

  22. [22]

    Crocoddyl: An efficient and versatile framework for multi-contact optimal control,

    C. Mastalli, R. Budhiraja, W. Merkt, G. Saurel, B. Hammoud, M. Naveau, J. Carpentier, L. Righetti, S. Vijayakumar, and N. Mansard, “Crocoddyl: An efficient and versatile framework for multi-contact optimal control,” inProceedings of the IEEE International Conference on Robotics and Automation, 2020, pp. 2536–2542

  23. [23]

    Whole-Body Model-Predictive Control of Legged Robots with MuJoCo,

    J. Z. Zhang, T. A. Howell, Z. Yi, C. Pan, G. Shi, G. Qu, T. Erez, Y . Tassa, and Z. Manchester, “Whole-Body Model-Predictive Control of Legged Robots with MuJoCo,” 2026, arXiv:2503.04613 [cs.RO]

  24. [24]

    Adaptive approximation of dynamics gradients via interpolation to speed up trajectory optimisation,

    D. Russell, R. Papallas, and M. Dogar, “Adaptive approximation of dynamics gradients via interpolation to speed up trajectory optimisation,” in2023 IEEE International Conference on Robotics and Automation (ICRA), May 2023

  25. [25]

    GPU Accelerated Batch Multi-Convex Trajectory Optimization for a Rectangular Holonomic Mobile Robot,

    F. Rastgar, H. Masnavi, K. Kruusamäe, A. Aabloo, and A. K. Singh, “GPU Accelerated Batch Multi-Convex Trajectory Optimization for a Rectangular Holonomic Mobile Robot,” 2021, arXiv:2109.13030 [cs.RO]

  26. [26]

    Primal-dual ilqr for gpu-accelerated learning and control in legged robots,

    L. Amatucci, J. Sousa-Pinto, G. Turrisi, D. Orban, V . Barasuol, and C. Semini, “Primal-dual ilqr for gpu-accelerated learning and control in legged robots,”IEEE Robotics and Automation Letters, vol. 11, pp. 1010–1017, 01 2026

  27. [27]

    GATO: GPU- Accelerated and Batched Trajectory Optimization for Scalable Edge Model Predictive Control,

    A. Du, E. Adabag, G. Bravo, and B. Plancher, “GATO: GPU- Accelerated and Batched Trajectory Optimization for Scalable Edge Model Predictive Control,” 2025, arXiv:2510.07625 [cs]

  28. [28]

    End-to-end training of deep visuomotor policies,

    S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,”The Journal of Machine Learning Research, vol. 17, no. 1, pp. 1334–1373, Jan. 2016

  29. [29]

    PLATO: Policy learning using adaptive trajectory optimization,

    G. Kahn, T. Zhang, S. Levine, and P. Abbeel, “PLATO: Policy learning using adaptive trajectory optimization,” in2017 IEEE International Conference on Robotics and Automation (ICRA), May 2017, pp. 3342– 3349

  30. [30]

    Mpc-net: A first principles guided policy search,

    J. Carius, F. Farshidian, and M. Hutter, “Mpc-net: A first principles guided policy search,”IEEE Robotics and Automation Letters, vol. 5, pp. 2897–2904, 02 2020

  31. [31]

    Differentiable mpc for end-to-end planning and control,

    B. Amos, I. D. J. Rodriguez, J. Sacks, B. Boots, and J. Z. Kolter, “Differentiable mpc for end-to-end planning and control,” inProceedings of the 32nd International Conference on Neural Information Processing Systems, ser. NIPS’18. Red Hook, NY , USA: Curran Associates Inc., 2018, p. 8299–8310

  32. [32]

    Exploring Model-based Planning with Policy Networks,

    T. Wang and J. Ba, “Exploring Model-based Planning with Policy Networks,” inProceedings of the International Conference on Learning Representations, Sep. 2019

  33. [33]

    Large scale model predictive control with neural networks and primal active sets,

    S. W. Chen, T. Wang, N. Atanasov, V . Kumar, and M. Morari, “Large scale model predictive control with neural networks and primal active sets,”Automatica, vol. 135, p. 109947, 2022

  34. [34]

    Residual reinforcement learning for robot control,

    T. Johannink, S. Bahl, A. Nair, J. Luo, A. Kumar, M. Loskyll, J. A. Ojea, E. Solowjow, and S. Levine, “Residual reinforcement learning for robot control,” in2019 International Conference on Robotics and Automation (ICRA). IEEE Press, 2019, p. 6023–6029

  35. [35]

    Plan online, learn offline: Efficient learning and exploration via model- based control,

    K. Lowrey, A. Rajeswaran, S. Kakade, E. Todorov, and I. Mordatch, “Plan online, learn offline: Efficient learning and exploration via model- based control,” inInternational Conference on Learning Representa- tions, 2019

  36. [36]

    TD-MPC2: Scalable, robust world models for continuous control,

    N. Hansen, H. Su, and X. Wang, “TD-MPC2: Scalable, robust world models for continuous control,” inThe Twelfth International Conference on Learning Representations, 2024

  37. [37]

    Accelerated policy learning with parallel differentiable simulation,

    J. Xu, V . Makoviychuk, Y . Narang, F. Ramos, W. Matusik, A. Garg, and M. Macklin, “Accelerated policy learning with parallel differentiable simulation,” inProceedings of the International Conference on Learning Representations, 2022

  38. [38]

    Contact-Implicit Model Predictive Control for Dexterous In-hand Manipulation: A Long-Horizon and Robust Approach,

    Y . Jiang, M. Yu, X. Zhu, M. Tomizuka, and X. Li, “Contact-Implicit Model Predictive Control for Dexterous In-hand Manipulation: A Long-Horizon and Robust Approach,” in2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2024

  39. [39]

    Dexterous contact-rich manipulation via the contact trust region,

    H. J. T. Suh, T. Pang, T. Zhao, and R. Tedrake, “Dexterous contact-rich manipulation via the contact trust region,”The International Journal of Robotics Research, p. 02783649251398875, 2026

  40. [40]

    Where to Touch, How to Contact: Hierarchical RL-MPC Framework for Geometry-Aware Long- Horizon Dexterous Manipulation,

    Z. Xie, Y . Xiang, M. Posa, and W. Jin, “Where to Touch, How to Contact: Hierarchical RL-MPC Framework for Geometry-Aware Long- Horizon Dexterous Manipulation,” 2026, arXiv:2601.10930 [cs]

  41. [41]

    Mujoco: A physics engine for model-based control,

    E. Todorov, T. Erez, and Y . Tassa, “Mujoco: A physics engine for model-based control,” in2012 IEEE/RSJ International Conference on Intelligent Robots and Systems. IEEE, 2012, pp. 5026–5033

  42. [42]

    Contact models in robotics: A comparative analysis,

    Q. Lidec, W. Jallet, L. Montaut, I. Laptev, C. Schmid, and J. Car- pentier, “Contact models in robotics: A comparative analysis,”IEEE Transactions on Robotics, vol. 40, pp. 3716–3733, 01 2024. APPENDIXI TRAININGPARAMETERS We list the full configuration used for each task. TABLE I CONFIGURATION PARAMETERS, FINGER TASK. Parameter Value / Description iLQR ho...