REVIEW 3 major objections 4 minor 42 references
Differentiating the converged contact residual, not the solver trace, keeps gradient memory almost flat as contact solves tighten, and that scale lets full-horizon plans be distilled into short-horizon residual control with much higher succ
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 04:55 UTC pith:VMFFHA4Q
load-bearing objection Solid MJX systems paper: residual IFT kills solver-trace memory, and residual-MPC distillation actually moves short-horizon success; gradient checks stay inside fixed contact sets. the 3 major comments →
Amortising Trajectory Optimisation for Residual MPC via Implicit Contact Differentiation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
An AD-assisted implicit backward pass on the stationarity residual of a regularised smooth contact solve matches converged whole-rollout finite-difference and unrolled-AD gradients at high iteration counts, while keeping compiled temporary memory nearly constant in solver effort and substantially lower as contacts and model dimension grow. That scalability enables optimiser distillation: batched full-horizon iLQR teachers feed a policy that guides short-horizon residual iLQR, raising six-step success by 28–98 percentage points over standard iLQR on Finger, Franka, and Unitree.
What carries the argument
AD-assisted implicit differentiation of the contact stationarity residual F(a, θ)=0 at the tolerance-converged acceleration: by the Implicit Function Theorem the local sensitivity is da*/dθ = −F_a⁻¹ F_θ, implemented as one transposed linear solve and a vector–Jacobian product without unrolling solver iterations or assembling a hand-derived KKT system. That derivative then powers residual MPC in which a distilled policy supplies the nominal action and short-horizon iLQR optimises only the residual.
Load-bearing premise
The derivative is trusted only inside a fixed contact set and smooth friction-cone region; if contacts appear, disappear, or switch friction mode, the local map can jump and the same guarantee does not automatically hold.
What would settle it
Re-run the whole-rollout central finite-difference check on trajectories that deliberately cross contact creation, removal, or friction-mode changes (instead of preserving the nominal contact schedule) and test whether implicit and finite-difference control gradients still agree through those events; large systematic disagreement would falsify the claim that the residual derivative is reliable for closed-loop contact-rich control.
If this is right
- Tighter contact solves no longer force a hard trade-off against batch size in GPU trajectory optimisation.
- Full-horizon iLQR can be kept online as a teacher without storing per-iteration contact traces.
- Short-horizon residual MPC guided by a distilled policy can succeed on contact tasks where plain short-horizon iLQR fails.
- The same residual interface can be reused for other regularised contact models that expose a differentiable stationarity map.
- Open-source release of the backward pass lets other planners swap unrolled contact AD for constant-in-K memory derivatives.
Where Pith is reading between the lines
- Memory headroom from dropping the solver trace is most valuable when many parallel full-horizon teachers are needed for continual distillation, not only for single-trajectory MPC.
- Tasks whose success hinges on discovering a long contact sequence (e.g. multi-push manipulation) are the natural stress test of whether policy guidance plus short residual horizons generalises beyond the three evaluated domains.
- Pairing the residual derivative with terminal-value or latent-model priors could further shrink the online horizon while keeping contact-aware feedback gains.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces an AD-assisted implicit backward pass for the regularised contact solve in Mujoco MJX: instead of unrolling the K-iteration Newton/CG solver trace, it differentiates the stationarity residual F(a,θ)=0 (Eq. 6) at the converged solution via the Implicit Function Theorem (Eqs. 16–17), implemented with jax.lax.custom_root and a dense QR adjoint solve. The authors validate whole-rollout control gradients against central finite differences and unrolled AD on a sphere bounce-to-target task (Fig. 2), measure compiled temporary memory versus solver effort, active contacts, and model dimension (Fig. 3), and then use IFT-enabled batched full-horizon iLQR as a teacher for "optimiser distillation": a policy supplies long-horizon nominal actions while short-horizon residual iLQR supplies contact-aware corrections (Eq. 19). Closed-loop evaluation on Finger, Franka, and Unitree over 160 randomised trials reports 28–98 percentage-point gains in six-step success over standard short-horizon iLQR (Fig. 5).
Significance. If the results hold, this is a practically useful contribution at the intersection of differentiable simulation and contact-rich MPC. Specific strengths worth crediting: (i) the derivation is clean and correctly notes that the Newton metric and line search drop out of the local solution derivative (Eqs. 12–14), so no hand-assembled KKT system is needed; (ii) gradient correctness is checked against an independent central-FD reference and against unrolled AD at matched K, not just against the method's own assumptions; (iii) the memory claims are direct systems measurements across three scaling axes; (iv) the closed-loop comparison isolates policy guidance by holding the differentiable model fixed between hybrid and standard iLQR; and (v) the implementation is released open source. The honest disclosure of the locality limitation in §III is also to the authors' credit. The distillation application is a natural and apparently effective use of the memory savings, and the Franka result (standard iLQR near zero success at short horizons vs. a working hybrid controller) is a meaningful demonstration if reproducible.
major comments (3)
- [§V-C / Tables I–III vs. §III-C, Eqs. (13)–(17)] Tables I–III (evaluation rows) set solver iterations = 1 for all three deployed tasks, but the IFT derivative in §III-C (Eqs. 15–17) is derived at the tolerance-converged root F(a*,θ)=0, and Eq. (13)–(14) explicitly relies on F vanishing at the evaluation point. With a single Newton iteration the forward solve has generally not converged, so the deployed adjoint differentiates the residual at a non-stationary point — the same under-solved regime the paper criticises for K=1 unrolled AD (Fig. 2, §III-D-a). This directly bears on the mechanism claimed to enable the closed-loop gains. Please report ||F|| residuals at K=1 on the evaluation tasks, justify if one iteration is near-converged under the chosen regularisation R, and/or show sensitivity of the Fig. 5 results to the evaluation-time solver budget.
- [§III-D-a, Fig. 2] The gradient-accuracy validation holds the contact schedule fixed for all 480 FD perturbations (Fig. 2 caption). Consequently FD, unrolled AD, and IFT are only compared within a single smooth region: unrolled AD differentiates the executed (fixed-branch) trace, schedule-preserving central FD converges to the same local one-sided derivative, and IFT provably returns the local-root derivative under the C¹ / Fa≻0 assumptions stated after Eq. (9). Their agreement is fully consistent with the global contact map being nonsmooth at contact creation/removal, yet the deployed use — receding-horizon iLQR with batched line searches — perturbs trajectories across impact timing at every iteration. The paper discloses this locality (§III intro), which is good, but no experiment bounds the error when the assumption fails. Please add a complementary check, e.g., FD with schedule-crossing perturbations n
- [§III-D-a, Fig. 2] Gradient validation is restricted to a single free-sphere rollout (one body, nV=6, zero nominal controls, Fig. 2), while the memory and trajectory-optimisation claims cover articulated systems up to 96 DoF with sustained multi-contact phases (Hand, Unitree). The mechanism is generic, but the headline accuracy claim ('IFT matches AD and FD in accuracy at high iteration counts', Abstract/Contribution 2) currently rests on one low-dimensional scenario. Please add at least one FD/AD/IFT agreement check on an articulated contact-rich model used later in the paper (e.g., Finger or a Unitree contact phase), or scope the accuracy claim explicitly to the validated setting.
minor comments (4)
- [Abstract / §V-C / Table III] Reconcile the abstract/§V-C claim of 'six-step success' with Table III, which lists the Franka evaluation iLQR horizon as 4 (Finger and Unitree use 6). State which horizon the 28–98 pp figure corresponds to per task.
- [§III-D-b, Fig. 3] No wall-clock comparison is reported. Since the motivation is a convergence–parallelism trade-off, please report forward/backward runtime for IFT vs. unrolled AD at matched K, at least for the Fig. 3 settings; memory-only evidence leaves open whether IFT trades time for memory (e.g., the dense QR solve, §III-C).
- [§V-B, Algorithm 1] §V-B and Algorithm 1 (line 5) retain only 'successful' teacher trajectories, but the retention criterion is not defined per task; please state it, and comment on potential distribution shift from imitating a filtered subset of teacher behaviour.
- [Throughout] Typos/consistency: 'continuos' (§II-A); mixed British/American spelling throughout ('amortising' vs. 'optimized', 'optimiser'); heading spacing in §III and §V; 'Admissible contact forces.' formatting in §III-A. Eq. (8): please define f_smooth and f_constraint explicitly relative to Eq. (6). Fig. 2 top panel: state which method the control-gradient magnitude curve corresponds to.
Circularity Check
No circularity: IFT contact derivatives and residual-MPC distillation are independently validated, not forced by construction.
full rationale
The paper’s load-bearing claims do not reduce to their inputs. The contact backward pass is the standard Implicit Function Theorem applied to MJX’s stationarity residual F(a⋆,θ)=0 (Eqs. 15–17); uniqueness and Fa≻0 follow from strong convexity of the regularised objective (Eqs. 5, 9), not from a self-defined target. Gradient agreement is checked against whole-rollout central finite differences and against unrolled AD at matched K (Fig. 2)—external numerical references, not fitted quantities renamed as predictions. Memory scaling is a direct systems measurement (Fig. 3). Trajectory optimisation and optimiser distillation imitate full-horizon iLQR actions into a policy and evaluate closed-loop success on randomised held-out trials (Fig. 5, success criteria in §V-C); success is not defined by the distillation loss. Citations (Blondel modular IFT, Todorov/MuJoCo contact, classical iLQR) are external prior art, not author uniqueness theorems that force the result. Locality of IFT across contact-mode switches is a correctness/scope limitation, not circularity. No self-definitional step, fitted-input-as-prediction, or load-bearing self-citation chain appears.
Axiom & Free-Parameter Ledger
free parameters (3)
- Contact regularisation R and solver tolerance / max iterations K =
K up to 10–20 in experiments; R from MJX defaults/task setup
- iLQR horizons, residual penalty R, and task cost weights =
e.g. Finger H_train=150, H_eval=6; Unitree height/velocity/rotation weights as in Table II
- Policy architecture and distillation hyperparameters =
[64,64] or [32,32,32]; lr 1e-3 decaying; epochs 100–150
axioms (5)
- standard math Implicit Function Theorem: if F(a*,θ)=0 and Fa is nonsingular and C¹, then a*(θ) is locally C¹ with da*/dθ = -Fa^{-1} Fθ.
- domain assumption MuJoCo regularised cone contact yields a unique minimiser of the strongly convex reduced objective L within a fixed contact set and smooth projection region (M≻0, R≻0).
- standard math Newton/CG forward iterations that reach the same local root do not change the IFT derivative of that root (P terms vanish when F=0).
- domain assumption Local linear–quadratic iLQR models with IFT Jacobians A_t, B_t are adequate for the reported contact-rich tasks when paired with line search.
- ad hoc to paper Successful full-horizon hybrid iLQR trajectories are sufficient supervision for a policy that remains useful as a short-horizon nominal under residual correction.
invented entities (2)
-
AD-assisted residual IFT backward for MJX contact via jax.lax.custom_root on stationarity residual F
independent evidence
-
Hybrid residual MPC with stop-gradient policy nominal and iLQR residual
independent evidence
read the original abstract
Differentiable simulation can accelerate contact-rich trajectory optimisation by exposing local sensitivities of task outcomes to controls. Existing approaches either use finite differences, which are expensive and step-size sensitive; differentiate iterative contact solvers by unrolling automatic differentiation (AD), which stores a growing computation trace; or require intricate, solver-specific KKT sensitivity derivations. We introduce an AD-assisted implicit derivative for regularised smooth contacts and apply it to Mujoco MJX, based on the Implicit Function Theorem (IFT). The method differentiates the stationarity residual at the tolerance-converged solution, avoiding both solver unrolling and hand-assembled KKT systems. IFT keeps compiled temporary memory nearly constant with solver effort, changing by less than 4$\%$ from one to ten iterations versus 10.6$\times$ growth for unrolled AD. IFT memory grows slower with active contacts and model dimension, using 20$\times$ less memory at 256 contacts and 6$\times$ less at 16 contacts and 96 DoF. We further introduce optimiser distillation for residual MPC, amortising batched full-horizon iLQR into a policy that guides short-horizon residual iLQR. Across Finger, Franka, and Unitree, this raises six-step success by 28-98 percentage points over standard iLQR.
Figures
Reference graph
Works this paper leans on
-
[1]
Optnet: Differentiable optimization as a layer in neural networks,
B. Amos and J. Z. Kolter, “Optnet: Differentiable optimization as a layer in neural networks,” inProceedings of the International Conference on Machine Learning, 2017
2017
-
[2]
Differentiable convex optimization layers,
A. Agrawal, B. Amos, S. Barratt, S. Boyd, S. Diamond, and J. Z. Kolter, “Differentiable convex optimization layers,” inAdvances in Neural Information Processing Systems, 2019
2019
-
[3]
Efficient and modular implicit differentiation,
M. Blondel, Q. Berthet, M. Cuturi, R. Frostig, S. Hoyer, F. Llinares- López, F. Pedregosa, and J.-P. Vert, “Efficient and modular implicit differentiation,” inAdvances in Neural Information Processing Systems, 2022
2022
-
[4]
Differentiable Implicit Soft-Body Physics,
J. Rojas, E. Sifakis, and L. Kavan, “Differentiable Implicit Soft-Body Physics,” 2021, arXiv:2102.05791 [cs.LG]
Pith/arXiv arXiv 2021
-
[5]
An implicit time-stepping scheme for rigid body dynamics with inelastic collisions and coulomb friction,
D. E. Stewart and J. C. Trinkle, “An implicit time-stepping scheme for rigid body dynamics with inelastic collisions and coulomb friction,” International Journal for Numerical Methods in Engineering, vol. 39, no. 15, pp. 2673–2691, 1996
1996
-
[6]
Optimization-based simulation of nonsmooth rigid multibody dynamics,
M. Anitescu, “Optimization-based simulation of nonsmooth rigid multibody dynamics,”Mathematical Programming, vol. 105, pp. 113– 143, 2006
2006
-
[7]
Convex and analytically-invertible dynamics with contacts and constraints: Theory and implementation in mujoco,
E. Todorov, “Convex and analytically-invertible dynamics with contacts and constraints: Theory and implementation in mujoco,” inProceedings of the IEEE International Conference on Robotics and Automation, 2014
2014
-
[8]
End-to-end differentiable physics for learning and control,
F. d. A. Belbute-Peres, K. A. Smith, K. R. Allen, J. B. Tenenbaum, and J. Z. Kolter, “End-to-end differentiable physics for learning and control,” inAdvances in Neural Information Processing Systems, 2018
2018
-
[9]
Fast and feature-complete differentiable physics for articulated rigid bodies with contact,
K. Werling, D. Omens, J. Lee, I. Exarchos, and C. K. Liu, “Fast and feature-complete differentiable physics for articulated rigid bodies with contact,” inRobotics: Science and Systems, 2021
2021
-
[10]
Dojo: A differentiable simulator for robotics,
T. A. Howell, S. Le Cleac’h, J. Brüdigam, J. Z. Kolter, M. Schwager, and Z. Manchester, “Dojo: A differentiable simulator for robotics,” in Proceedings of the Conference on Robot Learning, 2022
2022
-
[11]
Scalable differentiable physics for learning and control,
Y .-L. Qiao, J. Liang, V . Koltun, and M. C. Lin, “Scalable differentiable physics for learning and control,” inProceedings of the International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, vol. 119, 2020
2020
-
[12]
End-to-end and highly-efficient differentiable simulation for robotics,
Q. Le Lidec, L. Montaut, Y . de Mont-Marin, F. Schramm, and J. Carpentier, “End-to-end and highly-efficient differentiable simulation for robotics,” 2024, arXiv:2409.07107 [cs.RO]
Pith/arXiv arXiv 2024
-
[13]
Difftaichi: Differentiable programming for physical simulation,
Y . Hu, L. Anderson, T.-M. Li, Q. Sun, N. Carr, J. Ragan-Kelley, and F. Durand, “Difftaichi: Differentiable programming for physical simulation,” inProceedings of the International Conference on Learning Representations, 2020
2020
-
[14]
Brax: A differentiable physics engine for large scale rigid body simulation,
C. D. Freeman, E. Frey, A. Raichuk, S. Girgin, I. Mordatch, and O. Bachem, “Brax: A differentiable physics engine for large scale rigid body simulation,” inAdvances in Neural Information Processing Systems Datasets and Benchmarks Track, 2021
2021
-
[15]
Mujoco xla (mjx) documentation,
Google DeepMind, “Mujoco xla (mjx) documentation,” 2026, authori- tative software documentation, accessed May 2026
2026
-
[16]
Z. Zhong, X. Han, and E. B. Brikis, “Differentiable physics simulations with contacts: Do they have correct gradients with respect to position, velocity and control?” 2022, arXiv:2207.05060 [cs.LG]
Pith/arXiv arXiv 2022
-
[17]
Improving gradient computation for differentiable physics simulation with contacts,
Z. Zhong, X. Han, S. Dey, and E. B. Brikis, “Improving gradient computation for differentiable physics simulation with contacts,” in Proceedings of the Learning for Dynamics and Control Conference, ser. Proceedings of Machine Learning Research, vol. 211, 2023, pp. 128–141
2023
-
[18]
Differentiable Simulation of Hard Contacts with Soft Gradients for Learning and Control,
A. Paulus, A. R. Geist, P. Schumacher, V . Musil, S. Rappenecker, and G. Martius, “Differentiable Simulation of Hard Contacts with Soft Gradients for Learning and Control,” inProceedings of the International Conference on Learning Representations, Oct. 2025
2025
-
[19]
A direct method for trajectory optimization of rigid bodies through contact,
M. Posa, C. Cantu, and R. Tedrake, “A direct method for trajectory optimization of rigid bodies through contact,”The International Journal of Robotics Research, vol. 33, no. 1, pp. 69–81, 2014
2014
-
[20]
Contact-Implicit Trajectory Optimization with Hydroelastic Contact and iLQR,
V . Kurtz and H. Lin, “Contact-Implicit Trajectory Optimization with Hydroelastic Contact and iLQR,” in2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2022, pp. 8829–8834
2022
-
[21]
Synthesis and stabilization of complex behaviors through online trajectory optimization,
Y . Tassa, T. Erez, and E. Todorov, “Synthesis and stabilization of complex behaviors through online trajectory optimization,” in Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems, 2012
2012
-
[22]
Crocoddyl: An efficient and versatile framework for multi-contact optimal control,
C. Mastalli, R. Budhiraja, W. Merkt, G. Saurel, B. Hammoud, M. Naveau, J. Carpentier, L. Righetti, S. Vijayakumar, and N. Mansard, “Crocoddyl: An efficient and versatile framework for multi-contact optimal control,” inProceedings of the IEEE International Conference on Robotics and Automation, 2020, pp. 2536–2542
2020
-
[23]
Whole-Body Model-Predictive Control of Legged Robots with MuJoCo,
J. Z. Zhang, T. A. Howell, Z. Yi, C. Pan, G. Shi, G. Qu, T. Erez, Y . Tassa, and Z. Manchester, “Whole-Body Model-Predictive Control of Legged Robots with MuJoCo,” 2026, arXiv:2503.04613 [cs.RO]
arXiv 2026
-
[24]
Adaptive approximation of dynamics gradients via interpolation to speed up trajectory optimisation,
D. Russell, R. Papallas, and M. Dogar, “Adaptive approximation of dynamics gradients via interpolation to speed up trajectory optimisation,” in2023 IEEE International Conference on Robotics and Automation (ICRA), May 2023
2023
-
[25]
GPU Accelerated Batch Multi-Convex Trajectory Optimization for a Rectangular Holonomic Mobile Robot,
F. Rastgar, H. Masnavi, K. Kruusamäe, A. Aabloo, and A. K. Singh, “GPU Accelerated Batch Multi-Convex Trajectory Optimization for a Rectangular Holonomic Mobile Robot,” 2021, arXiv:2109.13030 [cs.RO]
Pith/arXiv arXiv 2021
-
[26]
Primal-dual ilqr for gpu-accelerated learning and control in legged robots,
L. Amatucci, J. Sousa-Pinto, G. Turrisi, D. Orban, V . Barasuol, and C. Semini, “Primal-dual ilqr for gpu-accelerated learning and control in legged robots,”IEEE Robotics and Automation Letters, vol. 11, pp. 1010–1017, 01 2026
2026
-
[27]
A. Du, E. Adabag, G. Bravo, and B. Plancher, “GATO: GPU- Accelerated and Batched Trajectory Optimization for Scalable Edge Model Predictive Control,” 2025, arXiv:2510.07625 [cs]
Pith/arXiv arXiv 2025
-
[28]
End-to-end training of deep visuomotor policies,
S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,”The Journal of Machine Learning Research, vol. 17, no. 1, pp. 1334–1373, Jan. 2016
2016
-
[29]
PLATO: Policy learning using adaptive trajectory optimization,
G. Kahn, T. Zhang, S. Levine, and P. Abbeel, “PLATO: Policy learning using adaptive trajectory optimization,” in2017 IEEE International Conference on Robotics and Automation (ICRA), May 2017, pp. 3342– 3349
2017
-
[30]
Mpc-net: A first principles guided policy search,
J. Carius, F. Farshidian, and M. Hutter, “Mpc-net: A first principles guided policy search,”IEEE Robotics and Automation Letters, vol. 5, pp. 2897–2904, 02 2020
2020
-
[31]
Differentiable mpc for end-to-end planning and control,
B. Amos, I. D. J. Rodriguez, J. Sacks, B. Boots, and J. Z. Kolter, “Differentiable mpc for end-to-end planning and control,” inProceedings of the 32nd International Conference on Neural Information Processing Systems, ser. NIPS’18. Red Hook, NY , USA: Curran Associates Inc., 2018, p. 8299–8310
2018
-
[32]
Exploring Model-based Planning with Policy Networks,
T. Wang and J. Ba, “Exploring Model-based Planning with Policy Networks,” inProceedings of the International Conference on Learning Representations, Sep. 2019
2019
-
[33]
Large scale model predictive control with neural networks and primal active sets,
S. W. Chen, T. Wang, N. Atanasov, V . Kumar, and M. Morari, “Large scale model predictive control with neural networks and primal active sets,”Automatica, vol. 135, p. 109947, 2022
2022
-
[34]
Residual reinforcement learning for robot control,
T. Johannink, S. Bahl, A. Nair, J. Luo, A. Kumar, M. Loskyll, J. A. Ojea, E. Solowjow, and S. Levine, “Residual reinforcement learning for robot control,” in2019 International Conference on Robotics and Automation (ICRA). IEEE Press, 2019, p. 6023–6029
2019
-
[35]
Plan online, learn offline: Efficient learning and exploration via model- based control,
K. Lowrey, A. Rajeswaran, S. Kakade, E. Todorov, and I. Mordatch, “Plan online, learn offline: Efficient learning and exploration via model- based control,” inInternational Conference on Learning Representa- tions, 2019
2019
-
[36]
TD-MPC2: Scalable, robust world models for continuous control,
N. Hansen, H. Su, and X. Wang, “TD-MPC2: Scalable, robust world models for continuous control,” inThe Twelfth International Conference on Learning Representations, 2024
2024
-
[37]
Accelerated policy learning with parallel differentiable simulation,
J. Xu, V . Makoviychuk, Y . Narang, F. Ramos, W. Matusik, A. Garg, and M. Macklin, “Accelerated policy learning with parallel differentiable simulation,” inProceedings of the International Conference on Learning Representations, 2022
2022
-
[38]
Contact-Implicit Model Predictive Control for Dexterous In-hand Manipulation: A Long-Horizon and Robust Approach,
Y . Jiang, M. Yu, X. Zhu, M. Tomizuka, and X. Li, “Contact-Implicit Model Predictive Control for Dexterous In-hand Manipulation: A Long-Horizon and Robust Approach,” in2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2024
2024
-
[39]
Dexterous contact-rich manipulation via the contact trust region,
H. J. T. Suh, T. Pang, T. Zhao, and R. Tedrake, “Dexterous contact-rich manipulation via the contact trust region,”The International Journal of Robotics Research, p. 02783649251398875, 2026
2026
-
[40]
Z. Xie, Y . Xiang, M. Posa, and W. Jin, “Where to Touch, How to Contact: Hierarchical RL-MPC Framework for Geometry-Aware Long- Horizon Dexterous Manipulation,” 2026, arXiv:2601.10930 [cs]
Pith/arXiv arXiv 2026
-
[41]
Mujoco: A physics engine for model-based control,
E. Todorov, T. Erez, and Y . Tassa, “Mujoco: A physics engine for model-based control,” in2012 IEEE/RSJ International Conference on Intelligent Robots and Systems. IEEE, 2012, pp. 5026–5033
2012
-
[42]
Contact models in robotics: A comparative analysis,
Q. Lidec, W. Jallet, L. Montaut, I. Laptev, C. Schmid, and J. Car- pentier, “Contact models in robotics: A comparative analysis,”IEEE Transactions on Robotics, vol. 40, pp. 3716–3733, 01 2024. APPENDIXI TRAININGPARAMETERS We list the full configuration used for each task. TABLE I CONFIGURATION PARAMETERS, FINGER TASK. Parameter Value / Description iLQR ho...
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.