Pith. sign in

REVIEW 26 cited by

Model Predictive Path Integral Control using Covariance Variable Importance Sampling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1509.01149 v3 pith:GBCEZUEZ submitted 2015-09-03 cs.SY cs.DCcs.ROcs.SY

classification cs.SYcs.DCcs.RO
keywords controlmodelpredictivesamplingalgorithmimportancediffusiongeneralized
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this paper we develop a Model Predictive Path Integral (MPPI) control algorithm based on a generalized importance sampling scheme and perform parallel optimization via sampling using a Graphics Processing Unit (GPU). The proposed generalized importance sampling scheme allows for changes in the drift and diffusion terms of stochastic diffusion processes and plays a significant role in the performance of the model predictive control algorithm. We compare the proposed algorithm in simulation with a model predictive control version of differential dynamic programming.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 26 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dream-MPC: Gradient-Based Model Predictive Control with Latent Imagination

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    Dream-MPC refines policy-generated trajectories by gradient ascent in a latent world model with uncertainty regularization and temporal amortization, improving base policy performance and beating gradient-free MPC on ...

  2. Stochastic Multiple Shooting Trajectory Optimization via Sequential Local Policy Evaluation

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A stochastic multiple-shooting optimizer that links short sampled control segments with local LQR feedback policies reaches terminal sets with fewer rollouts than MPPI and CEM on cartpole and VTOL landing benchmarks.

  3. Static In, Dynamic Out: Counterfactual Action Augmentation for Moving Object Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    SIDO morphs static demonstrations into counterfactual future-pose samples, training a goal-conditioned policy that, paired with a pose predictor, grasps objects whose motion was unseen during training.

  4. Adapting Generalist Vehicle Models for High-Speed MPC Across Terrains

    cs.RO 2026-07 conditional novelty 6.0 of 10

    History-conditioned fine-tuning with targeted synthetic rollouts from a per-terrain bicycle model roughly halves 6 m/s trajectory tracking error against a fine-tuned AnyCar baseline.

  5. Navigating the Crowd: Non-linear MPC with Social Forces Dynamics for Human-Aware Robot Navigation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Embedding Social Force Model dynamics and custom social costs inside nonlinear MPC yields joint human-robot trajectory prediction at 20 Hz and better social-compliance metrics than baselines in simulation.

  6. Natural Functional Gradients for Smooth Trajectory Optimization

    cs.RO 2026-05 unverdicted novelty 6.0 of 10

    A trajectory optimization method performs geometry-aware updates in function space via natural functional gradients and Monte-Carlo estimation on a smoothed surrogate objective to improve feasibility and smoothness in...

  7. Simultaneous Contact Selection and Planning for Contact-Rich Manipulation with Cascaded Optimization

    cs.RO 2026-05 unverdicted novelty 6.0 of 10

    SCSP is a cascaded optimization framework using a surrogate contact model and discrete-continuous search to enable simultaneous contact selection and planning for robust contact-rich manipulation.

  8. Slot-MPC: Goal-Conditioned Model Predictive Control with Object-Centric Representations

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Slot-MPC learns slot representations to build a differentiable object-centric dynamics model that supports efficient gradient-based MPC for robotic manipulation in novel situations.

  9. Model Predictive Path Integral Control as Preconditioned Gradient Descent

    math.OC 2026-03 unverdicted novelty 6.0 of 10

    MPPI is recovered as unit-step preconditioned gradient descent on a reduced free-energy objective over parametric sampling distributions, with descent guarantees when the preconditioned Hessian is bounded.

  10. Feedback Matters: Augmenting Autonomous Dissection with Visual and Topological Feedback

    cs.RO 2025-10 conditional novelty 6.0 of 10

    A stretch-based tissue connectivity estimator plus an exposure-maximizing controller and recovery planner raised autonomous dissection success on a da Vinci robot to 80%.

  11. Empowering Multi-Robot Cooperation via Sequential World Models

    cs.RO 2025-09 unverdicted novelty 6.0 of 10

    SeqWM introduces sequential autoregressive agent-wise world models for multi-robot MBRL, outperforming baselines in performance and sample efficiency on Bi-DexHands and Multi-Quadruped tasks with physical robot deployment.

  12. Continual Reinforcement Learning by Planning with Online World Models

    cs.LG 2025-07 conditional novelty 6.0 of 10

    An online follow-the-leader world model updated by analytic least squares, combined with CEM planning, solves sequential robotic tasks without forgetting on a new unified-dynamics benchmark.

  13. Robust Reward Alignment via Hypothesis Space Batch Cutting

    cs.LG 2025-02 conditional novelty 6.0 of 10

    HSBC learns rewards from preference batches by conservatively cutting the hypothesis space with a voting threshold, tolerating erroneous labels up to a bound set by a parameter gamma.

  14. GPD: Guided Polynomial Diffusion for Motion Planning

    cs.RO 2025-01 conditional novelty 6.0 of 10

    Guided diffusion over Bernstein polynomial coefficients generates smooth, collision-free manipulator trajectories with fewer denoising steps than waypoint-space diffusion.

  15. EVaDE : Event-Based Variational Thompson Sampling for Model-Based Reinforcement Learning

    cs.LG 2025-01 conditional novelty 6.0 of 10

    EVaDE inserts three Gaussian-dropout convolutional layers into SimPLe reward models, raising mean human-normalized Atari 100K score from 0.525 to 0.682 in the paper's runs.

  16. TD-MPC2: Scalable, Robust World Models for Continuous Control

    cs.LG 2023-10 conditional novelty 6.0 of 10

    TD-MPC2 scales an implicit world-model RL method to a 317M-parameter agent that masters 80 tasks across four domains with a single hyperparameter configuration.

  17. Is Conditional Generative Modeling all you need for Decision-Making?

    cs.LG 2022-11 unverdicted novelty 6.0 of 10

    Return-conditional diffusion models for policies outperform offline RL on benchmarks by circumventing dynamic programming and enable constraint or skill composition.

  18. RoboNet: Large-Scale Multi-Robot Learning

    cs.RO 2019-10 conditional novelty 6.0 of 10

    RoboNet is a multi-robot video dataset that enables pre-training of vision-based manipulation models which, after fine-tuning on a new robot, outperform robot-specific training that uses 4-20 times more data.

  19. Post-Hoc Robustness for Model-Based Reinforcement Learning

    cs.LG 2026-06 unverdicted novelty 5.0 of 10

    Introduces post-hoc robustification of model-based RL agents via adversarial model-predictive control at inference time to improve robustness in perturbed environments without additional neural network training.

  20. Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    MBDPO reformulates policy optimization as a diffusion process over searched trajectories in latent world models to reduce misalignment between search and value learning.

  21. WestWorld: A Knowledge-Encoded Scalable Trajectory World Model for Diverse Robotic Systems

    cs.LG 2026-03 unverdicted novelty 5.0 of 10

    WestWorld introduces a scalable trajectory world model with Sys-MoE routing via system embeddings and structural embeddings for physical knowledge, pretrained on 89 environments to improve zero-shot prediction and rea...

  22. D2 Actor Critic: Diffusion Actor Meets Distributional Critic

    cs.LG 2025-10 unverdicted novelty 5.0 of 10

    D2AC combines a diffusion actor with a distributional critic via fused distributional RL and clipped double Q-learning to reach state-of-the-art results on 18 hard control benchmarks including Humanoid, Dog, and Shadow Hand.

  23. Equivariant Action Sampling for Reinforcement Learning and Planning

    cs.RO 2024-12 conditional novelty 5.0 of 10

    Augmenting each sampled action with its full symmetry orbit makes finite-sample planning exactly equivariant and speeds up learning on several rotationally symmetric control tasks.

  24. Critique of World Model

    cs.LG 2025-07 conditional novelty 4.0 of 10

    The paper argues world models should simulate actionable possibilities and proposes GLP, a hierarchical generative architecture that closes the loop with observation reconstruction, but provides no experiments.

  25. M3PO: Massively Multi-Task Model-Based Policy Optimization

    cs.LG 2025-06 reject novelty 4.0 of 10

    A hybrid model-based RL algorithm combining TDMPC2-style world models, PPO clipped updates, and POME-style exploration bonuses reports strong benchmark results but omits reproducible evidence.

  26. Reinforcement Learning: From Algorithms To Foundation Models

    cs.AI 2026-07 conditional novelty 3.0 of 10

    A dissertation uniting the author's published results: non-exploitable Nash-DQN policies and the FightLadder benchmark for games, plus diffusion/consistency-model world models for RL — a compilation rather than new results.

Pith tools