REVIEW 26 cited by
Model Predictive Path Integral Control using Covariance Variable Importance Sampling
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In this paper we develop a Model Predictive Path Integral (MPPI) control algorithm based on a generalized importance sampling scheme and perform parallel optimization via sampling using a Graphics Processing Unit (GPU). The proposed generalized importance sampling scheme allows for changes in the drift and diffusion terms of stochastic diffusion processes and plays a significant role in the performance of the model predictive control algorithm. We compare the proposed algorithm in simulation with a model predictive control version of differential dynamic programming.
Forward citations
Cited by 26 Pith papers
-
Dream-MPC: Gradient-Based Model Predictive Control with Latent Imagination
Dream-MPC refines policy-generated trajectories by gradient ascent in a latent world model with uncertainty regularization and temporal amortization, improving base policy performance and beating gradient-free MPC on ...
-
Stochastic Multiple Shooting Trajectory Optimization via Sequential Local Policy Evaluation
A stochastic multiple-shooting optimizer that links short sampled control segments with local LQR feedback policies reaches terminal sets with fewer rollouts than MPPI and CEM on cartpole and VTOL landing benchmarks.
-
Static In, Dynamic Out: Counterfactual Action Augmentation for Moving Object Manipulation
SIDO morphs static demonstrations into counterfactual future-pose samples, training a goal-conditioned policy that, paired with a pose predictor, grasps objects whose motion was unseen during training.
-
Adapting Generalist Vehicle Models for High-Speed MPC Across Terrains
History-conditioned fine-tuning with targeted synthetic rollouts from a per-terrain bicycle model roughly halves 6 m/s trajectory tracking error against a fine-tuned AnyCar baseline.
-
Navigating the Crowd: Non-linear MPC with Social Forces Dynamics for Human-Aware Robot Navigation
Embedding Social Force Model dynamics and custom social costs inside nonlinear MPC yields joint human-robot trajectory prediction at 20 Hz and better social-compliance metrics than baselines in simulation.
-
Natural Functional Gradients for Smooth Trajectory Optimization
A trajectory optimization method performs geometry-aware updates in function space via natural functional gradients and Monte-Carlo estimation on a smoothed surrogate objective to improve feasibility and smoothness in...
-
Simultaneous Contact Selection and Planning for Contact-Rich Manipulation with Cascaded Optimization
SCSP is a cascaded optimization framework using a surrogate contact model and discrete-continuous search to enable simultaneous contact selection and planning for robust contact-rich manipulation.
-
Slot-MPC: Goal-Conditioned Model Predictive Control with Object-Centric Representations
Slot-MPC learns slot representations to build a differentiable object-centric dynamics model that supports efficient gradient-based MPC for robotic manipulation in novel situations.
-
Model Predictive Path Integral Control as Preconditioned Gradient Descent
MPPI is recovered as unit-step preconditioned gradient descent on a reduced free-energy objective over parametric sampling distributions, with descent guarantees when the preconditioned Hessian is bounded.
-
Feedback Matters: Augmenting Autonomous Dissection with Visual and Topological Feedback
A stretch-based tissue connectivity estimator plus an exposure-maximizing controller and recovery planner raised autonomous dissection success on a da Vinci robot to 80%.
-
Empowering Multi-Robot Cooperation via Sequential World Models
SeqWM introduces sequential autoregressive agent-wise world models for multi-robot MBRL, outperforming baselines in performance and sample efficiency on Bi-DexHands and Multi-Quadruped tasks with physical robot deployment.
-
Continual Reinforcement Learning by Planning with Online World Models
An online follow-the-leader world model updated by analytic least squares, combined with CEM planning, solves sequential robotic tasks without forgetting on a new unified-dynamics benchmark.
-
Robust Reward Alignment via Hypothesis Space Batch Cutting
HSBC learns rewards from preference batches by conservatively cutting the hypothesis space with a voting threshold, tolerating erroneous labels up to a bound set by a parameter gamma.
-
GPD: Guided Polynomial Diffusion for Motion Planning
Guided diffusion over Bernstein polynomial coefficients generates smooth, collision-free manipulator trajectories with fewer denoising steps than waypoint-space diffusion.
-
EVaDE : Event-Based Variational Thompson Sampling for Model-Based Reinforcement Learning
EVaDE inserts three Gaussian-dropout convolutional layers into SimPLe reward models, raising mean human-normalized Atari 100K score from 0.525 to 0.682 in the paper's runs.
-
TD-MPC2: Scalable, Robust World Models for Continuous Control
TD-MPC2 scales an implicit world-model RL method to a 317M-parameter agent that masters 80 tasks across four domains with a single hyperparameter configuration.
-
Is Conditional Generative Modeling all you need for Decision-Making?
Return-conditional diffusion models for policies outperform offline RL on benchmarks by circumventing dynamic programming and enable constraint or skill composition.
-
RoboNet: Large-Scale Multi-Robot Learning
RoboNet is a multi-robot video dataset that enables pre-training of vision-based manipulation models which, after fine-tuning on a new robot, outperform robot-specific training that uses 4-20 times more data.
-
Post-Hoc Robustness for Model-Based Reinforcement Learning
Introduces post-hoc robustification of model-based RL agents via adversarial model-predictive control at inference time to improve robustness in perturbed environments without additional neural network training.
-
Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization
MBDPO reformulates policy optimization as a diffusion process over searched trajectories in latent world models to reduce misalignment between search and value learning.
-
WestWorld: A Knowledge-Encoded Scalable Trajectory World Model for Diverse Robotic Systems
WestWorld introduces a scalable trajectory world model with Sys-MoE routing via system embeddings and structural embeddings for physical knowledge, pretrained on 89 environments to improve zero-shot prediction and rea...
-
D2 Actor Critic: Diffusion Actor Meets Distributional Critic
D2AC combines a diffusion actor with a distributional critic via fused distributional RL and clipped double Q-learning to reach state-of-the-art results on 18 hard control benchmarks including Humanoid, Dog, and Shadow Hand.
-
Equivariant Action Sampling for Reinforcement Learning and Planning
Augmenting each sampled action with its full symmetry orbit makes finite-sample planning exactly equivariant and speeds up learning on several rotationally symmetric control tasks.
-
Critique of World Model
The paper argues world models should simulate actionable possibilities and proposes GLP, a hierarchical generative architecture that closes the loop with observation reconstruction, but provides no experiments.
-
M3PO: Massively Multi-Task Model-Based Policy Optimization
A hybrid model-based RL algorithm combining TDMPC2-style world models, PPO clipped updates, and POME-style exploration bonuses reports strong benchmark results but omits reproducible evidence.
-
Reinforcement Learning: From Algorithms To Foundation Models
A dissertation uniting the author's published results: non-exploitable Nash-DQN policies and the FightLadder benchmark for games, plus diffusion/consistency-model world models for RL — a compilation rather than new results.
Discussion (0). Continue with ORCID to comment.