MODIP fine-tunes diffusion policies offline-to-online by training a world model, running MPC with terminal state values inside it to create targets, and using policy-independent TD critics, yielding gains over BC on D4RL and RoboMimic tasks.
Space odyssey: An experimental software security analysis of satellites
3 Pith papers cite this work. Polarity classification is still indexing.
3
Pith papers citing it
representative citing papers
Adapts PACE methodology to satellites with a graph-based state-transition model and dynamic resilience index, evaluating static, adaptive, and epsilon-greedy variants for improved survivability.
citing papers explorer
-
MODIP: Efficient Model-Based Optimization for Diffusion Policies
MODIP fine-tunes diffusion policies offline-to-online by training a world model, running MPC with terminal state values inside it to create targets, and using policy-independent TD critics, yielding gains over BC on D4RL and RoboMimic tasks.
-
Resilience Through Escalation: A Graph-Based PACE Architecture for Satellite Threat Response
Adapts PACE methodology to satellites with a graph-based state-transition model and dynamic resilience index, evaluating static, adaptive, and epsilon-greedy variants for improved survivability.
- Lagrangian Perturbation Diffusion Steering: Latent Reinforcement Learning for Generative Policies