DTWIL creates a surrogate demonstration set by aligning expert and constrained state trajectories through model predictive control and dynamic time warping, then trains an imitator on it.
State Alignment-based Imitation Learning
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Consider an imitation learning problem that the imitator and the expert have different dynamics models. Most of the current imitation learning methods fail because they focus on imitating actions. We propose a novel state alignment-based imitation learning method to train the imitator to follow the state sequences in expert demonstrations as much as possible. The state alignment comes from both local and global perspectives and we combine them into a reinforcement learning framework by a regularized policy update objective. We show the superiority of our method on standard imitation learning settings and imitation learning settings where the expert and imitator have different dynamics models.
fields
cs.RO 1years
2025 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Action-Constrained Imitation Learning
DTWIL creates a surrogate demonstration set by aligning expert and constrained state trajectories through model predictive control and dynamic time warping, then trains an imitator on it.