PI-VM trains a value network with TD updates on a recursive path integral target, giving faster and more stable stochastic optimal control than policy-based baselines.
Stochastic optimal control matching.Advances in Neural Information Processing Systems, 37:112459–112504, 2024
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Path Integral Value Matching for Linear Quadratic Stochastic Optimal Control
PI-VM trains a value network with TD updates on a recursive path integral target, giving faster and more stable stochastic optimal control than policy-based baselines.