A distributed homotopy primal-dual algorithm for multi-agent TD learning is proved to converge at O(log^2 T / T) under Markovian sampling, improving on the prior O(1/sqrt(T)) bound for GTD-type methods.
Residual algorithms: reinforcement learning with function approximation,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
math.OC 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Fast Multi-Agent Temporal-Difference Learning via Homotopy Stochastic Primal-Dual Optimization
A distributed homotopy primal-dual algorithm for multi-agent TD learning is proved to converge at O(log^2 T / T) under Markovian sampling, improving on the prior O(1/sqrt(T)) bound for GTD-type methods.