A temporal macro-action value factorization method with a segmented replay buffer improves asynchronous multi-agent RL performance, but the claimed proof that it generalizes standard IGM is invalid.
Improved Memory-Bounded Dynamic Programming for Decentralized POMDPs
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Memory-Bounded Dynamic Programming (MBDP) has proved extremely effective in solving decentralized POMDPs with large horizons. We generalize the algorithm and improve its scalability by reducing the complexity with respect to the number of observations from exponential to polynomial. We derive error bounds on solution quality with respect to this new approximation and analyze the convergence behavior. To evaluate the effectiveness of the improvements, we introduce a new, larger benchmark problem. Experimental results show that despite the high complexity of decentralized POMDPs, scalable solution techniques such as MBDP perform surprisingly well.
fields
cs.MA 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
ToMacVF : Temporal Macro-action Value Factorization for Asynchronous Multi-Agent Reinforcement Learning
A temporal macro-action value factorization method with a segmented replay buffer improves asynchronous multi-agent RL performance, but the claimed proof that it generalizes standard IGM is invalid.