Pith. sign in

REVIEW 1 cited by

Improved Memory-Bounded Dynamic Programming for Decentralized POMDPs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1206.5295 v1 pith:QJAD5ZGL submitted 2012-06-20 cs.AI

classification cs.AI
keywords decentralizedpomdpscomplexitydynamicmbdpmemory-boundedprogrammingrespect
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Memory-Bounded Dynamic Programming (MBDP) has proved extremely effective in solving decentralized POMDPs with large horizons. We generalize the algorithm and improve its scalability by reducing the complexity with respect to the number of observations from exponential to polynomial. We derive error bounds on solution quality with respect to this new approximation and analyze the convergence behavior. To evaluate the effectiveness of the improvements, we introduce a new, larger benchmark problem. Experimental results show that despite the high complexity of decentralized POMDPs, scalable solution techniques such as MBDP perform surprisingly well.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ToMacVF : Temporal Macro-action Value Factorization for Asynchronous Multi-Agent Reinforcement Learning

    cs.MA 2025-07 reject novelty 5.0 of 10

    A temporal macro-action value factorization method with a segmented replay buffer improves asynchronous multi-agent RL performance, but the claimed proof that it generalizes standard IGM is invalid.

Pith tools