Pith. sign in

REVIEW 2 cited by

Contrasting Centralized and Decentralized Critics in Multi-Agent Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.04402 v2 pith:Y3OPU3MW submitted 2021-02-08 cs.LG cs.AI

classification cs.LGcs.AI
keywords centralizeddecentralizedcriticcriticschoiceimplicationslearningmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Centralized Training for Decentralized Execution, where agents are trained offline using centralized information but execute in a decentralized manner online, has gained popularity in the multi-agent reinforcement learning community. In particular, actor-critic methods with a centralized critic and decentralized actors are a common instance of this idea. However, the implications of using a centralized critic in this context are not fully discussed and understood even though it is the standard choice of many algorithms. We therefore formally analyze centralized and decentralized critic approaches, providing a deeper understanding of the implications of critic choice. Because our theory makes unrealistic assumptions, we also empirically compare the centralized and decentralized critic methods over a wide set of environments to validate our theories and to provide practical advice. We show that there exist misconceptions regarding centralized critics in the current literature and show that the centralized critic design is not strictly beneficial, but rather both centralized and decentralized critics have different pros and cons that should be taken into account by algorithm designers.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 36 citations worldwide. Full citation record

  1. Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic

    cs.AI 2026-01 unverdicted novelty 6.0 of 10

    Multi-agent actor-critic methods with a centralized critic improve decentralized LLM collaboration over Monte Carlo baselines in long-horizon and sparse-reward settings.

  2. PAGNet: Pluggable Adaptive Generative Networks for Information Completion in Multi-Agent Communication

    cs.MA 2025-02 conditional novelty 5.0 of 10

    PAGNet learns per-agent communication weights and generates global states with a U-Net and GAN discriminator, reporting improved cooperative MARL performance on LBF, Hallway, and SMAC.

Pith tools