Pith. sign in

REVIEW 5 major objections 5 minor 49 references

Revisiting Communication Efficiency in Multi-Agent Reinforcement Learning from the Dimensional Analysis Perspective

T0 review · 5 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Dimensional redundancy and confounders persist in receiver-side message embeddings after sender-side gating, and DRMAC removes them to improve multi-agent communication.

desk verdict Plausible receiver-side complement to sender-side MARL communication, with an honest motivation and clean modules, but the evidence needs error bars, multi-map ablations, and a check that decorrelation preserves task-relevant information. read the letter →

arxiv 2501.02888 v2 pith:YP2MSC3Y submitted 2025-01-06 cs.MA

classification cs.MA
keywords Multi-AgentReinforcementLearningCommunicationEfficiencyDimensionalAnalysisRedundancyReductionRepresentationCooperativeSystemsConfounders
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that communication efficiency in multi-agent reinforcement learning is not fully solved by optimizing what and when agents send. Even after gating and content optimization, the integrated message embeddings formed at the receiving end still carry redundant dimensions and decision-irrelevant confounders that degrade decisions. The authors introduce DRMAC, which decorrelates the dimensions of these embeddings with a redundancy-reduction regularization term and learns a per-agent dimensional mask, via meta-learning, to down-weight confounding dimensions. They report consistent performance gains over existing communication methods on Hallway, Hallway-Group, and StarCraft tasks, and show that DRMAC's modules can be plugged into other MARL and communication algorithms to improve them. If true, communication efficiency can be improved at the receiving end rather than solely at the sender.

What carries the argument

The central mechanism is a redundancy-reduction regularizer applied to the cross-correlation matrix of two views of the same integrated message embedding, together with an Information Selective Network (ISN) that produces a dimensional mask. The regularizer, equation (1), pushes the diagonal of the cross-correlation matrix toward 1 and the off-diagonal entries toward 0, so each dimension of the embedding captures distinct information; the twin views come from the online integration encoder and a noise-perturbed copy of it. The ISN takes the agent's own observation as input and outputs a mask with values in [0,1], applied to the embedding by element-wise multiplication, and is trained through a meta-learning step using second-order gradients of the RL loss so it can adapt to which dimensions matter for the current decision.

What would settle it

On a diagnostic task where the optimal policy is known to depend on a specific, identifiable dimension of the integrated message embedding, train DRMAC while varying the noise scale used to create the twin view or removing the redundancy term; if performance collapses when that dimension is decorrelated or down-weighted, the assumption that redundancy reduction preserves decision-relevant information is contradicted.

Watch

Extended reading notes

Core claim

The central discovery is that, even after optimized and gated messages are sent, dimensional redundancy and confounders persist in the integrated message embeddings at the receiving end, and these negatively affect communication quality and decision-making. DRMAC mitigates both: a redundancy-reduction regularization term drives the cross-correlation matrix of twin message embeddings toward the identity matrix, decoupling the information carried by each dimension, and an Information Selective Network learns a dimensional mask that dynamically adjusts gradient weights so the policy focuses on decision-relevant dimensions while suppressing confounders. The paper reports that DRMAC consistently outperforms state-of-the-art communication baselines across a diverse set of cooperative multi-agent tasks, and that its key modules serve as a plug-and-play complement to existing communication strategies rather than a replacement.

Load-bearing premise

The load-bearing premise is that forcing the cross-correlation of two noise-corrupted views of the same integrated message toward the identity removes redundant dimensions without also discarding information the policy needs; if the Gaussian noise augmentation or the decorrelation destroys task-critical features, DRMAC's performance gains would disappear.

Editorial extensions

If this is right

  • Sender-side message filtering of content and timing alone does not resolve communication efficiency; receiver-side dimensional analysis is a necessary complementary step.
  • Decorrelation of message-embedding dimensions plus masking of decision-irrelevant dimensions improves cooperative performance across tasks with different numbers of agents and difficulty levels.
  • DRMAC's modules are plug-and-play: adding them to different MARL training algorithms (VDN, QMIX, MAPPO) and to communication methods (MASIA, SMS, TarMAC) improves those methods' performance.
  • Ablation shows that the redundancy-reduction term and the ISN mask are mutually reinforcing: removing either one degrades performance.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The motivating evidence that randomly masking dimensions sometimes improves performance suggests that a substantial fraction of embedding dimensions are actively harmful rather than merely redundant; a natural follow-up is to measure how that fraction changes with task complexity and message length.
  • The redundancy-reduction objective is adapted from self-supervised learning, so a targeted ablation of the Gaussian noise injection and the projection layer could reveal whether the invariance term, rather than the decorrelation, is the main driver of the gains.
  • The ISN is trained only by RL performance while the encoder is trained with the redundancy loss; one extension would be to let the mask influence which dimensions are decorrelated, instead of assuming all off-diagonal correlations are equally harmful.
  • Because the evaluation is limited to cooperative tasks with shared rewards, it remains open whether dimensional confounders play the same role in competitive or mixed-motive settings, where a receiver may want to distrust certain message dimensions.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper argues that in cooperative MARL, sender-side optimization of communication content and timing is insufficient because the receiver-side integrated message embeddings still exhibit dimensional redundancy and dimensional confounders. To address this, the authors propose DRMAC, which adds a Barlow-Twins-style redundancy-reduction regularization term on two views of integrated messages and an Information Selective Network that learns a per-agent dimensional mask via meta-learning. The method is evaluated on Hallway and SMAC tasks against TarMAC, MAIC, SMS, MASIA, and QMIX, and is also applied as a plug-and-play module on top of several baselines. The central claim is that DRMAC consistently outperforms existing methods by removing receiver-side dimensional redundancy and confounders while preserving decision-relevant information.

Significance. If the central claim is validated, DRMAC offers a useful complementary perspective to sender-side communication-efficiency methods and provides a simple, plug-and-play module that can be integrated with existing MARL algorithms. The motivating experiments in Figures 2 and 3 are a strength: they give a concrete demonstration that residual redundancy and harmful dimensions persist after sender-side optimization. The plug-and-play experiments in Figure 8 are also valuable because they show compatibility with multiple learning algorithms and communication baselines. However, the current empirical support is not yet sufficient: the headline comparisons are learning curves without error bars or significance tests, the component ablation is performed on a single SMAC map, and there is no direct evidence that decorrelation preserves task-relevant information. The contribution is potentially interesting, but the load-bearing claims require additional support before the paper can be accepted.

major comments (5)
  1. [Section 5.2, Appendix B] The headline claim that DRMAC 'consistently outperforms the baselines in almost all environments' is supported only by learning curves. Appendix B states that results are averaged over 5 random seeds, but no variance, error bars, numerical table, or significance test is reported anywhere in Section 5.2 or Figure 6. As a result, the reader cannot assess whether the observed advantages are statistically reliable or within run-to-run noise. Please add a quantitative summary (e.g., final mean and standard deviation per task) and, where practical, statistical comparisons for at least the main environments.
  2. [Section 4.1, Eq. (1)] The redundancy-reduction objective follows Barlow Twins, whose rationale assumes the two views share semantic content. Here the two views are produced by feeding the identical message through the online IIE and through a copy with Gaussian noise added to its weights; the invariance term therefore enforces invariance to weight perturbation, not to a content-preserving input augmentation. An encoder can satisfy Eq. (1) by becoming insensitive to its own parameter noise, which does not guarantee that decision-relevant information is retained. The paper provides no probe or information-theoretic check that task-relevant content survives in the decorrelated embedding, so the observed performance gains cannot be attributed specifically to redundancy removal as opposed to an incidental regularizing effect. Please add a measuring experiment (e.g., predicting the global state or the optimal action from the masked representation) or an ablation with a semantic augmentation.
  3. [Section 5.3, Fig. 7] The ablation study evaluates the two components on a single SMAC map, 1o_2r_vs_4r. The conclusion that redundancy reduction and ISN are mutually reinforcing is therefore based on one scenario; please extend the ablations to at least one additional task (for instance a Hallway variant or another SMAC map) to support the claimed generality.
  4. [Section 4.3, Eqs. (5)-(6)] The description of the ISN meta-learning update is ambiguous: the text says 'back-propagation of gradients is excluded' when computing the trial weights, yet immediately concludes that θ_ISN 'is refined through second-order gradient optimization.' If the trial weights are detached from the computation graph, the resulting meta-gradient is first-order; if they are not detached, the statement about excluding back-propagation is misleading. Please clarify exactly which computation graph is used and which order of derivatives is taken, since this determines both the optimization behavior and the reproducibility of the method.
  5. [Section 4.2, Section 5.3] The paper does not provide direct evidence that the learned ISN mask corresponds to decision-irrelevant dimensions. No analysis of the learned mask values, no comparison with masking the same number of random dimensions, and no visualization of which dimensions are suppressed is reported. Without such evidence, the improvement of DRMAC over DRMAC w/o ISN might reflect a generic input-dependent feature weighting rather than the removal of dimensional confounders, which is one of the paper's two central mechanisms.
minor comments (5)
  1. [Section 3] The Dec-POMDP tuple includes a message set M, but the formal formulation does not describe how messages are generated or how they enter the observation/reward structure; please make this explicit.
  2. [Section 4.3] The parameter set θ includes θ_IEE, while the encoder is abbreviated IIE in Section 4.1; please unify the notation.
  3. [Section 5.1, Appendix B] Section 5.1 states 'two well-known cooperative multi-agent environments, encompassing a total of eight tasks,' while Appendix B begins with 'The four test environments,' which is inconsistent.
  4. [Section 4.1, Eq. (2)] The normalization in Eq. (2) lacks an epsilon term for numerical stability; please add one or explain how zero-variance dimensions are handled.
  5. [Appendix A] The paper says to refer to Appendix A for implementation details, but Appendix A only reports DRMAC's architecture and hyperparameters, not a complete training recipe or the baseline configurations; please add a reproducibility statement or release code.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the decorrelation and masking modules are trained objectives, and the motivating baseline evidence is independent of DRMAC's construction.

full rationale

The paper's central claim is that receiver-side message embeddings retain dimensional redundancy and confounders even after sender-side filtering, and that DRMAC's redundancy-reduction loss and information-selective mask improve communication. This claim does not reduce to its own inputs by the paper's equations. The motivating evidence in Figs. 2 and 3 is obtained from the baseline methods MASIA and SMS before DRMAC is introduced: representations are visualized for dimensional similarity, and random dimensional masking is shown to sometimes improve baseline performance. Those observations are independent of the proposed objective. The method is a trained objective: Eq. (1) directly drives the cross-correlation matrix toward the identity, so the later demonstration that DRMAC's representation is less redundant is a verification of the training target, not a prediction made from a fitted parameter. Similarly, the ISN dimensional mask is learned via meta-learning on the RL loss in Eqs. (4)-(6), so its contribution to performance is the normal behavior of a trained module, and the ablation in Sec. 5.3 compares trained variants with a baseline rather than renaming a fitted value as a prediction. The only self-citation, reference [32] (IMMAC), appears in the related-work survey of when-to-communicate methods; it is not load-bearing for DRMAC's design or results, and no uniqueness theorem or external forcing is imported from the authors' prior work. The transfer of Barlow Twins [45] and the choice of Gaussian weight-noise augmentation are methodological assumptions, but they are not circular because the paper does not define the target result in terms of those assumptions. Experimental comparisons use official implementations of external baselines (TarMAC, MAIC, SMS, MASIA, QMIX), so the empirical support is not internal to a fitted quantity. No step satisfies the requirement of exhibiting a specific reduction of a prediction to an input by construction.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The central claim rests on the transfer of self-supervised redundancy reduction to RL message embeddings, the validity of the Gaussian-noise augmentation, and the stability of the meta-learned mask. These are assumptions external to the paper, not derived results. The only hand-set hyperparameters directly tied to the proposed losses are beta and lambda; the Gaussian noise scale and meta-learning rate are not reported.

free parameters (4)
  • beta = 0.01
    Weight balancing RL loss and redundancy reduction loss in L_tot, chosen by hand (Appendix A.2).
  • lambda = 0.0051
    Weight on off-diagonal terms of the cross-correlation objective in Eq. (1), chosen by hand (Appendix A.2).
  • Gaussian noise scale for target encoder
    Scale of noise injected into the copied IIE to create twin views is not reported; this augmentation strength affects L_RR in Section 4.1.
  • Meta-learning learning rate for trial weights
    Used in Eq. (6) for the one-step trial weights; value not reported in Appendix A.
assumptions (5)
  • standard math Tasks are Dec-POMDPs with shared rewards and messages exchanged among agents (Section 3).
    Standard formal model for cooperative MARL; not a contribution.
  • domain assumption Factorial coding, i.e., decorrelating representation dimensions, improves downstream task performance (Section 4.1, Eq. 1).
    Borrowed from Barlow Twins [45] and redundancy-reduction neuroscience [3]; transferred to RL message embeddings without proof.
  • domain assumption Adding Gaussian noise to a copy of the Information Integration Encoder produces valid augmented views of the same message (Section 4.1, Fig. 4).
    Augmentation choice is standard in SSL but not justified for message embeddings.
  • ad hoc to paper A one-step-unrolled, second-order meta-gradient update of the ISN mask reliably tracks which dimensions are decision-relevant in a non-stationary MARL setting (Section 4.3, Eqs. 5-6).
    This bi-level optimization scheme is proposed specifically for DRMAC and its stability is not analyzed.
  • ad hoc to paper Per-dimension multiplicative masks bounded in [0,1] can remove confounding information without hurting the policy (Section 4.2, Eq. 3).
    The mask is trained end-to-end on RL return, so this is a learnable assumption rather than a proven structural property.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Revisiting Communication Efficiency in Multi-Agent Reinforcement Learning from the Dimensional Analysis Perspective." pith.science (2026). https://pith.science/paper/YP2MSC3Y

@misc{pith2026250102888,
  author       = {Pith},
  title        = {Pith review of: Revisiting Communication Efficiency in Multi-Agent Reinforcement Learning from the Dimensional Analysis Perspective},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YP2MSC3Y}},
  note         = {Machine review of arXiv:2501.02888}
}
read the original abstract

In this work, we introduce a novel perspective, i.e., dimensional analysis, to address the challenge of communication efficiency in Multi-Agent Reinforcement Learning (MARL). Our findings reveal that simply optimizing the content and timing of communication at sending end is insufficient to fully resolve communication efficiency issues. Even after applying optimized and gated messages, dimensional redundancy and confounders still persist in the integrated message embeddings at receiving end, which negatively impact communication quality and decision-making. To address these challenges, we propose Dimensional Rational Multi-Agent Communication (DRMAC), designed to mitigate both dimensional redundancy and confounders in MARL. DRMAC incorporates a redundancy-reduction regularization term to encourage the decoupling of information across dimensions within the learned representations of integrated messages. Additionally, we introduce a dimensional mask that dynamically adjusts gradient weights during training to eliminate the influence of decision-irrelevant dimensions. We evaluate DRMAC across a diverse set of multi-agent tasks, demonstrating its superior performance over existing state-of-the-art methods in complex scenarios. Furthermore, the plug-and-play nature of DRMAC's key modules highlights its generalizable performance, serving as a valuable complement rather than a replacement for existing multi-agent communication strategies.

Figures

Figures reproduced from arXiv: 2501.02888 by the authors.

Figure 1
Figure 1. Existing communication methods typically focus [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. The visualizations depict the representations learned by SMS[ [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Experimental scatter plots were generated by MASIA and SMS with randomly masked dimensions on a challenging [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: The overview of our proposed DRMAC. embeddings to mitigate overlapping information and improve com￾munication efficiency. Concretely, for agent𝑖, the received messages are first fed into the Information Integration Encoder to obtain its representation 𝑧 𝑡 𝑖 . To obtain…
Figure 5
Figure 5. Figure 5: Multiple environments considered in our experiments. [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: Performance on multiple tasks. • RQ2. Which specific components of DRMAC are crucial to its overall performance? • RQ3. Can DRMAC be effectively integrated with various existing communication baselines to enhance their commu￾nication efficiency? 5.1 Setup Baselines To …
Figure 7
Figure 7. Figure 7: Performance of different DRMAC variants. [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: Performance of DRMAC integrated with various [PITH_FULL_IMAGE:figures/full_fig_p009_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

49 extracted references · 33 canonical work pages

  1. [1]

    OpenAI: Marcin Andrychowicz, Bowen Baker, Maciek Chociej, Rafal Jozefowicz, Bob McGrew, Jakub Pachocki, Arthur Petron, Matthias Plappert, Glenn Powell, Alex Ray, et al. 2020. Learning dexterous in-hand manipulation. The International Journal of Robotics Research 39, 1 (2020), 3–20

  2. [2]

    Devon Hjelm, and William Buchwalter

    Philip Bachman, R. Devon Hjelm, and William Buchwalter. 2019. Learning Repre- sentations by Maximizing Mutual Information Across Views. InAdvances in Neu- ral Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, Hanna M. Wallach, Hugo Larochelle, Alina...

  3. [3]

    Horace B Barlow et al. 1961. Possible principles underlying the transformation of sensory messages. Sensory communication 1, 01 (1961), 217–233

  4. [4]

    Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. 2020. Unsupervised Learning of Visual Features by Contrasting Cluster Assignments. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, Hugo Larochell...

  5. [5]

    Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey E. Hinton. 2020. A Simple Framework for Contrastive Learning of Visual Representations. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event (Proceedings of Machine Learning Research, Vol. 119) . PMLR, 1597–1607. http://proceedings.mlr....

  6. [6]

    Ching-Yao Chuang, Joshua Robinson, Yen-Chen Lin, Antonio Torralba, and Stefanie Jegelka. 2020. Debiased Contrastive Learning. In Advances in Neu- ral Information Processing Systems 33: Annual Conference on Neural Infor- mation Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual , Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-F...

  7. [7]

    Abhishek Das, Théophile Gervet, Joshua Romoff, Dhruv Batra, Devi Parikh, Mike Rabbat, and Joelle Pineau. 2019. Tarmac: Targeted multi-agent communication. In International Conference on Machine Learning . 1538–1546

  8. [8]

    Ziluo Ding, Tiejun Huang, and Zongqing Lu. 2020. Learning individually inferred communication for multi-agent cooperation. Advances in Neural Information Processing Systems 33 (2020), 22069–22079

Show all 49 references
  1. [9]

    Aleksandr Ermolov, Aliaksandr Siarohin, Enver Sangineto, and Nicu Sebe. 2021. Whitening for Self-Supervised Representation Learning. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event (Proceedings of Machine Learn...

  2. [10]

    Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Ávila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Rémi Munos, and Michal Valko

    Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H. Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Ávila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Rémi Munos, and Michal Valko. 2020. Bootstrap Your Own Latent -...

  3. [11]

    Cong Guan, Feng Chen, Lei Yuan, Chenghe Wang, Hao Yin, Zongzhang Zhang, and Yang Yu. 2022. Efficient Multi-agent Communication via Self-supervised Information Aggregation. Advances in Neural Information Processing Systems 35 (2022), 1020–1033

  4. [12]

    Girshick

    Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross B. Girshick. 2020. Momentum Contrast for Unsupervised Visual Representation Learning. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, W A, USA, June 13-19, 2020. Computer Vision ...

  5. [13]

    Jiechuan Jiang and Zongqing Lu. 2018. Learning attentional communication for multi-agent cooperation. In Advances in neural information processing systems . 7254–7264

  6. [14]

    Daewoo Kim, Sangwoo Moon, David Hostallero, Wan Ju Kang, Taeyoung Lee, Kyunghwan Son, and Yung Yi. 2019. Learning to schedule communication in multi-agent reinforcement learning. arXiv preprint arXiv:1902.01554 (2019)

  7. [15]

    Michael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto, Pieter Abbeel, and Ar- avind Srinivas. 2020. Reinforcement Learning with Augmented Data. InAdvances in Neural Information Processing Systems 33: Annual Conference on Neural In- formation Processing Systems 2020, NeurIPS 202...

  8. [16]

    Michael Laskin, Aravind Srinivas, and Pieter Abbeel. 2020. CURL: Contrastive Unsupervised Representations for Reinforcement Learning. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event (Proceedings of Machine Lear...

  9. [17]

    Edouard Leurent. 2018. A survey of state-action representations for autonomous driving. (2018)

  10. [18]

    Shikun Liu, Andrew Davison, and Edward Johns. 2019. Self-supervised generali- sation with meta auxiliary learning. Advances in Neural Information Processing Systems 32 (2019)

  11. [19]

    Ryan Lowe, Yi I Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch. 2017. Multi-agent actor-critic for mixed cooperative-competitive environments. In Advances in neural information processing systems . 6379–6390

  12. [20]

    Yaru Niu, Rohan R Paleja, and Matthew C Gombolay. 2021. Multi-Agent Graph- Attention Communication and Teaming.. In AAMAS. 964–973

  13. [21]

    Frans A Oliehoek, Christopher Amato, et al . 2016. A concise introduction to decentralized POMDPs. Vol. 1. Springer

  14. [22]

    Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy. 2016. Deep exploration via bootstrapped DQN. In Advances in neural information processing systems. 4026–4034

  15. [23]

    Roberta Raileanu, Maxwell Goldstein, Denis Yarats, Ilya Kostrikov, and Rob Fer- gus. 2021. Automatic Data Augmentation for Generalization in Reinforcement Learning. In Advances in Neural Information Processing Systems 34: Annual Confer- ence on Neural Information Processing Sy...

  16. [24]

    Tabish Rashid, Mikayel Samvelyan, Christian Schroeder De Witt, Gregory Far- quhar, Jakob Foerster, and Shimon Whiteson. 2018. QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning.arXiv preprint arXiv:1803.11485 (2018)

  17. [25]

    Joshua David Robinson, Ching-Yao Chuang, Suvrit Sra, and Stefanie Jegelka

  18. [26]

    Mikayel Samvelyan, Tabish Rashid, Christian Schroeder de Witt, Gregory Far- quhar, Nantas Nardelli, Tim GJ Rudner, Chia-Man Hung, Philip HS Torr, Jakob Foerster, and Shimon Whiteson. 2019. The starcraft multi-agent challenge. arXiv preprint arXiv:1902.04043 (2019)

  19. [27]

    David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al. 2018. A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play. Science 3...

  20. [28]

    David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al. 2017. Mastering the game of go without human knowledge. nature 550, 7676 (2017), 354–359

  21. [29]

    Amanpreet Singh, Tushar Jain, and Sainbayar Sukhbaatar. 2018. Learning when to communicate at scale in multiagent cooperative and competitive tasks. arXiv preprint arXiv:1812.09755 (2018)

  22. [30]

    Haolin Song, Mingxiao Feng, Wengang Zhou, and Houqiang Li. 2023. Ma2cl: Masked attentive contrastive learning for multi-agent reinforcement learning. arXiv preprint arXiv:2306.02006 (2023)

  23. [31]

    Sainbayar Sukhbaatar, Arthur Szlam, and Rob Fergus. 2016. Learning Multiagent Communication with Backpropagation. In Proceedings of the 30th International Conference on Neural Information Processing Systems (Barcelona, Spain) (NIPS’16). Curran Associates Inc., Red Hook, NY, US...

  24. [32]

    Chuxiong Sun, Bo Wu, Rui Wang, Xiaohui Hu, Xiaoya Yang, and Cong Cong

  25. [33]

    Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vini- cius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al. 2017. Value-decomposition networks for cooperative multi-agent learning. arXiv preprint arXiv:1706.05296 (2017)

  26. [34]

    In Proceedings of the 20th International Conference on Autonomous Agents and MultiAgent Systems (Virtual Event, United Kingdom) (AAMAS ’21)

    Intrinsic Motivated Multi-Agent Communication. In Proceedings of the 20th International Conference on Autonomous Agents and MultiAgent Systems (Virtual Event, United Kingdom) (AAMAS ’21). International Foundation for Autonomous Agents and Multiagent Systems, Richland, SC, 1668–1670

  27. [35]

    Yao-Hung Hubert Tsai, Yue Wu, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2021. Self-supervised Learning from a Multi-view Perspective. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . OpenReview.net. https:/...

  28. [36]

    Yonglong Tian, Dilip Krishnan, and Phillip Isola. 2020. Contrastive Multiview Coding. In Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XI (Lecture Notes in Computer Science, Vol. 12356), Andrea Vedaldi, Horst Bischof...

  29. [37]

    Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, An- drew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al. 2019. Grandmaster level in StarCraft II using multi-agent reinforcement learning. Nature 575, 7782 (2019), 350–354

  30. [38]

    Vikas Verma, Thang Luong, Kenji Kawaguchi, Hieu Pham, and Quoc V. Le. 2021. Towards Domain-Agnostic Contrastive Learning. In Proceedings of the 38th Inter- national Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event (Proceedings of Machine Learning Resea...

  31. [39]

    Tonghan Wang, Jianhao Wang, Chongyi Zheng, and Chongjie Zhang. 2020. Learn- ing Nearly Decomposable Value Functions Via Communication Minimization. In ICLR 2020 : Eighth International Conference on Learning Representations

  32. [40]

    Jianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu, and Chongjie Zhang. 2021. QPLEX: Duplex Dueling Multi-Agent Q-Learning. In ICLR 2021: The Ninth Inter- national Conference on Learning Representations

  33. [41]

    Di Xue, Lei Yuan, Zongzhang Zhang, and Yang Yu. 2022. Efficient Multi-Agent Communication via Shapley Message Value. In Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI-22 , Lud De Raedt (Ed.). International Joint Conferences on ...

  34. [42]

    Efros, and Trevor Darrell

    Tete Xiao, Xiaolong Wang, Alexei A. Efros, and Trevor Darrell. 2021. What Should Not Be Contrastive in Contrastive Learning. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . OpenReview.net. https://openreview.net/f...

  35. [43]

    Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, and Yi Wu. 2022. The surprising effectiveness of ppo in cooperative multi-agent games. Advances in Neural Information Processing Systems35 (2022), 24611–24624

  36. [44]

    Denis Yarats, Ilya Kostrikov, and Rob Fergus. 2021. Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels. In9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net. https://o...

  37. [45]

    Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stéphane Deny. 2021. Barlow Twins: Self-Supervised Learning via Redundancy Reduction. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event (Proceedings of Machine ...

  38. [46]

    Lei Yuan, Jianhao Wang, Fuxiang Zhang, Chenghe Wang, Zongzhang Zhang, Yang Yu, and Chongjie Zhang. 2022. Multi-agent incentive communication via decentralized teammate modeling. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36. 9466–9474

  39. [47]

    Sai Qian Zhang, Qi Zhang, and Jieyu Lin. 2020. Succinct and robust multi-agent communication with temporal message control. Advances in Neural Information Processing Systems 33 (2020), 17271–17282. A IMPLEMENTATION DETAILS A.1 Network architecture Module Architecture Informati...

  40. [48]

    Sai Qian Zhang, Qi Zhang, and Jieyu Lin. 2019. Efficient communication in multi-agent reinforcement learning via variance based control. In Advances in Neural Information Processing Systems . 3235–3244

  41. [2021]

    In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021

    Contrastive Learning with Hard Negative Samples. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net. https://openreview.net/forum?id=CR1XOQ0UTh-

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.