REVIEW 5 major objections 5 minor 49 references
Revisiting Communication Efficiency in Multi-Agent Reinforcement Learning from the Dimensional Analysis Perspective
T0 review · 5 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Dimensional redundancy and confounders persist in receiver-side message embeddings after sender-side gating, and DRMAC removes them to improve multi-agent communication.
desk verdict Plausible receiver-side complement to sender-side MARL communication, with an honest motivation and clean modules, but the evidence needs error bars, multi-map ablations, and a check that decorrelation preserves task-relevant information. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is a redundancy-reduction regularizer applied to the cross-correlation matrix of two views of the same integrated message embedding, together with an Information Selective Network (ISN) that produces a dimensional mask. The regularizer, equation (1), pushes the diagonal of the cross-correlation matrix toward 1 and the off-diagonal entries toward 0, so each dimension of the embedding captures distinct information; the twin views come from the online integration encoder and a noise-perturbed copy of it. The ISN takes the agent's own observation as input and outputs a mask with values in [0,1], applied to the embedding by element-wise multiplication, and is trained through a meta-learning step using second-order gradients of the RL loss so it can adapt to which dimensions matter for the current decision.
What would settle it
On a diagnostic task where the optimal policy is known to depend on a specific, identifiable dimension of the integrated message embedding, train DRMAC while varying the noise scale used to create the twin view or removing the redundancy term; if performance collapses when that dimension is decorrelated or down-weighted, the assumption that redundancy reduction preserves decision-relevant information is contradicted.
Extended reading notes
Core claim
The central discovery is that, even after optimized and gated messages are sent, dimensional redundancy and confounders persist in the integrated message embeddings at the receiving end, and these negatively affect communication quality and decision-making. DRMAC mitigates both: a redundancy-reduction regularization term drives the cross-correlation matrix of twin message embeddings toward the identity matrix, decoupling the information carried by each dimension, and an Information Selective Network learns a dimensional mask that dynamically adjusts gradient weights so the policy focuses on decision-relevant dimensions while suppressing confounders. The paper reports that DRMAC consistently outperforms state-of-the-art communication baselines across a diverse set of cooperative multi-agent tasks, and that its key modules serve as a plug-and-play complement to existing communication strategies rather than a replacement.
Load-bearing premise
The load-bearing premise is that forcing the cross-correlation of two noise-corrupted views of the same integrated message toward the identity removes redundant dimensions without also discarding information the policy needs; if the Gaussian noise augmentation or the decorrelation destroys task-critical features, DRMAC's performance gains would disappear.
Editorial extensions
If this is right
- Sender-side message filtering of content and timing alone does not resolve communication efficiency; receiver-side dimensional analysis is a necessary complementary step.
- Decorrelation of message-embedding dimensions plus masking of decision-irrelevant dimensions improves cooperative performance across tasks with different numbers of agents and difficulty levels.
- DRMAC's modules are plug-and-play: adding them to different MARL training algorithms (VDN, QMIX, MAPPO) and to communication methods (MASIA, SMS, TarMAC) improves those methods' performance.
- Ablation shows that the redundancy-reduction term and the ISN mask are mutually reinforcing: removing either one degrades performance.
Reading between the lines
- The motivating evidence that randomly masking dimensions sometimes improves performance suggests that a substantial fraction of embedding dimensions are actively harmful rather than merely redundant; a natural follow-up is to measure how that fraction changes with task complexity and message length.
- The redundancy-reduction objective is adapted from self-supervised learning, so a targeted ablation of the Gaussian noise injection and the projection layer could reveal whether the invariance term, rather than the decorrelation, is the main driver of the gains.
- The ISN is trained only by RL performance while the encoder is trained with the redundancy loss; one extension would be to let the mask influence which dimensions are decorrelated, instead of assuming all off-diagonal correlations are equally harmful.
- Because the evaluation is limited to cooperative tasks with shared rewards, it remains open whether dimensional confounders play the same role in competitive or mixed-motive settings, where a receiver may want to distrust certain message dimensions.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that in cooperative MARL, sender-side optimization of communication content and timing is insufficient because the receiver-side integrated message embeddings still exhibit dimensional redundancy and dimensional confounders. To address this, the authors propose DRMAC, which adds a Barlow-Twins-style redundancy-reduction regularization term on two views of integrated messages and an Information Selective Network that learns a per-agent dimensional mask via meta-learning. The method is evaluated on Hallway and SMAC tasks against TarMAC, MAIC, SMS, MASIA, and QMIX, and is also applied as a plug-and-play module on top of several baselines. The central claim is that DRMAC consistently outperforms existing methods by removing receiver-side dimensional redundancy and confounders while preserving decision-relevant information.
Significance. If the central claim is validated, DRMAC offers a useful complementary perspective to sender-side communication-efficiency methods and provides a simple, plug-and-play module that can be integrated with existing MARL algorithms. The motivating experiments in Figures 2 and 3 are a strength: they give a concrete demonstration that residual redundancy and harmful dimensions persist after sender-side optimization. The plug-and-play experiments in Figure 8 are also valuable because they show compatibility with multiple learning algorithms and communication baselines. However, the current empirical support is not yet sufficient: the headline comparisons are learning curves without error bars or significance tests, the component ablation is performed on a single SMAC map, and there is no direct evidence that decorrelation preserves task-relevant information. The contribution is potentially interesting, but the load-bearing claims require additional support before the paper can be accepted.
major comments (5)
- [Section 5.2, Appendix B] The headline claim that DRMAC 'consistently outperforms the baselines in almost all environments' is supported only by learning curves. Appendix B states that results are averaged over 5 random seeds, but no variance, error bars, numerical table, or significance test is reported anywhere in Section 5.2 or Figure 6. As a result, the reader cannot assess whether the observed advantages are statistically reliable or within run-to-run noise. Please add a quantitative summary (e.g., final mean and standard deviation per task) and, where practical, statistical comparisons for at least the main environments.
- [Section 4.1, Eq. (1)] The redundancy-reduction objective follows Barlow Twins, whose rationale assumes the two views share semantic content. Here the two views are produced by feeding the identical message through the online IIE and through a copy with Gaussian noise added to its weights; the invariance term therefore enforces invariance to weight perturbation, not to a content-preserving input augmentation. An encoder can satisfy Eq. (1) by becoming insensitive to its own parameter noise, which does not guarantee that decision-relevant information is retained. The paper provides no probe or information-theoretic check that task-relevant content survives in the decorrelated embedding, so the observed performance gains cannot be attributed specifically to redundancy removal as opposed to an incidental regularizing effect. Please add a measuring experiment (e.g., predicting the global state or the optimal action from the masked representation) or an ablation with a semantic augmentation.
- [Section 5.3, Fig. 7] The ablation study evaluates the two components on a single SMAC map, 1o_2r_vs_4r. The conclusion that redundancy reduction and ISN are mutually reinforcing is therefore based on one scenario; please extend the ablations to at least one additional task (for instance a Hallway variant or another SMAC map) to support the claimed generality.
- [Section 4.3, Eqs. (5)-(6)] The description of the ISN meta-learning update is ambiguous: the text says 'back-propagation of gradients is excluded' when computing the trial weights, yet immediately concludes that θ_ISN 'is refined through second-order gradient optimization.' If the trial weights are detached from the computation graph, the resulting meta-gradient is first-order; if they are not detached, the statement about excluding back-propagation is misleading. Please clarify exactly which computation graph is used and which order of derivatives is taken, since this determines both the optimization behavior and the reproducibility of the method.
- [Section 4.2, Section 5.3] The paper does not provide direct evidence that the learned ISN mask corresponds to decision-irrelevant dimensions. No analysis of the learned mask values, no comparison with masking the same number of random dimensions, and no visualization of which dimensions are suppressed is reported. Without such evidence, the improvement of DRMAC over DRMAC w/o ISN might reflect a generic input-dependent feature weighting rather than the removal of dimensional confounders, which is one of the paper's two central mechanisms.
minor comments (5)
- [Section 3] The Dec-POMDP tuple includes a message set M, but the formal formulation does not describe how messages are generated or how they enter the observation/reward structure; please make this explicit.
- [Section 4.3] The parameter set θ includes θ_IEE, while the encoder is abbreviated IIE in Section 4.1; please unify the notation.
- [Section 5.1, Appendix B] Section 5.1 states 'two well-known cooperative multi-agent environments, encompassing a total of eight tasks,' while Appendix B begins with 'The four test environments,' which is inconsistent.
- [Section 4.1, Eq. (2)] The normalization in Eq. (2) lacks an epsilon term for numerical stability; please add one or explain how zero-variance dimensions are handled.
- [Appendix A] The paper says to refer to Appendix A for implementation details, but Appendix A only reports DRMAC's architecture and hyperparameters, not a complete training recipe or the baseline configurations; please add a reproducibility statement or release code.
Circularity Check
No significant circularity: the decorrelation and masking modules are trained objectives, and the motivating baseline evidence is independent of DRMAC's construction.
full rationale
The paper's central claim is that receiver-side message embeddings retain dimensional redundancy and confounders even after sender-side filtering, and that DRMAC's redundancy-reduction loss and information-selective mask improve communication. This claim does not reduce to its own inputs by the paper's equations. The motivating evidence in Figs. 2 and 3 is obtained from the baseline methods MASIA and SMS before DRMAC is introduced: representations are visualized for dimensional similarity, and random dimensional masking is shown to sometimes improve baseline performance. Those observations are independent of the proposed objective. The method is a trained objective: Eq. (1) directly drives the cross-correlation matrix toward the identity, so the later demonstration that DRMAC's representation is less redundant is a verification of the training target, not a prediction made from a fitted parameter. Similarly, the ISN dimensional mask is learned via meta-learning on the RL loss in Eqs. (4)-(6), so its contribution to performance is the normal behavior of a trained module, and the ablation in Sec. 5.3 compares trained variants with a baseline rather than renaming a fitted value as a prediction. The only self-citation, reference [32] (IMMAC), appears in the related-work survey of when-to-communicate methods; it is not load-bearing for DRMAC's design or results, and no uniqueness theorem or external forcing is imported from the authors' prior work. The transfer of Barlow Twins [45] and the choice of Gaussian weight-noise augmentation are methodological assumptions, but they are not circular because the paper does not define the target result in terms of those assumptions. Experimental comparisons use official implementations of external baselines (TarMAC, MAIC, SMS, MASIA, QMIX), so the empirical support is not internal to a fitted quantity. No step satisfies the requirement of exhibiting a specific reduction of a prediction to an input by construction.
Assumptions & free parameters
free parameters (4)
- beta =
0.01
- lambda =
0.0051
- Gaussian noise scale for target encoder
- Meta-learning learning rate for trial weights
assumptions (5)
- standard math Tasks are Dec-POMDPs with shared rewards and messages exchanged among agents (Section 3).
- domain assumption Factorial coding, i.e., decorrelating representation dimensions, improves downstream task performance (Section 4.1, Eq. 1).
- domain assumption Adding Gaussian noise to a copy of the Information Integration Encoder produces valid augmented views of the same message (Section 4.1, Fig. 4).
- ad hoc to paper A one-step-unrolled, second-order meta-gradient update of the ISN mask reliably tracks which dimensions are decision-relevant in a non-stationary MARL setting (Section 4.3, Eqs. 5-6).
- ad hoc to paper Per-dimension multiplicative masks bounded in [0,1] can remove confounding information without hurting the policy (Section 4.2, Eq. 3).
Cite this review
Pith. "Pith review of Revisiting Communication Efficiency in Multi-Agent Reinforcement Learning from the Dimensional Analysis Perspective." pith.science (2026). https://pith.science/paper/YP2MSC3Y
@misc{pith2026250102888,
author = {Pith},
title = {Pith review of: Revisiting Communication Efficiency in Multi-Agent Reinforcement Learning from the Dimensional Analysis Perspective},
year = {2026},
howpublished = {\url{https://pith.science/paper/YP2MSC3Y}},
note = {Machine review of arXiv:2501.02888}
}
read the original abstract
In this work, we introduce a novel perspective, i.e., dimensional analysis, to address the challenge of communication efficiency in Multi-Agent Reinforcement Learning (MARL). Our findings reveal that simply optimizing the content and timing of communication at sending end is insufficient to fully resolve communication efficiency issues. Even after applying optimized and gated messages, dimensional redundancy and confounders still persist in the integrated message embeddings at receiving end, which negatively impact communication quality and decision-making. To address these challenges, we propose Dimensional Rational Multi-Agent Communication (DRMAC), designed to mitigate both dimensional redundancy and confounders in MARL. DRMAC incorporates a redundancy-reduction regularization term to encourage the decoupling of information across dimensions within the learned representations of integrated messages. Additionally, we introduce a dimensional mask that dynamically adjusts gradient weights during training to eliminate the influence of decision-irrelevant dimensions. We evaluate DRMAC across a diverse set of multi-agent tasks, demonstrating its superior performance over existing state-of-the-art methods in complex scenarios. Furthermore, the plug-and-play nature of DRMAC's key modules highlights its generalizable performance, serving as a valuable complement rather than a replacement for existing multi-agent communication strategies.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
OpenAI: Marcin Andrychowicz, Bowen Baker, Maciek Chociej, Rafal Jozefowicz, Bob McGrew, Jakub Pachocki, Arthur Petron, Matthias Plappert, Glenn Powell, Alex Ray, et al. 2020. Learning dexterous in-hand manipulation. The International Journal of Robotics Research 39, 1 (2020), 3–20
2020
-
[2]
Devon Hjelm, and William Buchwalter
Philip Bachman, R. Devon Hjelm, and William Buchwalter. 2019. Learning Repre- sentations by Maximizing Mutual Information Across Views. InAdvances in Neu- ral Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, Hanna M. Wallach, Hugo Larochelle, Alina...
work page 2019
-
[3]
Horace B Barlow et al. 1961. Possible principles underlying the transformation of sensory messages. Sensory communication 1, 01 (1961), 217–233
work page 1961
-
[4]
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. 2020. Unsupervised Learning of Visual Features by Contrasting Cluster Assignments. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, Hugo Larochell...
work page 2020
-
[5]
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey E. Hinton. 2020. A Simple Framework for Contrastive Learning of Visual Representations. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event (Proceedings of Machine Learning Research, Vol. 119) . PMLR, 1597–1607. http://proceedings.mlr....
2020
-
[6]
Ching-Yao Chuang, Joshua Robinson, Yen-Chen Lin, Antonio Torralba, and Stefanie Jegelka. 2020. Debiased Contrastive Learning. In Advances in Neu- ral Information Processing Systems 33: Annual Conference on Neural Infor- mation Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual , Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-F...
work page 2020
-
[7]
Abhishek Das, Théophile Gervet, Joshua Romoff, Dhruv Batra, Devi Parikh, Mike Rabbat, and Joelle Pineau. 2019. Tarmac: Targeted multi-agent communication. In International Conference on Machine Learning . 1538–1546
work page 2019
-
[8]
Ziluo Ding, Tiejun Huang, and Zongqing Lu. 2020. Learning individually inferred communication for multi-agent cooperation. Advances in Neural Information Processing Systems 33 (2020), 22069–22079
2020
Show all 49 references
-
[9]
Aleksandr Ermolov, Aliaksandr Siarohin, Enver Sangineto, and Nicu Sebe. 2021. Whitening for Self-Supervised Representation Learning. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event (Proceedings of Machine Learn...
2021
-
[10]
Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Ávila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Rémi Munos, and Michal Valko
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H. Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Ávila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Rémi Munos, and Michal Valko. 2020. Bootstrap Your Own Latent -...
2020
-
[11]
Cong Guan, Feng Chen, Lei Yuan, Chenghe Wang, Hao Yin, Zongzhang Zhang, and Yang Yu. 2022. Efficient Multi-agent Communication via Self-supervised Information Aggregation. Advances in Neural Information Processing Systems 35 (2022), 1020–1033
2022
-
[12]
Girshick
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross B. Girshick. 2020. Momentum Contrast for Unsupervised Visual Representation Learning. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, W A, USA, June 13-19, 2020. Computer Vision ...
2020
-
[13]
Jiechuan Jiang and Zongqing Lu. 2018. Learning attentional communication for multi-agent cooperation. In Advances in neural information processing systems . 7254–7264
2018
-
[14]
Daewoo Kim, Sangwoo Moon, David Hostallero, Wan Ju Kang, Taeyoung Lee, Kyunghwan Son, and Yung Yi. 2019. Learning to schedule communication in multi-agent reinforcement learning. arXiv preprint arXiv:1902.01554 (2019)
2019 arXiv
-
[15]
Michael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto, Pieter Abbeel, and Ar- avind Srinivas. 2020. Reinforcement Learning with Augmented Data. InAdvances in Neural Information Processing Systems 33: Annual Conference on Neural In- formation Processing Systems 2020, NeurIPS 202...
2020
-
[16]
Michael Laskin, Aravind Srinivas, and Pieter Abbeel. 2020. CURL: Contrastive Unsupervised Representations for Reinforcement Learning. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event (Proceedings of Machine Lear...
2020
-
[17]
Edouard Leurent. 2018. A survey of state-action representations for autonomous driving. (2018)
2018
-
[18]
Shikun Liu, Andrew Davison, and Edward Johns. 2019. Self-supervised generali- sation with meta auxiliary learning. Advances in Neural Information Processing Systems 32 (2019)
2019
-
[19]
Ryan Lowe, Yi I Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch. 2017. Multi-agent actor-critic for mixed cooperative-competitive environments. In Advances in neural information processing systems . 6379–6390
2017
-
[20]
Yaru Niu, Rohan R Paleja, and Matthew C Gombolay. 2021. Multi-Agent Graph- Attention Communication and Teaming.. In AAMAS. 964–973
2021
-
[21]
Frans A Oliehoek, Christopher Amato, et al . 2016. A concise introduction to decentralized POMDPs. Vol. 1. Springer
2016
-
[22]
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy. 2016. Deep exploration via bootstrapped DQN. In Advances in neural information processing systems. 4026–4034
2016
-
[23]
Roberta Raileanu, Maxwell Goldstein, Denis Yarats, Ilya Kostrikov, and Rob Fer- gus. 2021. Automatic Data Augmentation for Generalization in Reinforcement Learning. In Advances in Neural Information Processing Systems 34: Annual Confer- ence on Neural Information Processing Sy...
2021
-
[24]
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder De Witt, Gregory Far- quhar, Jakob Foerster, and Shimon Whiteson. 2018. QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning.arXiv preprint arXiv:1803.11485 (2018)
2018 arXiv
-
[25]
Joshua David Robinson, Ching-Yao Chuang, Suvrit Sra, and Stefanie Jegelka
-
[26]
Mikayel Samvelyan, Tabish Rashid, Christian Schroeder de Witt, Gregory Far- quhar, Nantas Nardelli, Tim GJ Rudner, Chia-Man Hung, Philip HS Torr, Jakob Foerster, and Shimon Whiteson. 2019. The starcraft multi-agent challenge. arXiv preprint arXiv:1902.04043 (2019)
2019 arXiv
-
[27]
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al. 2018. A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play. Science 3...
2018
-
[28]
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al. 2017. Mastering the game of go without human knowledge. nature 550, 7676 (2017), 354–359
2017
-
[29]
Amanpreet Singh, Tushar Jain, and Sainbayar Sukhbaatar. 2018. Learning when to communicate at scale in multiagent cooperative and competitive tasks. arXiv preprint arXiv:1812.09755 (2018)
2018 arXiv
-
[30]
Haolin Song, Mingxiao Feng, Wengang Zhou, and Houqiang Li. 2023. Ma2cl: Masked attentive contrastive learning for multi-agent reinforcement learning. arXiv preprint arXiv:2306.02006 (2023)
2023 arXiv
-
[31]
Sainbayar Sukhbaatar, Arthur Szlam, and Rob Fergus. 2016. Learning Multiagent Communication with Backpropagation. In Proceedings of the 30th International Conference on Neural Information Processing Systems (Barcelona, Spain) (NIPS’16). Curran Associates Inc., Red Hook, NY, US...
2016
-
[32]
Chuxiong Sun, Bo Wu, Rui Wang, Xiaohui Hu, Xiaoya Yang, and Cong Cong
-
[33]
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vini- cius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al. 2017. Value-decomposition networks for cooperative multi-agent learning. arXiv preprint arXiv:1706.05296 (2017)
2017 arXiv
-
[34]
In Proceedings of the 20th International Conference on Autonomous Agents and MultiAgent Systems (Virtual Event, United Kingdom) (AAMAS ’21)
Intrinsic Motivated Multi-Agent Communication. In Proceedings of the 20th International Conference on Autonomous Agents and MultiAgent Systems (Virtual Event, United Kingdom) (AAMAS ’21). International Foundation for Autonomous Agents and Multiagent Systems, Richland, SC, 1668–1670
-
[35]
Yao-Hung Hubert Tsai, Yue Wu, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2021. Self-supervised Learning from a Multi-view Perspective. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . OpenReview.net. https:/...
2021
-
[36]
Yonglong Tian, Dilip Krishnan, and Phillip Isola. 2020. Contrastive Multiview Coding. In Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XI (Lecture Notes in Computer Science, Vol. 12356), Andrea Vedaldi, Horst Bischof...
2020 doi
-
[37]
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, An- drew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al. 2019. Grandmaster level in StarCraft II using multi-agent reinforcement learning. Nature 575, 7782 (2019), 350–354
2019
-
[38]
Vikas Verma, Thang Luong, Kenji Kawaguchi, Hieu Pham, and Quoc V. Le. 2021. Towards Domain-Agnostic Contrastive Learning. In Proceedings of the 38th Inter- national Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event (Proceedings of Machine Learning Resea...
2021
-
[39]
Tonghan Wang, Jianhao Wang, Chongyi Zheng, and Chongjie Zhang. 2020. Learn- ing Nearly Decomposable Value Functions Via Communication Minimization. In ICLR 2020 : Eighth International Conference on Learning Representations
2020
-
[40]
Jianhao Wang, Zhizhou Ren, Terry Liu, Yang Yu, and Chongjie Zhang. 2021. QPLEX: Duplex Dueling Multi-Agent Q-Learning. In ICLR 2021: The Ninth Inter- national Conference on Learning Representations
2021
-
[41]
Di Xue, Lei Yuan, Zongzhang Zhang, and Yang Yu. 2022. Efficient Multi-Agent Communication via Shapley Message Value. In Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI-22 , Lud De Raedt (Ed.). International Joint Conferences on ...
2022 doi
-
[42]
Efros, and Trevor Darrell
Tete Xiao, Xiaolong Wang, Alexei A. Efros, and Trevor Darrell. 2021. What Should Not Be Contrastive in Contrastive Learning. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . OpenReview.net. https://openreview.net/f...
2021
-
[43]
Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, and Yi Wu. 2022. The surprising effectiveness of ppo in cooperative multi-agent games. Advances in Neural Information Processing Systems35 (2022), 24611–24624
2022
-
[44]
Denis Yarats, Ilya Kostrikov, and Rob Fergus. 2021. Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels. In9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net. https://o...
2021
-
[45]
Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stéphane Deny. 2021. Barlow Twins: Self-Supervised Learning via Redundancy Reduction. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event (Proceedings of Machine ...
2021
-
[46]
Lei Yuan, Jianhao Wang, Fuxiang Zhang, Chenghe Wang, Zongzhang Zhang, Yang Yu, and Chongjie Zhang. 2022. Multi-agent incentive communication via decentralized teammate modeling. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36. 9466–9474
2022
-
[47]
Sai Qian Zhang, Qi Zhang, and Jieyu Lin. 2020. Succinct and robust multi-agent communication with temporal message control. Advances in Neural Information Processing Systems 33 (2020), 17271–17282. A IMPLEMENTATION DETAILS A.1 Network architecture Module Architecture Informati...
2020
-
[48]
Sai Qian Zhang, Qi Zhang, and Jieyu Lin. 2019. Efficient communication in multi-agent reinforcement learning via variance based control. In Advances in Neural Information Processing Systems . 3235–3244
2019
-
[2021]
In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021
Contrastive Learning with Hard Negative Samples. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net. https://openreview.net/forum?id=CR1XOQ0UTh-
2021
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.