Pith. sign in

arXiv preprint arXiv:1910.01465 , year=

5 Pith papers cite this work. Polarity classification is still indexing.

5 Pith papers citing it
abstract

Many real world tasks require multiple agents to work together. Multi-agent reinforcement learning (RL) methods have been proposed in recent years to solve these tasks, but current methods often fail to efficiently learn policies. We thus investigate the presence of a common weakness in single-agent RL, namely value function overestimation bias, in the multi-agent setting. Based on our findings, we propose an approach that reduces this bias by using double centralized critics. We evaluate it on six mixed cooperative-competitive tasks, showing a significant advantage over current methods. Finally, we investigate the application of multi-agent methods to high-dimensional robotic tasks and show that our approach can be used to learn decentralized policies in this domain.

citation-role summary

background 1 method 1

citation-polarity summary

fields

cs.LG 5

years

2026 5

clear filters

representative citing papers

citing papers explorer

Showing 1 of 1 citing paper after filters.