Pith. sign in

REVIEW 1 cited by

Decentralised Learning in Systems with Many, Many Strategic Agents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1803.05028 v1 pith:EJRJ3QXJ submitted 2018-03-13 cs.MA

classification cs.MA
keywords agentssystemslearningnumberconvergenceinteractingmulti-agentpolicies
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Although multi-agent reinforcement learning can tackle systems of strategically interacting entities, it currently fails in scalability and lacks rigorous convergence guarantees. Crucially, learning in multi-agent systems can become intractable due to the explosion in the size of the state-action space as the number of agents increases. In this paper, we propose a method for computing closed-loop optimal policies in multi-agent systems that scales independently of the number of agents. This allows us to show, for the first time, successful convergence to optimal behaviour in systems with an unbounded number of interacting adaptive learners. Studying the asymptotic regime of N-player stochastic games, we devise a learning protocol that is guaranteed to converge to equilibrium policies even when the number of agents is extremely large. Our method is model-free and completely decentralised so that each agent need only observe its local state information and its realised rewards. We validate these theoretical results by showing convergence to Nash-equilibrium policies in applications from economics and control theory with thousands of strategically interacting agents.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Tacit Learning with Adaptive Information Selection for Cooperative Multi-Agent Reinforcement Learning

    cs.MA 2024-12 conditional novelty 5.0 of 10

    SICA combines selective state-space filtering with attention-based training-time communication and a regeneration module to let MARL agents coordinate without messages at execution time.

Pith tools