Pith. sign in

REVIEW 2 cited by

Empirically Evaluating Multiagent Learning Algorithms

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1401.8074 v1 pith:M5VRE6N4 submitted 2014-01-31 cs.GT cs.LG

classification cs.GTcs.LG
keywords algorithmslearningmultiagentexperimentsmanyresultsimplementationsmalt
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

There exist many algorithms for learning how to play repeated bimatrix games. Most of these algorithms are justified in terms of some sort of theoretical guarantee. On the other hand, little is known about the empirical performance of these algorithms. Most such claims in the literature are based on small experiments, which has hampered understanding as well as the development of new multiagent learning (MAL) algorithms. We have developed a new suite of tools for running multiagent experiments: the MultiAgent Learning Testbed (MALT). These tools are designed to facilitate larger and more comprehensive experiments by removing the need to build one-off experimental code. MALT also provides baseline implementations of many MAL algorithms, hopefully eliminating or reducing differences between algorithm implementations and increasing the reproducibility of results. Using this test suite, we ran an experiment unprecedented in size. We analyzed the results according to a variety of performance metrics including reward, maxmin distance, regret, and several notions of equilibrium convergence. We confirmed several pieces of conventional wisdom, but also discovered some surprising results. For example, we found that single-agent $Q$-learning outperformed many more complicated and more modern MAL algorithms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. STMARL: A Spatio-Temporal Multi-Agent Reinforcement Learning Approach for Cooperative Traffic Light Control

    cs.MA 2019-08 conditional novelty 5.0 of 10

    STMARL combines graph attention, LSTM memory, and distributed deep Q-learning to coordinate traffic lights, and reports lower average travel times than comparison methods in simulation.

  2. A Review of Cooperative Multi-Agent Deep Reinforcement Learning

    cs.LG 2019-08 conditional novelty 1.0 of 10

    A review that categorizes cooperative multi-agent deep RL into independent learners, observable critics, value factorization, consensus, and communication, with errors in the taxonomy and references.

Pith tools