Pith. sign in

REVIEW 3 cited by

PettingZoo: Gym for Multi-Agent Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2009.14471 v7 pith:M5PWVVLO submitted 2020-09-30 cs.LG cs.MAstat.ML

classification cs.LGcs.MAstat.ML
keywords pettingzoogamesmarllearninglibrarymodelmulti-agentreinforcement
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces the PettingZoo library and the accompanying Agent Environment Cycle ("AEC") games model. PettingZoo is a library of diverse sets of multi-agent environments with a universal, elegant Python API. PettingZoo was developed with the goal of accelerating research in Multi-Agent Reinforcement Learning ("MARL"), by making work more interchangeable, accessible and reproducible akin to what OpenAI's Gym library did for single-agent reinforcement learning. PettingZoo's API, while inheriting many features of Gym, is unique amongst MARL APIs in that it's based around the novel AEC games model. We argue, in part through case studies on major problems in popular MARL environments, that the popular game models are poor conceptual models of games commonly used in MARL and accordingly can promote confusing bugs that are hard to detect, and that the AEC games model addresses these problems.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-agent imitation learning with function approximation: Linear Markov games and beyond

    cs.LG 2026-02 conditional novelty 7.0 of 10

    In linear Markov games, behavior cloning's sample complexity hinges on a feature-level concentrability coefficient, and the interactive algorithm LSVI-UCB-ZERO-BC removes concentrability dependence entirely, scaling o...

  2. Artificial Generals Intelligence: Mastering Generals.io with Reinforcement Learning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A PPO agent trained with behavior cloning, self-play, and reward shaping reaches a 54.82% win rate against the previous best Generals.io bot and a reported top-25 human leaderboard position.

  3. Multi-Agent Reinforcement Learning in Cybersecurity: From Fundamentals to Applications

    cs.MA 2025-05 conditional novelty 3.0 of 10

    A narrative survey of multi-agent reinforcement learning for cyber defense, reviewing game-theoretic models, cyber gyms, and applications, concluding MARL is promising but faces scalability and simulation-to-real tran...

Pith tools