Pith. sign in

REVIEW 5 cited by

Emergent Coordination Through Competition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1902.07151 v2 pith:B6P3RAXV submitted 2019-02-19 cs.AI

classification cs.AI
keywords agentsbehaviorbehaviorscontinuousdemonstrateevaluationleadmulti-agent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study the emergence of cooperative behaviors in reinforcement learning agents by introducing a challenging competitive multi-agent soccer environment with continuous simulated physics. We demonstrate that decentralized, population-based training with co-play can lead to a progression in agents' behaviors: from random, to simple ball chasing, and finally showing evidence of cooperation. Our study highlights several of the challenges encountered in large scale multi-agent training in continuous control. In particular, we demonstrate that the automatic optimization of simple shaping rewards, not themselves conducive to co-operative behavior, can lead to long-horizon team behavior. We further apply an evaluation scheme, grounded by game theoretic principals, that can assess agent performance in the absence of pre-defined evaluation tasks or human baselines.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Expected Free Energy-based Planning as Variational Inference

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    EFE-based planning is formulated as variational free energy minimization with epistemic priors, decomposing into expected plan costs plus a complexity term.

  2. What Type of Inference is Active Inference?

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    EFE-based active inference planning is characterized as VFE on an augmented model plus entropy and planning corrections, with a derived message-passing implementation and grid-world validation.

  3. What Type of Inference is Active Inference?

    cs.AI 2026-06 accept novelty 7.0 of 10

    Proper EFE-based planning is VFE plus planning and epistemic entropy corrections, realized by channel-reparameterized message passing that captures novelty.

  4. Multi-Objective Multi-Agent Bandits: From Learning Efficiency to Fairness Optimization

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Pareto UCB1 Gossip achieves O(log T) Pareto regret while Simulated NSW UCB Gossip achieves O(T^{3/4}) regret under fairness, showing the cost of the fairness constraint.

  5. Arena: a toolkit for Multi-Agent Reinforcement Learning

    cs.LG 2019-07 accept novelty 6.0 of 10

    Arena introduces a modular Interface design that extends OpenAI Gym wrappers to support complex multi-agent RL scenarios including self-play and cooperative-competitive interactions.

Pith tools