Pith. sign in

REVIEW 13 cited by

Multi-agent Reinforcement Learning in Sequential Social Dilemmas

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1702.03037 v1 pith:YJGTBYAM submitted 2017-02-10 cs.MA cs.AIcs.GTcs.LG

classification cs.MAcs.AIcs.GTcs.LG
keywords dilemmassocialgamepoliciessequentialagentsgamesintroduce
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Matrix games like Prisoner's Dilemma have guided research on social dilemmas for decades. However, they necessarily treat the choice to cooperate or defect as an atomic action. In real-world social dilemmas these choices are temporally extended. Cooperativeness is a property that applies to policies, not elementary actions. We introduce sequential social dilemmas that share the mixed incentive structure of matrix game social dilemmas but also require agents to learn policies that implement their strategic intentions. We analyze the dynamics of policies learned by multiple self-interested independent learning agents, each using its own deep Q-network, on two Markov games we introduce here: 1. a fruit Gathering game and 2. a Wolfpack hunting game. We characterize how learned behavior in each domain changes as a function of environmental factors including resource abundance. Our experiments show how conflict can emerge from competition over shared resources and shed light on how the sequential nature of real world social dilemmas affects cooperation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 278 citations worldwide. Full citation record

  1. Why Study Emergent Behavior When You Can Regulate It? Aligning Multi-Agent Systems with Reward Prediction

    cs.MA 2026-08 conditional novelty 6.0 of 10

    MARP learns rewards from episode-level rankings of social outcomes and shows, in the Harvest Game, that this can steer decentralized agents toward chosen social objectives.

  2. Unsupervised Causal Abstractions Discovery

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Low-rank graphs induce latents that form causal abstractions, with identifiability results and a practical objective enabling unsupervised learning of high-level SCMs from low-level measurements.

  3. LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative Simulacra

    cs.MA 2025-07 reject novelty 6.0 of 10

    The LLM Economist framework couples persona-conditioned worker agents with an in-context RL planner to search US-bracket tax schedules, yet its Saez benchmark is derived from the planner's own solution and its headlin...

  4. Predicting human cooperation: sensitizing drift-diffusion model to interaction and external stimuli

    physics.soc-ph 2024-12 conditional novelty 6.0 of 10

    A Bayesian regressor-driven Drift-Diffusion Model predicts one-step-ahead response-time distributions and final earnings for a multiplayer Prisoner's Dilemma, and simulates how cooperation changes under strategic inte...

  5. Achieving Collective Welfare in Multi-Agent Reinforcement Learning via Suggestion Sharing

    cs.MA 2024-12 conditional novelty 6.0 of 10

    A suggestion-sharing MARL algorithm lets agents exchange optimized action proposals for each other, with a theoretical bound relating the surrogate objective to collective return.

  6. InvestESG: A multi-agent reinforcement learning benchmark for studying climate investment as a social dilemma

    cs.LG 2024-11 conditional novelty 6.0 of 10

    InvestESG is a MARL benchmark showing that ESG-conscious investors, not the disclosure mandate itself, drive corporate mitigation in long-run simulated markets.

  7. OpenSpiel: A Framework for Reinforcement Learning in Games

    cs.LG 2019-08 accept novelty 6.0 of 10

    OpenSpiel provides a unified, open-source API for coding, running, and evaluating many game types and algorithms for reinforcement learning and game theory in one framework.

  8. Memory-Induced Supra-Competitive Outcomes Between Deep Reinforcement Learning Agents in Optimal Trade Execution

    q-fin.CP 2026-05 unverdicted novelty 5.0 of 10

    In a two-agent Almgren-Chriss liquidation game, deep RL agents given intra-episode history of prices and own actions achieve supra-competitive outcomes more frequently and persistently than agents without such memory.

  9. Learning Incentive Structures for Cooperative Resilience in Multi-Agent Systems under Social Dilemmas

    cs.MA 2026-01 unverdicted novelty 5.0 of 10

    A method infers resilience-promoting reward functions via trajectory scoring and integrates them into MARL, with hybrid incentives shown to reduce collapse in disrupted resource environments.

  10. Homing through Reinforcement Learning

    cond-mat.soft 2026-02 reject novelty 4.0 of 10

    In a 2D Q-learning homing model, mean homing time is reported to be non-monotonic in rotational diffusion with a crossover at D_r≈12, and the learned policy is claimed to beat a stochastic-resetting ABP baseline.

  11. Reinforcement Learning for Bidding Strategy Optimization in Day-Ahead Energy Market

    math.OC 2024-11 reject novelty 3.0 of 10

    A DDPG agent learns offering curves for a day-ahead electricity seller from historical Italian PUN prices, but the paper only shows training curves and never demonstrates a validated profit improvement.

  12. AI Researchers Must Help Lead Arms Control to Mitigate Military AI Risks

    cs.CY 2026-06 unverdicted novelty 2.0 of 10

    AI researchers must lead technical research in arms control to mitigate risks from military AI systems, drawing lessons from nuclear deterrence.

  13. Serious Games: Human-AI Interaction, Evolution, and Coevolution

    cs.AI 2025-05 reject novelty 2.0 of 10

    A qualitative review applying Hawk-Dove, Iterated Prisoner's Dilemma, and War of Attrition to human-AI coevolution, with no new data or analysis.

Pith tools