Pith. sign in

REVIEW 18 cited by

Multi-agent Reinforcement Learning in Sequential Social Dilemmas

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1702.03037 v1 pith:YJGTBYAM submitted 2017-02-10 cs.MA cs.AIcs.GTcs.LG

classification cs.MAcs.AIcs.GTcs.LG
keywords dilemmassocialgamepoliciessequentialagentsgamesintroduce
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Matrix games like Prisoner's Dilemma have guided research on social dilemmas for decades. However, they necessarily treat the choice to cooperate or defect as an atomic action. In real-world social dilemmas these choices are temporally extended. Cooperativeness is a property that applies to policies, not elementary actions. We introduce sequential social dilemmas that share the mixed incentive structure of matrix game social dilemmas but also require agents to learn policies that implement their strategic intentions. We analyze the dynamics of policies learned by multiple self-interested independent learning agents, each using its own deep Q-network, on two Markov games we introduce here: 1. a fruit Gathering game and 2. a Wolfpack hunting game. We characterize how learned behavior in each domain changes as a function of environmental factors including resource abundance. Our experiments show how conflict can emerge from competition over shared resources and shed light on how the sequential nature of real world social dilemmas affects cooperation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 278 citations worldwide. Full citation record

  1. Why Study Emergent Behavior When You Can Regulate It? Aligning Multi-Agent Systems with Reward Prediction

    cs.MA 2026-08 conditional novelty 6.0 of 10

    MARP learns rewards from episode-level rankings of social outcomes and shows, in the Harvest Game, that this can steer decentralized agents toward chosen social objectives.

  2. Unsupervised Causal Abstractions Discovery

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Low-rank graphs induce latents that form causal abstractions, with identifiability results and a practical objective enabling unsupervised learning of high-level SCMs from low-level measurements.

  3. Emergence of Fair Leaders via Mediators in Multi-Agent Reinforcement Learning

    cs.MA 2025-08 conditional novelty 6.0 of 10

    A mediator that dynamically selects leaders in Stackelberg MARL can induce self-interested agents to adopt fair policies, improving fairness of returns.

  4. LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative Simulacra

    cs.MA 2025-07 reject novelty 6.0 of 10

    The LLM Economist framework couples persona-conditioned worker agents with an in-context RL planner to search US-bracket tax schedules, yet its Saez benchmark is derived from the planner's own solution and its headlin...

  5. Predicting human cooperation: sensitizing drift-diffusion model to interaction and external stimuli

    physics.soc-ph 2024-12 conditional novelty 6.0 of 10

    A Bayesian regressor-driven Drift-Diffusion Model predicts one-step-ahead response-time distributions and final earnings for a multiplayer Prisoner's Dilemma, and simulates how cooperation changes under strategic inte...

  6. Achieving Collective Welfare in Multi-Agent Reinforcement Learning via Suggestion Sharing

    cs.MA 2024-12 conditional novelty 6.0 of 10

    A suggestion-sharing MARL algorithm lets agents exchange optimized action proposals for each other, with a theoretical bound relating the surrogate objective to collective return.

  7. InvestESG: A multi-agent reinforcement learning benchmark for studying climate investment as a social dilemma

    cs.LG 2024-11 conditional novelty 6.0 of 10

    InvestESG is a MARL benchmark showing that ESG-conscious investors, not the disclosure mandate itself, drive corporate mitigation in long-run simulated markets.

  8. OpenSpiel: A Framework for Reinforcement Learning in Games

    cs.LG 2019-08 accept novelty 6.0 of 10

    OpenSpiel provides a unified, open-source API for coding, running, and evaluating many game types and algorithms for reinforcement learning and game theory in one framework.

  9. Memory-Induced Supra-Competitive Outcomes Between Deep Reinforcement Learning Agents in Optimal Trade Execution

    q-fin.CP 2026-05 unverdicted novelty 5.0 of 10

    In a two-agent Almgren-Chriss liquidation game, deep RL agents given intra-episode history of prices and own actions achieve supra-competitive outcomes more frequently and persistently than agents without such memory.

  10. Learning Incentive Structures for Cooperative Resilience in Multi-Agent Systems under Social Dilemmas

    cs.MA 2026-01 unverdicted novelty 5.0 of 10

    A method infers resilience-promoting reward functions via trajectory scoring and integrates them into MARL, with hybrid incentives shown to reduce collapse in disrupted resource environments.

  11. Fair Contracts in Principal-Agent Games with Heterogeneous Types

    cs.GT 2025-06 conditional novelty 5.0 of 10

    In a two-agent coin game, a principal trained to minimize variance in wealth learns linear contracts that equalize wealth across heterogeneous agents without reducing total welfare.

  12. Decentralized Multi-Player Q-Learning in Episodic Markov Decision Processes with Information Asymmetry

    cs.LG 2026-08 reject novelty 4.0 of 10

    Decentralized players using pre-agreed deterministic tie-breaking can match centralized Q-learning regret when either actions or rewards are shared, but the fully asymmetric setting rests on an exploration argument th...

  13. Homing through Reinforcement Learning

    cond-mat.soft 2026-02 reject novelty 4.0 of 10

    In a 2D Q-learning homing model, mean homing time is reported to be non-monotonic in rotational diffusion with a crossover at D_r≈12, and the learned policy is claimed to beat a stochastic-resetting ABP baseline.

  14. Reciprocity as the Foundational Substrate of Society: How Reciprocal Dynamics Scale into Social Systems

    cs.CY 2025-05 reject novelty 4.0 of 10

    The paper argues that reciprocity, not norms or institutions, is the scalable substrate of society, but it only sketches a framework without formal derivation or simulation.

  15. Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey

    cs.AI 2025-04 conditional novelty 4.0 of 10

    The paper surveys existing work on LLM meta-thinking and argues that multi-agent reinforcement learning is a promising missing ingredient for building self-correcting language models.

  16. Reinforcement Learning for Bidding Strategy Optimization in Day-Ahead Energy Market

    math.OC 2024-11 reject novelty 3.0 of 10

    A DDPG agent learns offering curves for a day-ahead electricity seller from historical Italian PUN prices, but the paper only shows training curves and never demonstrates a validated profit improvement.

  17. AI Researchers Must Help Lead Arms Control to Mitigate Military AI Risks

    cs.CY 2026-06 unverdicted novelty 2.0 of 10

    AI researchers must lead technical research in arms control to mitigate risks from military AI systems, drawing lessons from nuclear deterrence.

  18. Serious Games: Human-AI Interaction, Evolution, and Coevolution

    cs.AI 2025-05 reject novelty 2.0 of 10

    A qualitative review applying Hawk-Dove, Iterated Prisoner's Dilemma, and War of Attrition to human-AI coevolution, with no new data or analysis.

Pith tools