REVIEW 13 cited by
Multi-agent Reinforcement Learning in Sequential Social Dilemmas
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Matrix games like Prisoner's Dilemma have guided research on social dilemmas for decades. However, they necessarily treat the choice to cooperate or defect as an atomic action. In real-world social dilemmas these choices are temporally extended. Cooperativeness is a property that applies to policies, not elementary actions. We introduce sequential social dilemmas that share the mixed incentive structure of matrix game social dilemmas but also require agents to learn policies that implement their strategic intentions. We analyze the dynamics of policies learned by multiple self-interested independent learning agents, each using its own deep Q-network, on two Markov games we introduce here: 1. a fruit Gathering game and 2. a Wolfpack hunting game. We characterize how learned behavior in each domain changes as a function of environmental factors including resource abundance. Our experiments show how conflict can emerge from competition over shared resources and shed light on how the sequential nature of real world social dilemmas affects cooperation.
Forward citations
Cited by 13 Pith papers
-
Why Study Emergent Behavior When You Can Regulate It? Aligning Multi-Agent Systems with Reward Prediction
MARP learns rewards from episode-level rankings of social outcomes and shows, in the Harvest Game, that this can steer decentralized agents toward chosen social objectives.
-
Unsupervised Causal Abstractions Discovery
Low-rank graphs induce latents that form causal abstractions, with identifiability results and a practical objective enabling unsupervised learning of high-level SCMs from low-level measurements.
-
LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative Simulacra
The LLM Economist framework couples persona-conditioned worker agents with an in-context RL planner to search US-bracket tax schedules, yet its Saez benchmark is derived from the planner's own solution and its headlin...
-
Predicting human cooperation: sensitizing drift-diffusion model to interaction and external stimuli
A Bayesian regressor-driven Drift-Diffusion Model predicts one-step-ahead response-time distributions and final earnings for a multiplayer Prisoner's Dilemma, and simulates how cooperation changes under strategic inte...
-
Achieving Collective Welfare in Multi-Agent Reinforcement Learning via Suggestion Sharing
A suggestion-sharing MARL algorithm lets agents exchange optimized action proposals for each other, with a theoretical bound relating the surrogate objective to collective return.
-
InvestESG: A multi-agent reinforcement learning benchmark for studying climate investment as a social dilemma
InvestESG is a MARL benchmark showing that ESG-conscious investors, not the disclosure mandate itself, drive corporate mitigation in long-run simulated markets.
-
OpenSpiel: A Framework for Reinforcement Learning in Games
OpenSpiel provides a unified, open-source API for coding, running, and evaluating many game types and algorithms for reinforcement learning and game theory in one framework.
-
Memory-Induced Supra-Competitive Outcomes Between Deep Reinforcement Learning Agents in Optimal Trade Execution
In a two-agent Almgren-Chriss liquidation game, deep RL agents given intra-episode history of prices and own actions achieve supra-competitive outcomes more frequently and persistently than agents without such memory.
-
Learning Incentive Structures for Cooperative Resilience in Multi-Agent Systems under Social Dilemmas
A method infers resilience-promoting reward functions via trajectory scoring and integrates them into MARL, with hybrid incentives shown to reduce collapse in disrupted resource environments.
-
Homing through Reinforcement Learning
In a 2D Q-learning homing model, mean homing time is reported to be non-monotonic in rotational diffusion with a crossover at D_r≈12, and the learned policy is claimed to beat a stochastic-resetting ABP baseline.
-
Reinforcement Learning for Bidding Strategy Optimization in Day-Ahead Energy Market
A DDPG agent learns offering curves for a day-ahead electricity seller from historical Italian PUN prices, but the paper only shows training curves and never demonstrates a validated profit improvement.
-
AI Researchers Must Help Lead Arms Control to Mitigate Military AI Risks
AI researchers must lead technical research in arms control to mitigate risks from military AI systems, drawing lessons from nuclear deterrence.
-
Serious Games: Human-AI Interaction, Evolution, and Coevolution
A qualitative review applying Hawk-Dove, Iterated Prisoner's Dilemma, and War of Attrition to human-AI coevolution, with no new data or analysis.
Discussion (0). Continue with ORCID to comment.