REVIEW 32 cited by
Open Problems in Cooperative AI
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Problems of cooperation--in which agents seek ways to jointly improve their welfare--are ubiquitous and important. They can be found at scales ranging from our daily routines--such as driving on highways, scheduling meetings, and working collaboratively--to our global challenges--such as peace, commerce, and pandemic preparedness. Arguably, the success of the human species is rooted in our ability to cooperate. Since machines powered by artificial intelligence are playing an ever greater role in our lives, it will be important to equip them with the capabilities necessary to cooperate and to foster cooperation. We see an opportunity for the field of artificial intelligence to explicitly focus effort on this class of problems, which we term Cooperative AI. The objective of this research would be to study the many aspects of the problems of cooperation and to innovate in AI to contribute to solving these problems. Central goals include building machine agents with the capabilities needed for cooperation, building tools to foster cooperation in populations of (machine and/or human) agents, and otherwise conducting AI research for insight relevant to problems of cooperation. This research integrates ongoing work on multi-agent systems, game theory and social choice, human-machine interaction and alignment, natural-language processing, and the construction of social tools and platforms. However, Cooperative AI is not the union of these existing areas, but rather an independent bet about the productivity of specific kinds of conversations that involve these and other areas. We see opportunity to more explicitly focus on the problem of cooperation, to construct unified theory and vocabulary, and to build bridges with adjacent communities working on cooperation, including in the natural, social, and behavioural sciences.
Forward citations
Cited by 32 Pith papers
-
Multi-Player Discrete-Bidding Games; Determinacy, Equilibria, and Complexity
Under linear tie-breaking, multi-player discrete-bidding games are determined, admit pure Nash equilibria and mean-payoff values, and deciding the winner is already PSPACE-hard for unary reachability.
-
Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety
Changing only the consequence-allocation rule in multi-agent AI shifts collective fatality by 22–58 percentage points across seven model populations, with identity salience in rule text causally driving targeted exploitation.
-
Who Is Really Playing? Strategic Interaction in AI-Guided Populations
A folk theorem for LLMs proves that all feasible and individually rational outcomes can be sustained as ε-equilibria in repeated games where LLMs advise client populations, despite indirect observation.
-
Verbalized Bayesian Persuasion
VBP solves Bayesian persuasion in natural language by treating LLMs as sender and receiver in a mediator-augmented game and searching prompt strategies with Prompt-PSRO.
-
Multi-Agent AI Safety as an Institutional Design Problem
In synthetic delegation workflows, identical final violation rates hide different mechanisms: prompts prevent prohibited attempts, provenance-aware guards block and recover, and a local policy guard fails when transfo...
-
Calibrating Artificial Guilt: Neurally Grounded Reward Shaping for Prosocial Multi-Agent Reinforcement Learning
A happiness-regression contrast from the SoDec dataset is used as a reward-shaping weight in a two-agent Social Lottery, yielding a safe rate of 0.459 versus a human 0.484, but the contrast is statistically indistingu...
-
Strategy, Not Payoffs: A Behavioural Embedding of Normal-Form Games
A two-feature game embedding (Nash entropy and best-response switching) predicts cross-game transfer of fine-tuned LLMs on held-out games, outperforming game identity and published structural embeddings.
-
The Agentic Web Requires New Normative Infrastructure
The web's anti-bot regime should be replaced by a framework that presumptively lets user-authorized AI agents act for their principals, requires platforms to disclose access policies, and permits agent blocking only w...
-
Inferring Hidden Motives: A Utility Bayesian Model of learning the values of others
People update beliefs about others' social preferences as graded, continuous values rather than discrete types, and the best-fitting account uses a seven-parameter utility function estimated from repeated dictator games.
-
Emergence of Fair Leaders via Mediators in Multi-Agent Reinforcement Learning
A mediator that dynamically selects leaders in Stackelberg MARL can induce self-interested agents to adopt fair policies, improving fairness of returns.
-
Network reciprocity turns cheap talk into a force for cooperation
In spatial populations, conditional cooperators that pay a cognitive cost can act as catalysts that make cheap talk evolutionarily effective.
-
How large language models judge and influence human cooperation
LLMs' implicit social norms for judging cooperation vary by model and version, and these differences change predicted long-term cooperation in indirect reciprocity models.
-
Learning from Active Human Involvement through Proxy Value Propagation
A reward-free human-in-the-loop RL method that labels human demonstrations with high Q values and intervened agent actions with low Q values, then propagates these values through TD learning to train policies across d...
-
Governing AI Agents
Agency law and principal-agent theory can frame the governance problems posed by AI agents and justify new principles of inclusivity, visibility, and liability.
-
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations
A frequency-aware mean-field model of incremental Boltzmann Q-learning in the Prisoner's Dilemma predicts that apparent stable cooperation is a long metastable transient and that high discount factors induce oscillati...
-
Emergence of Reputation-Based Cooperation in LLM Agents
AI agents evolve donation strategies resembling Image Scoring, and the steepness of their discrimination against uncooperative opponents predicts resistance to free-riders.
-
Steganalysis of Adaptive Covert Collusion in Tool-Using Agent Populations: A Black-Box, Cross-Principal Approach
Proposes an encoding-agnostic, black-box detector for covert agent collusion and a capacity-theoretic frontier showing low-rate channels are undetectable, but all empirical numbers are placeholders pending measurement.
-
Draining the Energy Commons: Self-Defeating Over-Appropriation as a Coordination Failure in Agentic LLM Collectives
LLM prosumers deplete a shared renewable reserve exactly when demand exceeds peak replacement, acting like impatient open-access users even when sustaining the reserve is feasible.
-
Trust or Check? Understanding the (Evolutionary) Dynamics of User Trust in AI Systems
In an evolutionary game where trust is reduced monitoring, safe and widely adopted AI is the stable outcome only when punishment for unsafe development exceeds the cost of safety and monitoring is affordable.
-
Non-coercive extortion in game theory
An agent can profit by committing to give a co-player an outcome-contingent reward that worsens a target player's equilibrium, and win-win 2x2 games are the most vulnerable.
-
HKGAI-V1: Towards Regional Sovereign Large Language Model for Hong Kong
A DeepSeek-based model fine-tuned for Hong Kong outperforms general models on Hong Kong benchmarks, but most of those benchmarks are self-authored and unreleased.
-
A theory of appropriateness with applications to generative artificial intelligence
A theory that human and AI behavior is guided by context-dependent appropriateness implemented as predictive pattern completion, with norms as conventional sanctioning patterns.
-
The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem
Dominant control-based AI alignment falls short for potential AGI subjects; a parenting model drawing on Turing's child machines should foster gradual autonomy and cooperative coexistence.
-
Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions
Current XAI methods for DNNs and LLMs rest on paradoxes and false assumptions that demand a paradigm shift to verification protocols, scientific foundations, context-aware design, and faithful model analysis rather th...
-
The Theory of Strategic Evolution: Games with Endogenous Players and the Seven Laws of Strategic Replicators
A theory of strategic evolution says multi-level systems of self-reproducing optimizers are stable only under a small-gain condition, and stable AI alignment requires bounding self-modification.
-
Language Games as the Pathway to Artificial Superhuman Intelligence
A position paper arguing that open-ended language games with fluid roles, varied rewards, and evolving rules can drive expanded data reproduction and thus a path to artificial superhuman intelligence.
-
Towards a Theory of AI Personhood
The paper outlines agency, theory of mind, and self-awareness as necessary conditions for AI personhood, reviews inconclusive evidence, and argues that AI personhood would make control-focused alignment ethically problematic.
-
Agentic LLMs in the Supply Chain: Towards Autonomous Multi-Agent Consensus-Seeking
LLM-powered agents that negotiate with neighboring echelons reduce bullwhip and costs in a simulated supply chain, but the results rest on single runs and manually tuned prompts.
-
Towards Transparent Ethical AI: A Roadmap for Trustworthy Robotic Systems
The paper argues transparency is fundamental to trustworthy robotics and proposes a framework connecting technical transparency tools to ethical outcomes such as accountability and informed consent.
-
An Outlook on the Opportunities and Challenges of Multi-Agent AI Systems
The paper formalizes multi-agent AI systems and argues, with toy experiments, that they beat single agents only under narrow conditions on task decomposition, data diversity, and feedback.
-
Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models
A study measures how often GPT-4, Claude, LLaMA, and Gemini agree on PhD-level statistics questions, finding that Claude and GPT-4 produce questions with higher inter-model agreement, but the reliability metric relies...
-
Modeling human reputation-seeking behavior in a spatio-temporally complex public good provision game
A reputation-motivated multi-agent RL model reproduces human groups' cooperation under identifiability and its collapse under anonymity in the Clean Up public goods game.
Discussion (0). Continue with ORCID to comment.