Pith. sign in

REVIEW 2 cited by

Bottom-Up Reputation Promotes Cooperation with Multi-Agent Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.01971 v1 pith:4U7MRUCI submitted 2025-02-04 cs.MA

classification cs.MA
keywords reputationcooperationlearningbottom-upmulti-agentsocialagentsdilemma
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reputation serves as a powerful mechanism for promoting cooperation in multi-agent systems, as agents are more inclined to cooperate with those of good social standing. While existing multi-agent reinforcement learning methods typically rely on predefined social norms to assign reputations, the question of how a population reaches a consensus on judgement when agents hold private, independent views remains unresolved. In this paper, we propose a novel bottom-up reputation learning method, Learning with Reputation Reward (LR2), designed to promote cooperative behaviour through rewards shaping based on assigned reputation. Our agent architecture includes a dilemma policy that determines cooperation by considering the impact on neighbours, and an evaluation policy that assigns reputations to affect the actions of neighbours while optimizing self-objectives. It operates using local observations and interaction-based rewards, without relying on centralized modules or predefined norms. Our findings demonstrate the effectiveness and adaptability of LR2 across various spatial social dilemma scenarios. Interestingly, we find that LR2 stabilizes and enhances cooperation not only with reward reshaping from bottom-up reputation but also by fostering strategy clustering in structured populations, thereby creating environments conducive to sustained cooperation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. The Reciprocity Gradient

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    The reciprocity gradient allows agents to learn near-optimal context-sensitive policies by analytically propagating reward gradients through reputation chains in multi-agent settings.

  2. Reinforcement learning with reputation-based adaptive exploration promotes the evolution of cooperation

    physics.comp-ph 2026-04 unverdicted novelty 5.0 of 10

    Coupling exploration rates to local reputation differences and using asymmetric reputation updates in Q-learning promotes the evolution of cooperation in multi-agent evolutionary games.

Pith tools