MARP learns rewards from episode-level rankings of social outcomes and shows, in the Harvest Game, that this can steer decentralized agents toward chosen social objectives.
Journal of Economic Literature , volume=
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.MA 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Why Study Emergent Behavior When You Can Regulate It? Aligning Multi-Agent Systems with Reward Prediction
MARP learns rewards from episode-level rankings of social outcomes and shows, in the Harvest Game, that this can steer decentralized agents toward chosen social objectives.