Pith. sign in

REVIEW 1 cited by

Influence-Based Multi-Agent Exploration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.05512 v1 pith:KNV2I6SC submitted 2019-10-12 cs.LG cs.MAstat.ML

classification cs.LGcs.MAstat.ML
keywords explorationinfluenceagentsedtieitimulti-agentcoordinatedinteraction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Intrinsically motivated reinforcement learning aims to address the exploration challenge for sparse-reward tasks. However, the study of exploration methods in transition-dependent multi-agent settings is largely absent from the literature. We aim to take a step towards solving this problem. We present two exploration methods: exploration via information-theoretic influence (EITI) and exploration via decision-theoretic influence (EDTI), by exploiting the role of interaction in coordinated behaviors of agents. EITI uses mutual information to capture influence transition dynamics. EDTI uses a novel intrinsic reward, called Value of Interaction (VoI), to characterize and quantify the influence of one agent's behavior on expected returns of other agents. By optimizing EITI or EDTI objective as a regularizer, agents are encouraged to coordinate their exploration and learn policies to optimize team performance. We show how to optimize these regularizers so that they can be easily integrated with policy gradient reinforcement learning. The resulting update rule draws a connection between coordinated exploration and intrinsic reward distribution. Finally, we empirically demonstrate the significant strength of our method in a variety of multi-agent scenarios.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When is Routing Meaningful? Diversity and Robustness in Language Model Societies

    cs.MA 2026-07 conditional novelty 6.0 of 10

    Routing is meaningful only with behavioural diversity (adapted Hierarchic Social Entropy) and assignment stability under query perturbations; accuracy alone can hide vacuous or brittle routing, and fewer than ten mode...

Pith tools