Pith. sign in

REVIEW 1 cited by

Toward Risk-based Optimistic Exploration for Cooperative Multi-Agent Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.01768 v1 pith:NTZ6DD6S submitted 2023-03-03 cs.LG cs.MA

classification cs.LGcs.MA
keywords explorationmulti-agentoptimisticcooperativedistributionallearningreinforcementtake
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

The multi-agent setting is intricate and unpredictable since the behaviors of multiple agents influence one another. To address this environmental uncertainty, distributional reinforcement learning algorithms that incorporate uncertainty via distributional output have been integrated with multi-agent reinforcement learning (MARL) methods, achieving state-of-the-art performance. However, distributional MARL algorithms still rely on the traditional $\epsilon$-greedy, which does not take cooperative strategy into account. In this paper, we present a risk-based exploration that leads to collaboratively optimistic behavior by shifting the sampling region of distribution. Initially, we take expectations from the upper quantiles of state-action values for exploration, which are optimistic actions, and gradually shift the sampling region of quantiles to the full distribution for exploitation. By ensuring that each agent is exposed to the same level of risk, we can force them to take cooperatively optimistic actions. Our method shows remarkable performance in multi-agent settings requiring cooperative exploration based on quantile regression appropriately controlling the level of risk.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Tackling Uncertainties in Multi-Agent Reinforcement Learning through Integration of Agent Termination Dynamics

    cs.LG 2025-01 conditional novelty 4.0 of 10

    Adding a penalty for ally deaths to distributional multi-agent Q-learning improves win rates on StarCraft II and driving benchmarks compared with six baseline algorithms.

Pith tools