Pith. sign in

REVIEW 3 cited by

Online Learning for Cooperative Multi-Player Multi-Armed Bandits

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.03818 v1 pith:2A223536 submitted 2021-09-07 cs.LG

classification cs.LG
keywords asymmetryinformationplayerssettingactionsalgorithmregretreward
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

We introduce a framework for decentralized online learning for multi-armed bandits (MAB) with multiple cooperative players. The reward obtained by the players in each round depends on the actions taken by all the players. It's a team setting, and the objective is common. Information asymmetry is what makes the problem interesting and challenging. We consider three types of information asymmetry: action information asymmetry when the actions of the players can't be observed but the rewards received are common; reward information asymmetry when the actions of the other players are observable but rewards received are IID from the same distribution; and when we have both action and reward information asymmetry. For the first setting, we propose a UCB-inspired algorithm that achieves $O(\log T)$ regret whether the rewards are IID or Markovian. For the second section, we offer an environment such that the algorithm given for the first setting gives linear regret. For the third setting, we show that a variation of the `explore then commit' algorithm achieves almost log regret.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Robust Multi-Agent Bandits with Heavy-Tailed Rewards and Information Asymmetry

    cs.LG 2026-08 conditional novelty 6.0 of 10

    Three decentralized multi-agent bandit algorithms achieve near-centralized regret under heavy-tailed rewards across different information asymmetry regimes.

  2. Coordinating the Unknown Lipschitz Constant in Multiplayer Bandits

    cs.LG 2026-08 conditional novelty 6.0 of 10

    Cooperative multiplayer Lipschitz bandits with an unknown Lipschitz constant achieve T^(Md+1)/(Md+2)-scale regret in three information structures, with dithering synchronizing the players' discretizations.

  3. Decentralized Multi-Player Q-Learning in Episodic Markov Decision Processes with Information Asymmetry

    cs.LG 2026-08 reject novelty 4.0 of 10

    Decentralized players using pre-agreed deterministic tie-breaking can match centralized Q-learning regret when either actions or rewards are shared, but the fully asymmetric setting rests on an exploration argument th...

Pith tools