Pith. sign in

REVIEW 5 cited by

Optimal Cooperative Multiplayer Learning Bandits with Noisy Rewards and No Communication

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.06210 v1 pith:WXFRDHNO submitted 2023-11-10 cs.LG cs.MAstat.ML

classification cs.LGcs.MAstat.ML
keywords playersactionsalgorithmlearningoptimalrewardasymmetrycannot
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

We consider a cooperative multiplayer bandit learning problem where the players are only allowed to agree on a strategy beforehand, but cannot communicate during the learning process. In this problem, each player simultaneously selects an action. Based on the actions selected by all players, the team of players receives a reward. The actions of all the players are commonly observed. However, each player receives a noisy version of the reward which cannot be shared with other players. Since players receive potentially different rewards, there is an asymmetry in the information used to select their actions. In this paper, we provide an algorithm based on upper and lower confidence bounds that the players can use to select their optimal actions despite the asymmetry in the reward information. We show that this algorithm can achieve logarithmic $O(\frac{\log T}{\Delta_{\bm{a}}})$ (gap-dependent) regret as well as $O(\sqrt{T\log T})$ (gap-independent) regret. This is asymptotically optimal in $T$. We also show that it performs empirically better than the current state of the art algorithm for this environment.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Robust Multi-Agent Bandits with Heavy-Tailed Rewards and Information Asymmetry

    cs.LG 2026-08 conditional novelty 6.0 of 10

    Three decentralized multi-agent bandit algorithms achieve near-centralized regret under heavy-tailed rewards across different information asymmetry regimes.

  2. Coordinating the Unknown Lipschitz Constant in Multiplayer Bandits

    cs.LG 2026-08 conditional novelty 6.0 of 10

    Cooperative multiplayer Lipschitz bandits with an unknown Lipschitz constant achieve T^(Md+1)/(Md+2)-scale regret in three information structures, with dithering synchronizing the players' discretizations.

  3. Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework

    cs.MA 2025-06 conditional novelty 6.0 of 10

    A decentralized UCB variant learns unknown coordination thresholds and avoids decoy arms, reportedly approaching oracle-level cumulative reward in a small simulated environment.

  4. DCM Bandits: Multiplayer Information Asymmetric Cascading Bandits for Multiple Clicks

    cs.LG 2026-08 conditional novelty 5.0 of 10

    A new multiplayer formulation of Dependent Click Model bandits with action and reward asymmetry, three decentralized algorithms, and sublinear regret guarantees.

  5. Decentralized Multi-Player Q-Learning in Episodic Markov Decision Processes with Information Asymmetry

    cs.LG 2026-08 reject novelty 4.0 of 10

    Decentralized players using pre-agreed deterministic tie-breaking can match centralized Q-learning regret when either actions or rewards are shared, but the fully asymmetric setting rests on an exploration argument th...

Pith tools