A decentralized policy for heterogeneous multiplayer bandits is claimed to achieve O(log^{1+δ}T + W) regret under adversarial zero-reward attacks using one-bit communication.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
stat.ML 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Heterogeneous Multi-Player Multi-Armed Bandits Robust To Adversarial Attacks
A decentralized policy for heterogeneous multiplayer bandits is claimed to achieve O(log^{1+δ}T + W) regret under adversarial zero-reward attacks using one-bit communication.