Pith. sign in

REVIEW 3 major objections 5 minor 51 references

Multiplayer Bandit Learning, from Competition to Cooperation

T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read This paper claims that in a two-player one-armed bandit, competition lowers the exploration threshold below the solo Gittins index, cooperation raises it above, and neutral players can each beat the solo optimum by observing each other's…

desk verdict A useful unifying framework with mostly solid qualitative results, but Theorem 4 Part 2 has two real proof gaps and the almost-sure claims outrun the argument; worth refereeing, not yet citable as stated. read the letter →

arxiv 1908.01135 v4 pith:NT6QK3LV submitted 2019-08-03 cs.GT cs.LGecon.TH

classification cs.GTcs.LGecon.TH
keywords multi-armedbanditsstrategicexperimentationexploration-exploitationtradeoffzero-sumgamescooperativeGittinsindexNashequilibriumperfectBayesian
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper studies a two-player, one-armed bandit game in which each player sees the other's arm choices but not the other's rewards. It asks how the exploration-exploitation tradeoff changes when the players are competing, cooperating, or neutral, and it answers with three threshold statements relative to the Gittins index of the risky arm. Competing players stop exploring above a threshold below the solo optimum; cooperating players explore above the solo optimum; and neutral players, in a middle range of arm quality, each earn strictly more than a single optimal player because they can infer information from each other's actions. The paper also proves that competing and neutral players eventually settle on the same arm in every Nash equilibrium, while cooperating players need not.

What carries the argument

The load-bearing objects are the Gittins index $g=g(\mu,\beta)$, the threshold known-arm probability at which a single player is indifferent between the known and risky arms; the copycat strategy, in which a player stays on the known arm until the opponent explores and then repeats the opponent's previous move one round later; and the induced threshold $p^*\leq (m\beta+g)/(1+\beta)$ where copying destroys the value of exploration. The cooperative result uses a lagged-copy arrangement that gives the team two observations per experiment. The neutral learning result is carried by a perfect Bayesian deviation argument: if a player received only the single-player optimum, she could switch, upon the opponent's first exploration, to the zero-sum copycat strategy and win in the remaining subgame, contradicting equilibrium. The long-term convergence results use a concentration inequality on the empirical mean of the risky arm to show that an infinitely exploring player eventually identifies the better arm and that the other player follows.

What would settle it

Compute, for a fixed prior $\mu$ and discount $\beta$, the value of the zero-sum game that starts with Alice forced to play the risky arm in round 0 and Bob forced to play the known arm. If for some $p\in(p^*,g)$ Bob's value is not strictly positive, the pivot of the neutral-learning theorem fails. A simulation counterpart is to truncate the game at a large horizon, solve for the perfect Bayesian equilibrium by dynamic programming, and check whether each player's expected reward still exceeds the single-player optimum on that interval.

Watch

Extended reading notes

Core claim

The paper's central claim is that the value of information in the two-player game is alignment-dependent. In the zero-sum regime ($\lambda=-1$) there is a threshold $p^*<g$, with the copycat bound $p^*\leq (m\beta+g)/(1+\beta)$, such that for every $p>p^*$ neither player ever explores in equilibrium; a player who experiments first hands the opponent a one-round-lagged copy of the information, and above $p^*$ that erases the explorer's advantage. Below the threshold $m+\beta w/2$, however, both players explore in the first round, so competition does not reduce play to pure myopia. In the fully cooperative regime ($\lambda=1$), one player can explore while the other copies with a delay, effectively turning one experiment into two observations, and this makes the team explore for some $p>g$. In the neutral regime ($\lambda=0$), for $p$ between the competing threshold $p^*$ and the solo Gittins index $g$, every perfect Bayesian equilibrium gives each player strictly more expected reward than a single player using an optimal strategy; the mechanism is that after the opponent's first exploration, a player can switch to a winning zero-sum strategy against her. Finally, in every Nash equilibrium competing and neutral players converge to the same arm with probability 1, whereas cooperating players have equilibria with infinitely many switches.

Load-bearing premise

The neutral-learning theorem rests on the unproved subgame property that after one player explores, the other can guarantee a strictly positive expected advantage for every $p$ in $(p^*,g)$; the paper's copycat argument proves this only for $p$ above the larger cutoff $(m\beta+g)/(1+\beta)$, so the interval between the two cutoffs is where the claim is unsupported.

Editorial extensions

If this is right

  • Competing players will not explore the risky arm for any $p$ above $p^*$, so head-to-head rivalry can freeze experimentation even when a solo learner would continue.
  • Competing players still explore for all sufficiently small $p$, so the zero-sum interaction does not collapse to always playing the safer arm.
  • Two cooperating players can explore for values of $p$ where a single player would stop, because a lagged-copy arrangement makes one experiment yield two observations.
  • Neutral players who observe actions but not rewards can each beat the single-player optimum in every perfect Bayesian equilibrium for $p\in(p^*,g)$, and the probability that no one explores decays exponentially.
  • In every Nash equilibrium, competing and neutral players settle on the same arm almost surely, while cooperating players can have equilibria with one player switching arms infinitely often.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the copycat bound can be sharpened to cover the whole interval $(p^*,g)$, the neutral-learning theorem would pin the end of mutually beneficial learning exactly at the competition threshold; a direct numerical check is to compute the value of the zero-sum subgame after a forced exploration.
  • The same copycat mechanism suggests that any rule raising the cost of imitation, such as a temporary protection for the arm a player first explored, would widen the exploration region for competing and neutral players.
  • The exponential decay of the no-exploration probability gives a quantitative handle for algorithm designers: a finite-time agent can estimate equilibrium exploration probabilities and decide when observation has made further own-experimentation unnecessary.
  • The non-convergence example for cooperating players indicates that aligned payoffs alone do not ensure coordinated specialization; equilibrium selection or communication would be needed to make teams settle on a single arm.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper studies a two-player one-armed bandit in which, each round, a player chooses a predictable left arm (known success probability p) or a risky right arm (prior μ), observes own reward and the other player's action but not the other player's reward, and has utility Γ_i + λΓ_j. The three regimes are competing (λ=-1), neutral (λ=0), and cooperating (λ=1). The main claims are: competing players explore less than a single player for p above a threshold p* ≤ (mβ+g)/(1+β), yet still explore for p below about m+βw/2 (Theorems 1 and 2); cooperating players explore for some p above the Gittins index g (Theorem 3); neutral players explore with probability one for p<g and, in every perfect Bayesian equilibrium for p∈(p*,g), each player strictly beats the single-player optimum (Theorem 4); and competing and neutral players eventually settle on the same arm in every Nash equilibrium while cooperating players may oscillate (Theorems 9-10 and Proposition 2). Finite-horizon analogues and improved bounds for a uniform prior are also given.

Significance. If the proofs are completed, this is a valuable contribution to multiplayer bandit learning and strategic experimentation. The λ-interpolation gives a clean comparative framework, and the paper makes concrete falsifiable predictions: competition reduces exploration, cooperation increases it, neutral players can profit from observing each other's actions, and long-run agreement fails only outside the [-1,0] range. The paper also has genuine technical strengths: the copycat strategy is explicit, the concentration lemma (Lemma 4) is clean, and the thresholds are defined intrinsically rather than fitted to data. The main caveat is that the proof of the headline welfare result for neutral players, Theorem 4 Part 2, currently rests on two unjustified steps, and the algebra in Theorem 1 needs correction. These are fixable in principle, but they are load-bearing.

major comments (3)
  1. [Section 5, proof of Theorem 4 Part 2] The invocation 'By Theorem 1' does not cover all p∈(p*,g). Theorem 1's proof gives Bob a strict advantage via the copycat strategy only when p > (mβ+g)/(1+β), and the theorem only establishes p* ≤ (mβ+g)/(1+β). Thus the interval p∈(p*, (mβ+g)/(1+β)] is left unsupported. A separate argument is needed showing that in the zero-sum subgame where Alice is forced to play R at round k, Bob can guarantee a strictly positive net advantage for every p in (p*,g).
  2. [Section 5, proof of Theorem 4 Part 2] The assertion 'Since the equilibrium is perfect Bayesian, we have E(Γ'_A) ≥ α/(1−β)' does not follow. Alice's strategy S_A is a best response to S_B, not to the deviating strategy S'_B; under S'_B, which plays left until Alice explores, Alice may lose information that she exploited in the original equilibrium, so her payoff could fall below the single-player optimum α/(1−β). Without this lower bound, the inequality E(Γ'_B)>E(Γ'_A) does not imply that Bob's deviation beats α/(1−β), which is the contradiction the proof needs.
  3. [Section 3, Eqs. (7)-(8)] The algebraic identity leading to inequality (8) is incorrect. From the displayed expressions for E(Γ_A) and E(Γ_B) one obtains E(Γ_A)-E(Γ_B) = (m-p(1+β))β^k + (1-β)∑_{t=k+1}^∞ E(γ_A(t))β^t, not the displayed expression with an extra factor β^k on the tail sum. Consequently inequality (8) does not follow as stated. A corrected derivation changes the no-exploration threshold to (m+βg)/(1+β) (under the same bound on the tail), and Theorem 3's use of (8) inherits the problem. The stated theorems may still be true, but the proof must be redone.
minor comments (5)
  1. [Section 1.1 and Section 6.1] The model restricts λ to [-1,1], but Proposition 3 analyzes λ<-1; please clarify whether that proposition is intended as an out-of-model remark or whether the model should allow λ outside [-1,1].
  2. [Section 1.2.1, Theorem 1] The definition of p* as sup{p: arm R is explored in some Nash equilibrium} already makes 'for all p>p* the players do not explore' true by definition; the content of the theorem is the upper bound p*≤(mβ+g)/(1+β). The statement could be rephrased to avoid this redundancy.
  3. [Section 5, proof of Theorem 4 Part 1] The displayed chain from the equilibrium bound to the inequality Φ_k ≥ (α−p)/(1−p−β^k) is compressed; the term handling the case where exploration has already occurred is omitted in the first displayed inequality and only appears implicitly in the next line. Please spell out the derivation.
  4. [Section 6.1, Proposition 3] The sentence 'there is a perfect Bayesian equilibrium in which Bob visits both arms infinitely often whenever' ends abruptly; the trailing 'whenever' should be removed or completed.
  5. [References and typos] Several small typos remain: 'Rotschild' should be 'Rothschild' in Section 7, and 'Salomon' in the bibliography entry [RSV] should be 'Solan'. The reference [RSV] also lacks a year.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the central derivations are self-contained against the external Gittins-index benchmark.

full rationale

The paper's main claims are derived from the model definition, the Gittins index (an external benchmark), and equilibrium reasoning, rather than from fitted parameters or prior results by the authors. The thresholds p* and ~p are defined in terms of equilibrium exploration, but the nontrivial content, such as p* <= (m*beta + g)/(1 + beta) and ~p >= m + beta*w/2, is proved directly via copycat and value arguments; the no-exploration property above p* is a definitional consequence, not a fitted prediction. The neutral-player learning theorem (Theorem 4) is argued from the single-player Gittins optimum and a deviation argument, and the long-run convergence theorems use concentration and martingale-style bounds; none of these reduces to the statements being proved. Two steps in the proof of Theorem 4 Part 2 are not fully justified: the invocation of Theorem 1 for Bob's advantage in the range p in (p*, (m*beta + g)/(1 + beta)], and the assertion E(Gamma'_A) >= alpha/(1 - beta) under (S_A, S'_B) from perfect Bayesian rationality. These are correctness or rigor gaps, not circularity. The paper does not fit parameters to data, rename a known result, or import a uniqueness theorem from the authors' prior work.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The paper postulates no new physical or mathematical entities; it analyzes a standard one-armed bandit with two players and a cooperation parameter. The free parameters of the model (mu, beta, p, lambda) are inputs, not fitted quantities, and the thresholds p*, ~p, p-hat and p-circle are derived rather than fit.

assumptions (5)
  • domain assumption Gittins index theorem for the one-armed bandit, and the reduction of the single-player problem to comparing the risky arm's index g with the safe arm's success probability p.
    The paper imports the Gittins index framework from BJK56 and GJ74 without proof and uses it as the baseline for single-player optimality in Theorems 1, 3 and 4.
  • standard math Sion's minimax theorem guarantees the value of the zero-sum game.
    Invoked in Section 3 to assert the zero-sum game has a value and to support the definition of optimal play.
  • domain assumption Existence of perfect Bayesian equilibria in the neutral game (Fudenberg-Levine).
    Theorem 4 and Theorem 8 rely on PBE existence and on the property that continuation strategies are optimal from every information set.
  • domain assumption The information structure: players observe each other's actions but not rewards.
    This is the defining modeling choice of the paper and the key distinction from perfect-monitoring strategic experimentation papers like BH99 and CKR05.
  • domain assumption The prior mu has no atom at the safe arm's success probability p, i.e. mu(p) = 0.
    Theorems 5, 9 and 10 require this condition to avoid pathological ties in beliefs about whether the risky arm equals the safe arm.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Multiplayer Bandit Learning, from Competition to Cooperation." pith.science (2026). https://pith.science/paper/NT6QK3LV

@misc{pith2026190801135,
  author       = {Pith},
  title        = {Pith review of: Multiplayer Bandit Learning, from Competition to Cooperation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NT6QK3LV}},
  note         = {Machine review of arXiv:1908.01135}
}
abstract

The stochastic multi-armed bandit model captures the tradeoff between exploration and exploitation. We study the effects of competition and cooperation on this tradeoff. Suppose there are $k$ arms and two players, Alice and Bob. In every round, each player pulls an arm, receives the resulting reward, and observes the choice of the other player but not their reward. Alice's utility is $\Gamma_A + \lambda \Gamma_B$ (and similarly for Bob), where $\Gamma_A$ is Alice's total reward and $\lambda \in [-1, 1]$ is a cooperation parameter. At $\lambda = -1$ the players are competing in a zero-sum game, at $\lambda = 1$, they are fully cooperating, and at $\lambda = 0$, they are neutral: each player's utility is their own reward. The model is related to the economics literature on strategic experimentation, where usually players observe each other's rewards. With discount factor $\beta$, the Gittins index reduces the one-player problem to the comparison between a risky arm, with a prior $\mu$, and a predictable arm, with success probability $p$. The value of $p$ where the player is indifferent between the arms is the Gittins index $g = g(\mu,\beta) > m$, where $m$ is the mean of the risky arm. We show that competing players explore less than a single player: there is $p^* \in (m, g)$ so that for all $p > p^*$, the players stay at the predictable arm. However, the players are not myopic: they still explore for some $p > m$. On the other hand, cooperating players explore more than a single player. We also show that neutral players learn from each other, receiving strictly higher total rewards than they would playing alone, for all $ p\in (p^*, g)$, where $p^*$ is the threshold from the competing case. Finally, we show that competing and neutral players eventually settle on the same arm in every Nash equilibrium, while this can fail for cooperating players.

Figures

Figures reproduced from arXiv: 1908.01135 by the authors.

Figure 1
Figure 1. Different regions in which players explore depending on the success probability p of the left arm as a function of the prior µ and the discount factor β. Here m is the mean of µ, while g = g(µ, β) is the Gittins index of the right arm, pe is the threshold where for all p < pe competing players explore, p ∗ the threshold where for all p > p∗ competing players do not explore, pb the threshold where for all p < pb coop… view at source ↗
Figure 2
Figure 2. Trajectories of the players on the main line under strategies [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Depiction of the intervals induced by the sequence [PITH_FULL_IMAGE:figures/full_fig_p019_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Trajectories of the players on the main line under strategies [PITH_FULL_IMAGE:figures/full_fig_p021_4.png]
Figure 5
Figure 5. Figure 5: Bounds as a function of p ∈ [0.5, 1] where the right arm has a uniform prior and β → 1: the red line shows the lower bound on Alice’s net gain given by the function lb(p) = min{1/6 + p 3/6 − p 2/2, 5/6 − 3p/2} (Proposition 4) when in round zero she starts at the right …

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

51 extracted references · 47 canonical work pages

  1. [1]

    Multi-Player Bandits: The Adversarial Case

    P. Alatur, K. Y. Levy, and A. Krause. Multi-player bandits: The adversarial case. arXiv preprint arXiv:1902.08036 , 2019

  2. [2]

    The perils of exploration under competition: A computational modeling approach

    Guy Aridor, Kevin Liu, Aleksandrs Slivkins, and Zhiwei Steven Wu. The perils of exploration under competition: A computational modeling approach. In Proceedings of the 2019 ACM Conference on Economics and Computation, EC 2019, Phoenix, AZ, USA, June 24-28, 2019. , pages 171--172, 2019

  3. [3]

    Aumann and Michael Maschler

    Robert J. Aumann and Michael Maschler. Repeated Games with Incomplete Information . MIT Press, 1995

  4. [4]

    Avner and S

    O. Avner and S. Mannor. Concurrent bandits and cognitive radio networks. In ECML/PKDD , 2014

  5. [5]

    The evolution of eusociality

    M Anderson. The evolution of eusociality. Annual Review of Ecology and Systematics , 15(1):165--189, 1984

  6. [6]

    Mutual observability and the convergence of actions in a multi-person two-armed bandit model

    Masaki Aoyagi. Mutual observability and the convergence of actions in a multi-person two-armed bandit model. Journal of Economic Theory , 82:405--424, 1998

  7. [7]

    Corrigendum: Mutual observability and the convergence of actions in a multi-person two-armed bandit model

    Masaki Aoyagi. Corrigendum: Mutual observability and the convergence of actions in a multi-person two-armed bandit model. 2011

  8. [8]

    Toward a theory of discounted repeated games with imperfect monitoring

    Dilip Abreu, David Pearce, and Ennio Stacchetti. Toward a theory of discounted repeated games with imperfect monitoring. Econometrica , 58(5):1041--1063, 1990

Show all 51 references
  1. [9]

    Robert J. Aumann. Agreeing to disagree. The Annals of Statistics , 4(6):1236--1239, 1976

  2. [10]

    Multi-armed bandit learning in iot networks: Learning helps even in non-stationary settings

    R \' e mi Bonnefoi, Lilian Besson, Christophe Moy, Emilie Kaufmann, and Jacques Palicot. Multi-armed bandit learning in iot networks: Learning helps even in non-stationary settings. In Cognitive Radio Oriented Wireless Networks - 12th International Conference, CROWNCOM 2017, L...

  3. [11]

    Regret analysis of stochastic and nonstochastic multi-armed bandit problems

    Sebastien Bubeck and Nicolo Cesa-Bianchi. Regret analysis of stochastic and nonstochastic multi-armed bandit problems. Foundations and Trends in Machine Learning , 5(1):1--122, 2012

  4. [12]

    Berry and Bert Fristedt

    Donald A. Berry and Bert Fristedt. Bandit problems . Monographs on Statistics and Applied Probability. Chapman & Hall, London, 1985. Sequential allocation of experiments

  5. [13]

    Bolton and C

    P. Bolton and C. Harris. Strategic experimentation. Econometrica , 67(2):349--374, 1999

  6. [14]

    R. N. Bradt, S. M. Johnson, and S. Karlin. On sequential designs for maximizing the sum of n observations. Annals of Mathematical Statistics , 27:1060--1074, 1956

  7. [15]

    Distributed multiplayer bandits - a games of thrones approach

    Ilai Bistritz and Amir Leshem. Distributed multiplayer bandits - a games of thrones approach. In Proceedings of the Conference on Neural Information Processing Systems (NeurIPS) , pages 7222--7232, 2018

  8. [16]

    Non-stochastic multi-player multi-armed bandits: Optimal rate with collision information, sublinear without

    S \' e bastien Bubeck, Yuanzhi Li, Yuval Peres, and Mark Sellke. Non-stochastic multi-player multi-armed bandits: Optimal rate with collision information, sublinear without. CoRR , abs/1904.12233, 2019

  9. [17]

    Matthew Weinberg

    Mark Braverman, Jieming Mao, Jon Schneider, and S. Matthew Weinberg. Multi-armed bandit problems with strategic arms. In Conference on Learning Theory, COLT 2019, 25-28 June 2019, Phoenix, AZ, USA , pages 383--416, 2019

  10. [18]

    Jacobus J. Boomsma. Kin selection versus sexual selection: Why the ends do not meet. Current Biology , 17(16):R673 -- R683, 2007

  11. [19]

    SIC-MMAB: synchronisation involves communication in multiplayer multi-armed bandits

    Etienne Boursier and Vianney Perchet. SIC-MMAB: synchronisation involves communication in multiplayer multi-armed bandits. CoRR , abs/1809.08151, 2018

  12. [20]

    The impact of market structure and learning on the tradeoff between r and d competition and cooperation

    David Besanko and Jianjun Wu. The impact of market structure and learning on the tradeoff between r and d competition and cooperation. Journal of Industrial Economics , 61(1):166--201, 2013

  13. [21]

    Strategic experimentation with exponential bandits

    Martin Cripps, Godfrey Keller, and Sven Rady. Strategic experimentation with exponential bandits. Econometrica , 73(1):39--68, 2005

  14. [22]

    Price of competition and dueling games

    Sina Dehghani, Mohammad Taghi Hajiaghayi, Hamid Mahini, and Saeed Seddighin. Price of competition and dueling games. In 43rd International Colloquium on Automata, Languages, and Programming, ICALP 2016, July 11-15, 2016, Rome, Italy , pages 21:1--21:14, 2016

  15. [23]

    Cooperative and noncooperative research and development in duopoly with spillovers

    Claude D'Aspremont and Alexis Jacquemin. Cooperative and noncooperative research and development in duopoly with spillovers. The American Economic Review , 78(5):1133--1137, 1988

  16. [24]

    Incentivizing exploration

    Peter Frazier, David Kempe, Jon Kleinberg, and Robert Kleinberg. Incentivizing exploration. In Proceedings of the Fifteenth ACM Conference on Economics and Computation , EC '14, pages 5--22, New York, NY, USA, 2014. ACM

  17. [25]

    Subgame-perfect equilibria of finite- and infinite-horizon games

    Drew Fudenberg and David Levine. Subgame-perfect equilibria of finite- and infinite-horizon games. Journal of Economic Theory , 31(2):251 -- 268, 1983

  18. [26]

    Multi-armed Bandit Allocation Indices

    John Gittins, Kevin Glazebrook, and Richard Weber. Multi-armed Bandit Allocation Indices . Wiley, 2011

  19. [27]

    J. C. Gittins. Bandit processes and dynamic allocation indices. Journal of the Royal Statistical Society, Series B , pages 148--177, 1979

  20. [28]

    Gittins and D.M

    J.C. Gittins and D.M. Jones. A dynamic allocation index for the sequential design of experiments. In J. Gani, editor, Progress in Statistics , pages 241--266. North-Holland, Amsterdam, 1974

  21. [29]

    Christoph Gruter, Ellouise Leadbeater, and Francis L. W. Ratnieks. Social learning: The importance of copying others. Current Biology , 20(16), 2010

  22. [30]

    W. D. Hamilton. The genetical evolution of social behaviour. i, ii. Journal of Theoretical Biology , 7(1):1--16, 1964

  23. [31]

    Distributed exploration in multi-armed bandits

    Eshcar Hillel, Zohar S Karnin, Tomer Koren, Ronny Lempel, and Oren Somekh. Distributed exploration in multi-armed bandits. In C. J. C. Burges, L. Bottou, M. Welling, Z. Ghahramani, and K. Q. Weinberger, editors, Advances in Neural Information Processing Systems 26 , pages 854-...

  24. [32]

    Strategic experimentation with private payoffs

    Paul Heidhues, Sven Rady, and Philipp Strack. Strategic experimentation with private payoffs. Journal of Economic Theory , 159:531--551, 2015

  25. [33]

    Dueling algorithms

    Nicole Immorlica, Adam Tauman Kalai, Brendan Lucier, Ankur Moitra, Andrew Postlewaite, and Moshe Tennenholtz. Dueling algorithms. In Proceedings of the 43rd ACM Symposium on Theory of Computing, STOC 2011, San Jose, CA, USA, 6-8 June 2011 , pages 215--224, 2011

  26. [34]

    Decentralized learning for multiplayer multiarmed bandits

    Dileep Kalathil. Decentralized learning for multiplayer multiarmed bandits. IEEE Transactions on Information Theory , 60(4), 2014

  27. [35]

    wisdom of the crowd

    Ilan Kremer, Yishay Mansour, and Motty Perry. Implementing the "wisdom of the crowd". In Proceedings of the Fourteenth ACM Conference on Electronic Commerce , EC '13, pages 605--606, New York, NY, USA, 2013. ACM

  28. [36]

    Karlin and Yuval Peres

    Anna R. Karlin and Yuval Peres. Game theory, alive . American Mathematical Society, Providence, RI, 2017

  29. [37]

    Negatively correlated bandits

    Nicolas Klein and Sven Rady. Negatively correlated bandits. The Review of Economic Studies , 78(2):693--732, 2011

  30. [38]

    Vincent Poor

    Lifeng Lai, Hai Jiang, and H. Vincent Poor. Medium access in cognitive radio networks: A competitive multi-armed bandit framework. 2008 42nd Asilomar Conference on Signals, Systems and Computers , pages 98--102, 2008

  31. [39]

    Multiplayer bandits without observing collision information

    G \' a bor Lugosi and Abbas Mehrabian. Multiplayer bandits without observing collision information. CoRR , abs/1808.08416, 2018

  32. [40]

    Bandit Algorithms

    Tor Lattimore and Csaba Szepesvari. Bandit Algorithms . http://downloads.tor-lattimore.com/banditbook/book.pdf, 2019

  33. [41]

    Distributed learning in multi-armed bandit with multiple players

    Keqin Liu and Qing Zhao. Distributed learning in multi-armed bandit with multiple players. Trans. Sig. Proc. , 58(11):5667--5681, November 2010

  34. [42]

    Bayesian incentive-compatible bandit exploration

    Yishay Mansour, Aleksandrs Slivkins, and Vasilis Syrgkanis. Bayesian incentive-compatible bandit exploration. In Proceedings of the Sixteenth ACM Conference on Economics and Computation , EC '15, pages 565--582, New York, NY, USA, 2015. ACM

  35. [43]

    Competing bandits: Learning under competition

    Yishay Mansour, Aleksandrs Slivkins, and Zhiwei Steven Wu. Competing bandits: Learning under competition. In 9th Innovations in Theoretical Computer Science Conference, ITCS 2018, January 11-14, 2018, Cambridge, MA, USA , pages 48:1--48:27, 2018

  36. [44]

    Game Theory

    Michael Maschler, Eilon Solan, and Shmuel Zamir. Game Theory . Cambridge University Press, 2013

  37. [45]

    Nowak, Corina E

    Martin A. Nowak, Corina E. Tarnita, and Edward O. Wilson. The evolution of eusociality. Nature , 466(7310):1057–1062, 2010

  38. [46]

    A two-armed bandit theory of market pricing

    Michael Rothschild. A two-armed bandit theory of market pricing. Journal of Economic Theory , 9(2):185 -- 202, 1974

  39. [47]

    Multi-player bandits - a musical chairs approach

    Jonathan Rosenski, Ohad Shamir, and Liran Szlak. Multi-player bandits - a musical chairs approach. In Proceedings of the International Conference on Machine Learning (ICML) , pages 155--163, 2016

  40. [48]

    On games of strategic experimentation

    Dinah Rosenberg, Antoine Salomon, and Nicolas Vieille. On games of strategic experimentation. Games and Economic Behavior , 82:31--51

  41. [49]

    Social learning in one-arm bandit problems

    Dinah Rosenberg, Eilon Solan, and Nicolas Vieille. Social learning in one-arm bandit problems. Econometrica , 75(6):1591--1611, 2007

  42. [50]

    On general minimax theorems

    Maurice Sion. On general minimax theorems. Pacific Journal of Mathematics , 8(1):171--176, 1958

  43. [51]

    Introduction to multi-armed bandits

    Aleksandrs Slivkins. Introduction to multi-armed bandits . Foundations and Trends in ML, 2019. draft

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.