Pith. sign in

REVIEW 2 cited by

Bandit Social Learning: Exploration under Myopic Behavior

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.07425 v5 pith:THVR3NSU submitted 2023-02-15 cs.GT cs.DScs.LG

classification cs.GTcs.DScs.LG
keywords banditresultslearningagentsbayesianfailuresalgorithmalgorithms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study social learning dynamics motivated by reviews on online platforms. The agents collectively follow a simple multi-armed bandit protocol, but each agent acts myopically, without regards to exploration. We allow the greedy (exploitation-only) algorithm, as well as a wide range of behavioral biases. Specifically, we allow myopic behaviors that are consistent with (parameterized) confidence intervals for the arms' expected rewards. We derive stark learning failures for any such behavior, and provide matching positive results. The learning-failure results extend to Bayesian agents and Bayesian bandit environments. In particular, we obtain general, quantitatively strong results on failure of the greedy bandit algorithm, both for ``frequentist" and ``Bayesian" versions. Failure results known previously are quantitatively weak, and either trivial or very specialized. Thus, we provide a theoretical foundation for designing non-trivial bandit algorithms, \ie algorithms that intentionally explore, which has been missing from the literature. Our general behavioral model can be interpreted as agents' optimism or pessimism. The matching positive results entail a maximal allowed amount of optimism. Moreover, we find that no amount of pessimism helps against the learning failures, whereas even a small-but-constant fraction of extreme optimists avoids the failures and leads to near-optimal regret rates.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Envy-Free Allocation of Indivisible Goods via Noisy Queries

    cs.GT 2026-02 conditional novelty 7.0 of 10

    With Gaussian noise on valuation queries, two-agent envy-free allocation has query complexity Θ~(m^{5/2}/Δ²) when the optimal envy gap Δ is not too small.

  2. On Incentivized Exploration beyond Bayesianism and Full-Information

    cs.GT 2026-07 conditional novelty 6.0 of 10

    A new Pareto-optimal behavior model shows exactly when sublinear regret is achievable in incentivized exploration with private external information and multiple priors.

Pith tools