Pith. sign in

REVIEW 1 cited by

Thompson Sampling-like Algorithms for Stochastic Rising Bandits

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.12092 v2 pith:7IOHORUP submitted 2025-05-17 stat.ML cs.LG

classification stat.MLcs.LG
keywords algorithmsregretsettingts-likeapproachesbanditcomplexityexpected
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Stochastic rising rested bandit (SRRB) is a setting where the arms' expected rewards increase as they are pulled. It models scenarios in which the performances of the different options grow as an effect of an underlying learning process (e.g., online model selection). Even if the bandit literature provides specifically crafted algorithms based on upper-confidence bounds for such a setting, no study about Thompson sampling TS-like algorithms has been performed so far. The strong regularity of the expected rewards in the SRRB setting suggests that specific instances may be tackled effectively using adapted and sliding-window TS approaches. This work provides novel regret analyses for such algorithms in SRRBs, highlighting the challenges and providing new technical tools of independent interest. Our results allow us to identify under which assumptions TS-like algorithms succeed in achieving sublinear regret and which properties of the environment govern the complexity of the regret minimization problem when approached with TS. Furthermore, we provide a regret lower bound based on a complexity index we introduce. Finally, we conduct numerical simulations comparing TS-like algorithms with state-of-the-art approaches for SRRBs in synthetic and real-world settings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Periodic Bootstrap Thompson Sampling For Periodically Non-Stationary Bandit Problems

    cs.LG 2026-07 conditional novelty 3.0 of 10

    A Thompson Sampling variant that periodically resets its beliefs and runs forced exploration reduces cumulative regret versus plain Thompson Sampling in simulated periodic bandit problems.

Pith tools