Pith. sign in

REVIEW 2 cited by

Cascading Bandits for Large-Scale Recommendation Problems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1603.05359 v2 pith:Z6GQ7GPG submitted 2016-03-17 cs.LG stat.ML

classification cs.LGstat.ML
keywords itemsitemlearningalgorithmalgorithmsattractionattractivebandits
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Most recommender systems recommend a list of items. The user examines the list, from the first item to the last, and often chooses the first attractive item and does not examine the rest. This type of user behavior can be modeled by the cascade model. In this work, we study cascading bandits, an online learning variant of the cascade model where the goal is to recommend $K$ most attractive items from a large set of $L$ candidate items. We propose two algorithms for solving this problem, which are based on the idea of linear generalization. The key idea in our solutions is that we learn a predictor of the attraction probabilities of items from their features, as opposing to learning the attraction probability of each item independently as in the existing work. This results in practical learning algorithms whose regret does not depend on the number of items $L$. We bound the regret of one algorithm and comprehensively evaluate the other on a range of recommendation problems. The algorithm performs well and outperforms all baselines.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Cascading Bandits Robust to Adversarial Corruptions

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Cascading bandits can be made robust to adversarial click corruption using multi-instance position-based elimination, with regret logarithmic in time and linear in the corruption budget.

  2. Federated Linear Dueling Bandits

    cs.LG 2025-02 reject novelty 6.0 of 10

    A new federated linear dueling bandit algorithm with claimed sublinear regret, but the key proof step is invalid.

Pith tools