Pith. sign in

REVIEW 1 cited by

Recommendation System-based Upper Confidence Bound for Online Advertising

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.04190 v1 pith:7GJEM5EG submitted 2019-09-09 cs.IR cs.LGstat.ML

classification cs.IRcs.LGstat.ML
keywords methodrecommendationadvertisingonlineproposedboundconfidencelearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

In this paper, the method UCB-RS, which resorts to recommendation system (RS) for enhancing the upper-confidence bound algorithm UCB, is presented. The proposed method is used for dealing with non-stationary and large-state spaces multi-armed bandit problems. The proposed method has been targeted to the problem of the product recommendation in the online advertising. Through extensive testing with RecoGym, an OpenAI Gym-based reinforcement learning environment for the product recommendation in online advertising, the proposed method outperforms the widespread reinforcement learning schemes such as $\epsilon$-Greedy, Upper Confidence (UCB1) and Exponential Weights for Exploration and Exploitation (EXP3).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Theory to Practice with RAVEN-UCB: Addressing Non-Stationarity in Multi-Armed Bandits through Variance Adaptation

    cs.LG 2025-06 reject novelty 4.0 of 10

    RAVEN-UCB proposes a variance-adaptive UCB algorithm for non-stationary bandits, but the proof of its main regret bound is mathematically invalid.

Pith tools