A Thompson sampling bandit with a Bayesian filter over progressively revealed engagement signals improves cold-start recommendation before long-term rewards are observed, with regret bounded by the Value of Progressive Feedback.
context feature vectors
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
method 1
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1roles
method 1polarities
use method 1representative citing papers
citing papers explorer
-
Impatient Bandits: Optimizing for the Long-Term Without Delay
A Thompson sampling bandit with a Bayesian filter over progressively revealed engagement signals improves cold-start recommendation before long-term rewards are observed, with regret bounded by the Value of Progressive Feedback.