A Thompson sampling bandit with a Bayesian filter over progressively revealed engagement signals improves cold-start recommendation before long-term rewards are observed, with regret bounded by the Value of Progressive Feedback.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Impatient Bandits: Optimizing for the Long-Term Without Delay
A Thompson sampling bandit with a Bayesian filter over progressively revealed engagement signals improves cold-start recommendation before long-term rewards are observed, with regret bounded by the Value of Progressive Feedback.