R3S adds an uncertainty penalty from a diffusion world model and a decay-weighted entropy penalty to offline reward shaping, reporting small cumulative-reward gains on Coat, Yahoo, and KuaiRand.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.IR 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Reward Balancing Revisited: Enhancing Offline Reinforcement Learning for Recommender Systems
R3S adds an uncertainty penalty from a diffusion world model and a decay-weighted entropy penalty to offline reward shaping, reporting small cumulative-reward gains on Coat, Yahoo, and KuaiRand.