SL-MGAC combines supervised reward prediction, user-group decomposition, and actor-critic RL to allocate live streams in a feed; offline and online tests report gains, but the reward predictor is partly fed the true reward bin.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.IR 1years
2024 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Supervised Learning-enhanced Multi-Group Actor Critic for Live Stream Allocation in Feed
SL-MGAC combines supervised reward prediction, user-group decomposition, and actor-critic RL to allocate live streams in a feed; offline and online tests report gains, but the reward predictor is partly fed the true reward bin.