A position paper sketching how RL (DQN, PPO, RLHF) could optimize conversational product recommendation, without any validation.
A vehicle routing problem with dynamic demands and restricted failures solved using stochastic predictive control
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.IR 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Optimizing Conversational Product Recommendation via Reinforcement Learning
A position paper sketching how RL (DQN, PPO, RLHF) could optimize conversational product recommendation, without any validation.