LAAC uses an LLM as a reference policy in an adversarial actor-critic setup, with regularization that keeps untested suggestions grounded, improving diversity, novelty, and accuracy on MovieLens-1M.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Large Language Model-Enhanced Reinforcement Learning for Diverse and Novel Recommendations
LAAC uses an LLM as a reference policy in an adversarial actor-critic setup, with regularization that keeps untested suggestions grounded, improving diversity, novelty, and accuracy on MovieLens-1M.