STARec trains LLM user agents to first rank fast, then reflect on mismatches and rewrite the user profile, using teacher distillation plus GRPO; on MovieLens-1M and Amazon CDs it reportedly beats full-data baselines with a small fraction of the data.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
STARec: An Efficient Agent Framework for Recommender Systems via Autonomous Deliberate Reasoning
STARec trains LLM user agents to first rank fast, then reflect on mismatches and rewrite the user profile, using teacher distillation plus GRPO; on MovieLens-1M and Amazon CDs it reportedly beats full-data baselines with a small fraction of the data.