PGCR trains a policy to generate modified user states and an encoder to ignore the changed parts, claiming this finds causally relevant features, and reports improved offline RL recommendation performance.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.IR 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Policy-Guided Causal State Representation for Offline Reinforcement Learning Recommendation
PGCR trains a policy to generate modified user states and an encoder to ignore the changed parts, claiming this finds causally relevant features, and reports improved offline RL recommendation performance.