CLTV selects source trajectories that resemble a small target dataset (using learned transition scores and a KL-plus-return trajectory value) and trains the offline RL agent on those trajectories together with the target data.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Enhancing Offline Reinforcement Learning with Curriculum Learning-Based Trajectory Valuation
CLTV selects source trajectories that resemble a small target dataset (using learned transition scores and a KL-plus-return trajectory value) and trains the offline RL agent on those trajectories together with the target data.