LLMDPD conditions an offline policy-diffusion model on LLM-embedded text task descriptions and a transformer-encoded trajectory prompt, reporting improved success on unseen Meta-World and D4RL tasks, though the evaluation protocol limits the strength of the claim.
Ad4rl: Autonomous driving benchmarks for of- fline reinforcement learning with value-based dataset
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
LLM-Driven Policy Diffusion: Enhancing Generalization in Offline Reinforcement Learning
LLMDPD conditions an offline policy-diffusion model on LLM-embedded text task descriptions and a transformer-encoded trajectory prompt, reporting improved success on unseen Meta-World and D4RL tasks, though the evaluation protocol limits the strength of the claim.