Appending autoencoder-derived trajectory encodings to the state space improves offline RL transfer to new CartPole dynamics compared with BCQ, though the effect is small in some environments.
Varibad: A very good method for Bayes-adaptive deep RL via meta-learning
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
TEA: Trajectory Encoding Augmentation for Robust and Transferable Policies in Offline Reinforcement Learning
Appending autoencoder-derived trajectory encodings to the state space improves offline RL transfer to new CartPole dynamics compared with BCQ, though the effect is small in some environments.