A three-component framework (policy re-evaluation, value alignment, constrained fine-tuning) improves stable fine-tuning from offline RL policies to SAC, TD3, and PPO.
Uncertainty-aware model-based offline reinforcement learning for automated driving
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.LG 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
support 1representative citing papers
citing papers explorer
-
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL
A three-component framework (policy re-evaluation, value alignment, constrained fine-tuning) improves stable fine-tuning from offline RL policies to SAC, TD3, and PPO.