A small, bounded residual policy, trained online with SAC and introduced gradually, refines frozen Behavior Transformer and Diffusion Policy models to near-perfect success on ManiSkill and Adroit.
24, we experimented with all the aforementioned Q-function architectures in SAC fine-tuning experiments
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.RO 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
A small, bounded residual policy, trained online with SAC and introduced gradually, refines frozen Behavior Transformer and Diffusion Policy models to near-perfect success on ManiSkill and Adroit.