REVIEW 2 cited by
Model-Based Offline Planning with Trajectory Pruning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The recent offline reinforcement learning (RL) studies have achieved much progress to make RL usable in real-world systems by learning policies from pre-collected datasets without environment interaction. Unfortunately, existing offline RL methods still face many practical challenges in real-world system control tasks, such as computational restriction during agent training and the requirement of extra control flexibility. The model-based planning framework provides an attractive alternative. However, most model-based planning algorithms are not designed for offline settings. Simply combining the ingredients of offline RL with existing methods either provides over-restrictive planning or leads to inferior performance. We propose a new light-weighted model-based offline planning framework, namely MOPP, which tackles the dilemma between the restrictions of offline learning and high-performance planning. MOPP encourages more aggressive trajectory rollout guided by the behavior policy learned from data, and prunes out problematic trajectories to avoid potential out-of-distribution samples. Experimental results show that MOPP provides competitive performance compared with existing model-based offline planning and RL approaches.
Forward citations
Cited by 2 Pith papers
-
Nonreciprocal current induced by dissipation in time-reversal symmetric systems
Dissipation induces nonreciprocal current in time-reversal symmetric noncentrosymmetric systems via interband processes, inversely proportional to lifetime and linked to the shift vector.
-
M$^3$PC: Test-time Model Predictive Control for Pretrained Masked Trajectory Model
M3PC runs model predictive control at test time on a pretrained masked trajectory Transformer, improving offline RL returns and enabling goal reaching without extra model training.
Discussion (0). Continue with ORCID to comment.