Pith. sign in

REVIEW 2 cited by

H2O+: An Improved Framework for Hybrid Offline-and-Online RL with Dynamics Gaps

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.12716 v2 pith:53MWVM3D submitted 2023-09-22 cs.LG cs.AIcs.RO

classification cs.LGcs.AIcs.RO
keywords offlinesimulationlearningonlinedatadynamicsenvironmentsflexibility
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Solving real-world complex tasks using reinforcement learning (RL) without high-fidelity simulation environments or large amounts of offline data can be quite challenging. Online RL agents trained in imperfect simulation environments can suffer from severe sim-to-real issues. Offline RL approaches although bypass the need for simulators, often pose demanding requirements on the size and quality of the offline datasets. The recently emerged hybrid offline-and-online RL provides an attractive framework that enables joint use of limited offline data and imperfect simulator for transferable policy learning. In this paper, we develop a new algorithm, called H2O+, which offers great flexibility to bridge various choices of offline and online learning methods, while also accounting for dynamics gaps between the real and simulation environment. Through extensive simulation and real-world robotics experiments, we demonstrate superior performance and flexibility over advanced cross-domain online and offline RL algorithms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hybrid Cross-domain Robust Reinforcement Learning

    cs.LG 2025-05 conditional novelty 6.0 of 10

    HYDRO combines a small offline robust RL dataset with a mismatched online simulator, filtering simulator samples by uncertainty and gap to the worst-case model to improve robust policy performance.

  2. Revealing the Challenges of Sim-to-Real Transfer in Model-Based Reinforcement Learning via Latent Space Modeling

    cs.LG 2025-06 reject novelty 5.0 of 10

    A latent-space extension of MBPO for measuring and mitigating the sim-to-real gap performs inconsistently across MuJoCo perturbations, and its gap metric is non-monotonic.

Pith tools