Pith. sign in

REVIEW 3 cited by

Hybrid Transfer Reinforcement Learning: Provable Sample Efficiency from Shifted-Dynamics Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.03810 v1 pith:CGGTOGKI submitted 2024-11-06 cs.LG stat.ML

classification cs.LGstat.ML
keywords datasampledynamicsonlinetargetefficiencyenvironmentlearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Online Reinforcement learning (RL) typically requires high-stakes online interaction data to learn a policy for a target task. This prompts interest in leveraging historical data to improve sample efficiency. The historical data may come from outdated or related source environments with different dynamics. It remains unclear how to effectively use such data in the target task to provably enhance learning and sample efficiency. To address this, we propose a hybrid transfer RL (HTRL) setting, where an agent learns in a target environment while accessing offline data from a source environment with shifted dynamics. We show that -- without information on the dynamics shift -- general shifted-dynamics data, even with subtle shifts, does not reduce sample complexity in the target environment. However, with prior information on the degree of the dynamics shift, we design HySRL, a transfer algorithm that achieves problem-dependent sample complexity and outperforms pure online RL. Finally, our experimental results demonstrate that HySRL surpasses state-of-the-art online RL baseline.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Contextual Online Pricing with (Biased) Offline Data

    cs.LG 2025-07 conditional novelty 7.0 of 10

    Contextual online pricing can safely incorporate biased offline data, achieving the standard square-root-of-T worst-case regret and better rates when the offline bias is small.

  2. Transfer Learning Across Policy Regimes in Adaptive Multi-Agent Systems

    cs.MA 2026-06 accept novelty 6.0 of 10

    Restricting a learner to structural knowledge from a prior policy regime improves small-sample performance when the target preserves that structure and produces negative transfer when a threshold break moves the targe...

  3. Pessimism Principle Can Be Effective: Towards a Framework for Zero-Shot Transfer Reinforcement Learning

    cs.LG 2025-05 conditional novelty 5.0 of 10

    A pessimism-based framework for zero-shot transfer RL builds conservative proxies from robust MDPs, yielding lower-bound performance guarantees and distributed algorithms that mitigate negative transfer.

Pith tools