REVIEW 3 cited by
Hybrid Transfer Reinforcement Learning: Provable Sample Efficiency from Shifted-Dynamics Data
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Online Reinforcement learning (RL) typically requires high-stakes online interaction data to learn a policy for a target task. This prompts interest in leveraging historical data to improve sample efficiency. The historical data may come from outdated or related source environments with different dynamics. It remains unclear how to effectively use such data in the target task to provably enhance learning and sample efficiency. To address this, we propose a hybrid transfer RL (HTRL) setting, where an agent learns in a target environment while accessing offline data from a source environment with shifted dynamics. We show that -- without information on the dynamics shift -- general shifted-dynamics data, even with subtle shifts, does not reduce sample complexity in the target environment. However, with prior information on the degree of the dynamics shift, we design HySRL, a transfer algorithm that achieves problem-dependent sample complexity and outperforms pure online RL. Finally, our experimental results demonstrate that HySRL surpasses state-of-the-art online RL baseline.
Forward citations
Cited by 3 Pith papers
-
Contextual Online Pricing with (Biased) Offline Data
Contextual online pricing can safely incorporate biased offline data, achieving the standard square-root-of-T worst-case regret and better rates when the offline bias is small.
-
Transfer Learning Across Policy Regimes in Adaptive Multi-Agent Systems
Restricting a learner to structural knowledge from a prior policy regime improves small-sample performance when the target preserves that structure and produces negative transfer when a threshold break moves the targe...
-
Pessimism Principle Can Be Effective: Towards a Framework for Zero-Shot Transfer Reinforcement Learning
A pessimism-based framework for zero-shot transfer RL builds conservative proxies from robust MDPs, yielding lower-bound performance guarantees and distributed algorithms that mitigate negative transfer.
Discussion (0). Continue with ORCID to comment.