Pith. sign in

REVIEW 2 cited by

Blending Imitation and Reinforcement Learning for Robust Policy Improvement

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.01737 v3 pith:FOGN255A submitted 2023-10-03 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords learningoraclesperformancepolicyimitationrobustacrossdomains
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While reinforcement learning (RL) has shown promising performance, its sample complexity continues to be a substantial hurdle, restricting its broader application across a variety of domains. Imitation learning (IL) utilizes oracles to improve sample efficiency, yet it is often constrained by the quality of the oracles deployed. which actively interleaves between IL and RL based on an online estimate of their performance. RPI draws on the strengths of IL, using oracle queries to facilitate exploration, an aspect that is notably challenging in sparse-reward RL, particularly during the early stages of learning. As learning unfolds, RPI gradually transitions to RL, effectively treating the learned policy as an improved oracle. This algorithm is capable of learning from and improving upon a diverse set of black-box oracles. Integral to RPI are Robust Active Policy Selection (RAPS) and Robust Policy Gradient (RPG), both of which reason over whether to perform state-wise imitation from the oracles or learn from its own value function when the learner's performance surpasses that of the oracles in a specific state. Empirical evaluations and theoretical analysis validate that RPI excels in comparison to existing state-of-the-art methodologies, demonstrating superior performance across various benchmark domains.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DexTrack: Towards Generalizable Neural Tracking Control for Dexterous Manipulation from Human References

    cs.RO 2025-02 conditional novelty 6.0 of 10

    A neural controller combining RL and imitation learning on iteratively mined demonstrations tracks human kinematic references for dexterous manipulation, yielding over 10% higher success rates than prior baselines.

  2. DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization

    cs.LG 2025-02 conditional novelty 5.0 of 10

    SPO fine-tunes an LLM to rewrite drug molecules into analogs that score higher on docking, drug-likeness, solubility, and synthesizability, beating baselines on two protein targets.

Pith tools