Pith. sign in

REVIEW 3 cited by

Offline Learning from Demonstrations and Unlabeled Experience

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.13885 v1 pith:HFMSPIDM submitted 2020-11-27 cs.LG cs.AIcs.ROstat.ML

classification cs.LGcs.AIcs.ROstat.ML
keywords learningunlabeledofflineexperiencedataorilrewardrobot
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Behavior cloning (BC) is often practical for robot learning because it allows a policy to be trained offline without rewards, by supervised learning on expert demonstrations. However, BC does not effectively leverage what we will refer to as unlabeled experience: data of mixed and unknown quality without reward annotations. This unlabeled data can be generated by a variety of sources such as human teleoperation, scripted policies and other agents on the same robot. Towards data-driven offline robot learning that can use this unlabeled experience, we introduce Offline Reinforced Imitation Learning (ORIL). ORIL first learns a reward function by contrasting observations from demonstrator and unlabeled trajectories, then annotates all data with the learned reward, and finally trains an agent via offline reinforcement learning. Across a diverse set of continuous control and simulated robotic manipulation tasks, we show that ORIL consistently outperforms comparable BC agents by effectively leveraging unlabeled experience.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RENEW: Towards Learning World Models and Repairing Model Exploitation from Preferences

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Binary preferences over single imagined transitions can supervise a world model's dynamics, and uncertainty-directed querying (RENEW) reduces label cost on small discrete and control benchmarks, under a synthetic oracle.

  2. Value from Observations: Towards Large-Scale Imitation Learning via Self-Improvement

    cs.LG 2025-07 conditional novelty 6.0 of 10

    VfO trains a state-value function on action-free expert demonstrations mixed with lower-quality background data, then uses advantage-weighted regression on the background data to improve the agent, approaching oracle ...

  3. TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning

    cs.LG 2025-06 conditional novelty 5.0 of 10

    TROFI learns a reward model from ranked trajectories, labels an offline dataset with it, and trains a TD3+BC policy, matching ground-truth-reward performance on many D4RL tasks without a hand-coded reward or expert de...

Pith tools