REVIEW 3 cited by
Offline Learning from Demonstrations and Unlabeled Experience
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Behavior cloning (BC) is often practical for robot learning because it allows a policy to be trained offline without rewards, by supervised learning on expert demonstrations. However, BC does not effectively leverage what we will refer to as unlabeled experience: data of mixed and unknown quality without reward annotations. This unlabeled data can be generated by a variety of sources such as human teleoperation, scripted policies and other agents on the same robot. Towards data-driven offline robot learning that can use this unlabeled experience, we introduce Offline Reinforced Imitation Learning (ORIL). ORIL first learns a reward function by contrasting observations from demonstrator and unlabeled trajectories, then annotates all data with the learned reward, and finally trains an agent via offline reinforcement learning. Across a diverse set of continuous control and simulated robotic manipulation tasks, we show that ORIL consistently outperforms comparable BC agents by effectively leveraging unlabeled experience.
Forward citations
Cited by 3 Pith papers
-
RENEW: Towards Learning World Models and Repairing Model Exploitation from Preferences
Binary preferences over single imagined transitions can supervise a world model's dynamics, and uncertainty-directed querying (RENEW) reduces label cost on small discrete and control benchmarks, under a synthetic oracle.
-
Value from Observations: Towards Large-Scale Imitation Learning via Self-Improvement
VfO trains a state-value function on action-free expert demonstrations mixed with lower-quality background data, then uses advantage-weighted regression on the background data to improve the agent, approaching oracle ...
-
TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning
TROFI learns a reward model from ranked trajectories, labels an offline dataset with it, and trains a TD3+BC policy, matching ground-truth-reward performance on many D4RL tasks without a hand-coded reward or expert de...
Discussion (0). Continue with ORCID to comment.