Pith. sign in

REVIEW 2 cited by

Extrapolating Beyond Suboptimal Demonstrations via Inverse Reinforcement Learning from Observations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1904.06387 v5 pith:62SZW72Y submitted 2019-04-12 cs.LG stat.ML

classification cs.LGstat.ML
keywords learningdemonstrationsdemonstratorreinforcementrewardt-rexbeyondinverse
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A critical flaw of existing inverse reinforcement learning (IRL) methods is their inability to significantly outperform the demonstrator. This is because IRL typically seeks a reward function that makes the demonstrator appear near-optimal, rather than inferring the underlying intentions of the demonstrator that may have been poorly executed in practice. In this paper, we introduce a novel reward-learning-from-observation algorithm, Trajectory-ranked Reward EXtrapolation (T-REX), that extrapolates beyond a set of (approximately) ranked demonstrations in order to infer high-quality reward functions from a set of potentially poor demonstrations. When combined with deep reinforcement learning, T-REX outperforms state-of-the-art imitation learning and IRL methods on multiple Atari and MuJoCo benchmark tasks and achieves performance that is often more than twice the performance of the best demonstration. We also demonstrate that T-REX is robust to ranking noise and can accurately extrapolate intention by simply watching a learner noisily improve at a task over time.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Training Language Models to Self-Correct via Reinforcement Learning

    cs.LG 2024-09 unverdicted novelty 6.0 of 10

    SCoRe uses multi-turn online RL with regularization on self-generated traces to improve LLM self-correction, achieving 15.6% and 9.1% gains on MATH and HumanEval for Gemini models.

  2. Learning Implicit Social Navigation Behavior using Deep Inverse Reinforcement Learning

    cs.RO 2025-01 conditional novelty 4.0 of 10

    S-MEDIRL, a deep inverse RL method with a bilateral filtering smoothing loss and demonstration extrapolation, learns to yield and avoid deadlock in a narrow crossing, reaching about 92% success.

Pith tools