Pith. sign in

REVIEW 1 cited by

Learning to Search Better Than Your Teacher

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1502.02206 v2 pith:ZCXYMDVH submitted 2015-02-08 cs.LG stat.ML

classification cs.LGstat.ML
keywords referencelearningpolicysearchstructuredapplicationscomparedguarantees
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Methods for learning to search for structured prediction typically imitate a reference policy, with existing theoretical guarantees demonstrating low regret compared to that reference. This is unsatisfactory in many applications where the reference policy is suboptimal and the goal of learning is to improve upon it. Can learning to search work even when the reference is poor? We provide a new learning to search algorithm, LOLS, which does well relative to the reference policy, but additionally guarantees low regret compared to deviations from the learned policy: a local-optimality guarantee. Consequently, LOLS can improve upon the reference policy, unlike previous algorithms. This enables us to develop structured contextual bandits, a partial information structured prediction setting with many potential applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. An Empirical Comparison on Imitation Learning and Reinforcement Learning for Paraphrase Generation

    cs.CL 2019-08 conditional novelty 5.0 of 10

    DAgger-style imitation learning outperforms REINFORCE reinforcement learning for paraphrase generation with a pointer-generator, and the best model reaches state-of-the-art scores on Quora.

Pith tools