Pith. sign in

REVIEW 2 cited by

Classical Structured Prediction Losses for Sequence to Sequence Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1711.04956 v5 pith:QCYXMCVF submitted 2017-11-14 cs.CL

classification cs.CL
keywords sequencemodelsbeambeenclassicallikelossesneural
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

There has been much recent work on training neural attention models at the sequence-level using either reinforcement learning-style methods or by optimizing the beam. In this paper, we survey a range of classical objective functions that have been widely used to train linear models for structured prediction and apply them to neural sequence to sequence models. Our experiments show that these losses can perform surprisingly well by slightly outperforming beam search optimization in a like for like setup. We also report new state of the art results on both IWSLT'14 German-English translation as well as Gigaword abstractive summarization. On the larger WMT'14 English-French translation task, sequence-level training achieves 41.5 BLEU which is on par with the state of the art.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Transformer Dissection: A Unified Understanding of Transformer's Attention via the Lens of Kernel

    cs.LG 2019-08 conditional novelty 6.0 of 10

    Transformer attention is reframed as kernel smoothing, and a product of symmetric kernels for features and positions achieves competitive performance on neural machine translation and sequence prediction.

  2. Neural Text Generation with Unlikelihood Training

    cs.LG 2019-08 conditional novelty 6.0 of 10

    Training neural language models with an unlikelihood objective that penalizes repeated and frequent tokens reduces degenerate, repetitive text while preserving quality.

Pith tools