Incorporating Copying Mechanism in Sequence-to-Sequence Learning

Hang Li; Jiatao Gu; Victor O.K. Li; Zhengdong Lu

arxiv: 1603.06393 · v3 · pith:4SMREGGWnew · submitted 2016-03-21 · 💻 cs.CL · cs.AI· cs.LG· cs.NE

Incorporating Copying Mechanism in Sequence-to-Sequence Learning

Jiatao Gu , Zhengdong Lu , Hang Li , Victor O.K. Li This is my paper

classification 💻 cs.CL cs.AIcs.LGcs.NE

keywords copyingcopynetsequencelearningseq2seqdataexampleinput

0 comments

read the original abstract

We address an important problem in sequence-to-sequence (Seq2Seq) learning referred to as copying, in which certain segments in the input sequence are selectively replicated in the output sequence. A similar phenomenon is observable in human language communication. For example, humans tend to repeat entity names or even long phrases in conversation. The challenge with regard to copying in Seq2Seq is that new machinery is needed to decide when to perform the operation. In this paper, we incorporate copying into neural network-based Seq2Seq learning and propose a new model called CopyNet with encoder-decoder structure. CopyNet can nicely integrate the regular way of word generation in the decoder with the new copying mechanism which can choose sub-sequences in the input sequence and put them at proper places in the output sequence. Our empirical study on both synthetic data sets and real world data sets demonstrates the efficacy of CopyNet. For example, CopyNet can outperform regular RNN-based model with remarkable margins on text summarization tasks.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Pointer Sentinel Mixture Models
cs.CL 2016-09 conditional novelty 7.0

Pointer sentinel-LSTM mixes context copying with softmax prediction to reach 70.9 perplexity on Penn Treebank using fewer parameters than standard LSTMs.