REVIEW 4 cited by
Leveraging Pre-trained Checkpoints for Sequence Generation Tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Unsupervised pre-training of large neural models has recently revolutionized Natural Language Processing. By warm-starting from the publicly released checkpoints, NLP practitioners have pushed the state-of-the-art on multiple benchmarks while saving significant amounts of compute time. So far the focus has been mainly on the Natural Language Understanding tasks. In this paper, we demonstrate the efficacy of pre-trained checkpoints for Sequence Generation. We developed a Transformer-based sequence-to-sequence model that is compatible with publicly available pre-trained BERT, GPT-2 and RoBERTa checkpoints and conducted an extensive empirical study on the utility of initializing our model, both encoder and decoder, with these checkpoints. Our models result in new state-of-the-art results on Machine Translation, Text Summarization, Sentence Splitting, and Sentence Fusion.
Forward citations
Cited by 4 Pith papers
-
Text Summarization with Pretrained Encoders
BERTSUM adapts BERT for summarization with per-sentence [CLS] tokens, interval segment embeddings, separate optimizers, and two-stage fine-tuning, reaching state-of-the-art ROUGE scores on CNN/DailyMail, NYT, and XSum.
-
Ad Headline Generation using Self-Critical Masked Language Model
Self-critical policy-gradient training of a multi-product BERT MLM produces ad headlines that beat LSTM+RL baselines and human submissions on overlap metrics and blind quality/grammar audits.
-
COBART: Controlled, Optimized, Bidirectional and Auto-Regressive Transformer for Ad Headline Generation
Prefixing BART encoder input with bucketized CTR and length control tokens lets a single fine-tuned model generate ad headlines with controllable length and higher estimated CTR than prior baselines.
-
Pre-training A Neural Language Model Improves The Sample Efficiency of an Emergency Room Classification Model
Self-supervised pre-training on unlabeled clinical notes reduced the labeled data needed for ER trauma classification by roughly 10x, but the estimate is weakened by test-set-based model selection and missing variance.
Discussion (0). Continue with ORCID to comment.