Pith. sign in

REVIEW 1 cited by

Z-Code++: A Pre-trained Language Model Optimized for Abstractive Summarization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.09770 v2 pith:PXUYRHHM submitted 2022-08-21 cs.CL cs.AI

classification cs.CLcs.AI
keywords modelsummarizationpre-trainedtextlanguagez-codeabstractivecorpora
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents Z-Code++, a new pre-trained language model optimized for abstractive text summarization. The model extends the state of the art encoder-decoder model using three techniques. First, we use a two-phase pre-training process to improve model's performance on low-resource summarization tasks. The model is first pre-trained using text corpora for language understanding, and then is continually pre-trained on summarization corpora for grounded text generation. Second, we replace self-attention layers in the encoder with disentangled attention layers, where each word is represented using two vectors that encode its content and position, respectively. Third, we use fusion-in-encoder, a simple yet effective method of encoding long sequences in a hierarchical manner. Z-Code++ creates new state of the art on 9 out of 13 text summarization tasks across 5 languages. Our model is parameter-efficient in that it outperforms the 600x larger PaLM-540B on XSum, and the finetuned 200x larger GPT3-175B on SAMSum. In zero-shot and few-shot settings, our model substantially outperforms the competing models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Parameter-Efficient Fine-Tuning for Foundation Models

    cs.CL 2025-01 conditional novelty 2.0 of 10

    A survey that categorizes and summarizes parameter-efficient fine-tuning methods across large language, vision, and multimodal models.

Pith tools