Pith. sign in

REVIEW 3 major objections 6 minor 120 references

Survey on Abstractive Text Summarization: Dataset, Models, and Metrics

T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Small test: AI summary fact errors down from 30 percent

desk verdict Useful but uneven survey; the empirical claim of reduced factual inconsistency rests on a metric mismatch. read the letter →

arxiv 2412.17165 v1 pith:N3E2FOF3 submitted 2024-12-22 cs.AI

classification cs.AI
keywords abstractivetextsummarizationfactualconsistencytransformermodelsevaluationmetricsmulti-documentlongdocumentdatasets
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey maps the current landscape of abstractive text summarization, the task of having a model paraphrase a document into a short, fluent, factually consistent summary, by reviewing datasets, models, and evaluation metrics. Its experimental part runs public transformer checkpoints on small samples of short, long, and multi-document inputs and reports that the outputs score high on automatic metrics and on a GPT-2-based fact-checking model. The central empirical claim is that factual inconsistency has dropped significantly relative to the roughly 30 percent rate cited in earlier work. The authors argue that pretraining objectives, finetuning domain, model size, and distilled knowledge are the main drivers of the improvement. A careful reader should note that the long and multi-document fact checks were run with the reference summary standing in for the source evidence.

What carries the argument

The load-bearing machinery is the transformer encoder-decoder architecture with pretraining objectives such as masked language modeling, gap-sentence generation, and causal language modeling, plus the sparse and local-global attention variants that extend input length to about 16,000 tokens. On top of this, the experiment uses a GPT-2-based fact-checking classifier that returns the probability that a generated summary, treated as a claim, is entailed by an evidence text. For long and multi-document cases, the evidence is substituted with the reference summary when the source is too long, and that substitution is what allows the conclusion about reduced factual inconsistency to be drawn at all.

What would settle it

Run the same fact-checking protocol on a much larger sample and, for long and multi-document inputs, compare FactCheck scores when the evidence is the original source document versus the reference summary; if source-evidence scores fall substantially below reference-evidence scores, the claimed reduction in factual inconsistency is largely an artifact of the evidence substitution.

Watch

Extended reading notes

Core claim

The paper claims that, on its test cases, factual inconsistency in abstractive summarization has fallen to a small fraction of the roughly 30 percent rate reported by earlier studies, with average FactCheck scores of 0.93 to 1.00 for short-document summaries, 0.795 to 0.986 for long-document summaries, and 0.835 to 0.93 for multi-document summaries. It also claims that model size, knowledge distillation, and finetuning on multiple domains tend to improve scores, and that pretrained transformer models such as BART, PEGASUS, Longformer, LongT5, PRIMERA, CENTRUM, and REFLECT can produce fluent summaries, with repetition largely controlled by no-repeat n-gram settings. The survey's broader claim is that the field can now be mapped by task type, dataset domain, and evaluation dimension, even though factual faithfulness remains an open challenge for long and multi-document inputs.

Load-bearing premise

The conclusion that factual inconsistency has dropped depends on assuming the reference summary is factually consistent and can serve as evidence for the source, and on a very small number of test documents.

Editorial extensions

If this is right

  • If the reported FactCheck scores hold up, modern pretrained summarizers may have reduced the hallucination problem to a small fraction of its earlier level, at least for short news-like documents.
  • Finetuning on multiple domains and using distilled versions of large models appear to be reliable routes to better summarization scores, not just larger parameter counts.
  • Factual consistency of long and multi-document summaries cannot yet be measured directly, so better evidence-based checkers would be needed before deploying these models.
  • ROUGE-style overlap metrics remain the default evaluation, with semantic and factual metrics playing a supporting role, so model rankings could shift if factuality were weighted more heavily.
  • The taxonomy of tasks by input length and document count gives practitioners a way to choose a model family appropriate to their use case.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The 'reduced significantly' conclusion is fragile because it rests on only 7, 2, and 10 test documents, and a larger evaluation could move the averages considerably.
  • For long and multi-document outputs, using the reference summary as evidence likely inflates FactCheck scores, because the reference is already a clean paraphrase; comparing source-evidence scores would quantify this inflation.
  • A testable extension would be to run the same protocol across a broad multi-domain benchmark to see whether the apparent drop in factual inconsistency is domain-dependent, especially outside news.
  • If the evidence-substitution gap turns out to be large, the practical takeaway changes from 'hallucination is mostly solved' to 'we still lack a reliable way to measure faithfulness for long inputs.'
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper is a survey of abstractive text summarization, covering task definitions, extractive/abstractive/hybrid approaches, transformer-based models for short, long, and multi-document inputs, datasets, and automatic evaluation metrics. It also reports small-scale experiments on publicly available fine-tuned checkpoints (BART, PEGASUS, Longformer, LongT5, PRIMERA, CENTRUM, REFLECT) evaluated with ROUGE, METEOR, CHRF, BertScore, and a FactCheck probability score. The central empirical claim, stated in Sections 6.1, 6.2, and 7, is that factual inconsistency in generated summaries has "reduced significantly" relative to the approximately 30% inconsistency rate reported by references [46-49].

Significance. If the survey's map of the field is accurate, it provides a useful organized introduction to abstractive summarization models, datasets, and metrics, particularly for readers outside the area. A notable strength is that the authors release the code, data, and per-sample results, which supports reproducibility. The empirical sections are best read as illustrative demonstrations of public checkpoints rather than as rigorous comparative evaluations. The claimed large reduction in factual inconsistency would be significant if supported, but it is not supported by the current evidence because the FactCheck score is an uncalibrated average probability that is not commensurable with the 30% inconsistency baseline, and the sample sizes of 7, 2, and 10 documents are far too small to support the word "significantly".

major comments (3)
  1. [§6.1, §6.2, §7; Tables 2-4] The conclusion that factual inconsistency "has reduced significantly compared to the ~30% factual inconsistency reported by [46-49]" is not supported by the reported FactCheck numbers. The FactCheck column is an average probability assigned by a FEVER-trained GPT-2 NLI model, while a 30% inconsistency rate is a proportion of summaries classified as inconsistent under some scoring scheme. An average probability of 0.93 can coexist with a high proportion of low-probability summaries, so the comparison is meaningful only if the paper specifies and applies a threshold (e.g., probability below 0.5 counts as inconsistent) and reports the resulting rate, along with some check of the model's calibration. Without this, the average FactCheck value in Tables 2, 3, and 4 cannot be converted into an inconsistency percentage, and the central claim in Section 7 is not established.
  2. [§6] For long and multi-document summaries, the reference summary of the source document is used as evidence for the factuality check, under the stated assumption that the reference summary is factually consistent in entities and entity relations. This assumption is load-bearing for the long- and multi-document claims in Tables 3 and 4. Reference summaries are not guaranteed to be factually complete or consistent, and the paper provides no validation of this assumption. Consequently, the FactCheck scores for long and multi-document outputs measure consistency with the reference summary rather than faithfulness to the source documents, and the strong claim of reduced factual inconsistency cannot be drawn from them.
  3. [§6.1 and §6.2] The sample sizes are 7 short documents, 2 long documents, and 10 multi-document clusters. The paper reports only means, with no confidence intervals, standard deviations, per-model statistical tests, or per-sample dispersion. Statements such as "M s2 shows 100% factually consistent summaries score" and "the factuality problem ... has reduced significantly" are therefore not statistically supported; the study is too small to detect meaningful differences among models or to compare reliably with historical rates.
minor comments (6)
  1. [Abstract and author affiliation] The affiliation for Flavio Bertini is given as "University of Parma" in the affiliation block but the contact line says "flavio.bertini@unipr.it"; the reader report uses "Favio" as a first name. Please verify the spelling and institutional attribution.
  2. [Keyword line] The keyword line contains the typo "Estractive" (should be "Extractive").
  3. [Section 2.3, paragraph on abstractive summarization] The sentence "but the factuality of the summries produce is still a challenge" contains spelling errors and should be rewritten, e.g., "but the factuality of the summaries produced is still a challenge."
  4. [Section 6.2] The text states that "M l2 outperforms other finetuned model" when referring to the multi-document results; the model labels are M m1, M m2, and M m3, so this should read "M m2."
  5. [Sections 3.1 and 5] The paper uses inconsistent capitalization and spacing for model names (e.g., "P EGASU SLARGE", "Tranformer", "Rouge"). A consistent formatting pass across the text and tables would improve readability.
  6. [Table 1] The table caption calls the table "The breakdown of the experimented finetuned models," but several listed models are not covered by the later experiments (e.g., Gigaword, Wikihow, Reddit TIFU, BookSum); consider either expanding the experiments or revising the table caption to indicate which models and datasets were actually tested.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the paper performs no derivation or parameter fitting, and its empirical claim, while methodologically fragile, is not defined in terms of the conclusion it asserts.

full rationale

This is a survey with small illustrative experiments, not a derivation chain. No model is fitted by the authors; they evaluate publicly available Hugging Face checkpoints (BART, PEGASUS, Longformer, LongT5, PRIMERA, CENTRUM, REFLECT) with external metrics (ROUGE, METEOR, CHRF, BertScore) and an external fact-checking model (fractalego/fact-checking, a GPT-2 model trained on FEVER). The paper therefore contains no self-definitional step, no fitted parameter renamed as a prediction, and no uniqueness claim imported from the authors' own prior work. The only self-citation is to the authors' GitHub repository for data and code, which is data-availability boilerplate and is not load-bearing for any stated conclusion. The central empirical claim that factual inconsistency 'has reduced significantly compared to the ~30% factual inconsistency reported by [46-49]' is unsupported because an average continuous FactCheck probability is compared to a proportion of inconsistent summaries without applying a classification threshold, and the FactCheck model's calibration for summarization is unknown. That is a validity or correctness problem, not circularity: the measured score is not defined in terms of the claim, and no input quantity is reused as the output. The paper also explicitly discloses that for long and multi-document cases the reference summary was used as evidence under an assumption of factual consistency; this weakens the faithfulness measurement but does not make the reported score equivalent to the conclusion by construction. Per the stated rules, an unsupported comparison is not a circular step, and the honest finding is therefore no significant circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The experimental claims rest on external assumptions rather than new axioms: the fact-checking model [85] is taken as a valid measure of factual consistency; for long and multi-document cases the reference summary is taken as factually consistent evidence; the small test samples are treated as representative; and automatic metrics are treated as quality indicators. None of these is derived in the paper, and the first two directly shape the reported FactCheck scores.

assumptions (4)
  • domain assumption The GPT-2-based fact-checking model from [85] provides a valid measure of summary factual consistency.
    Used in Section 6 to compute FactCheck scores; the paper does not validate this checker against human judgments on summarization outputs.
  • ad hoc to paper The reference summary of a document is factually consistent and can serve as evidence for long and multi-document factuality checks.
    Stated in Section 6: 'it is assumed that the reference of the source document is a factual consistent summary.' This assumption is load-bearing for FactCheck values in Tables 3 and 4.
  • ad hoc to paper Seven short, two long, and ten multi-document samples are representative enough to compare model performance.
    No sampling protocol is provided in Sections 6.1 and 6.2, so representativeness is assumed rather than demonstrated.
  • domain assumption Automatic metrics (ROUGE, METEOR, CHRF, BERTScore) and the fact-checking score capture the quality dimensions claimed in Section 4.
    The paper itself notes human evaluation is the gold standard; experimental conclusions rely on automatic metrics without human validation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Survey on Abstractive Text Summarization: Dataset, Models, and Metrics." pith.science (2026). https://pith.science/paper/N3E2FOF3

@misc{pith2026241217165,
  author       = {Pith},
  title        = {Pith review of: Survey on Abstractive Text Summarization: Dataset, Models, and Metrics},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/N3E2FOF3}},
  note         = {Machine review of arXiv:2412.17165}
}
read the original abstract

The advancements in deep learning, particularly the introduction of transformers, have been pivotal in enhancing various natural language processing (NLP) tasks. These include text-to-text applications such as machine translation, text classification, and text summarization, as well as data-to-text tasks like response generation and image-to-text tasks such as captioning. Transformer models are distinguished by their attention mechanisms, pretraining on general knowledge, and fine-tuning for downstream tasks. This has led to significant improvements, particularly in abstractive summarization, where sections of a source document are paraphrased to produce summaries that closely resemble human expression. The effectiveness of these models is assessed using diverse metrics, encompassing techniques like semantic overlap and factual correctness. This survey examines the state of the art in text summarization models, with a specific focus on the abstractive summarization approach. It reviews various datasets and evaluation metrics used to measure model performance. Additionally, it includes the results of test cases using abstractive summarization models to underscore the advantages and limitations of contemporary transformer-based models. The source codes and the data are available at https://github.com/gospelnnadi/Text-Summarization-SOTA-Experiment.

Figures

Figures reproduced from arXiv: 2412.17165 by the authors.

Figure 1
Figure 1. World report on the number of citable documents, reported by Scimagojr. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Scopus Report of all scientific articles Covid pubblished from year 2015 to 2024. [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. 1. Extractive: One of the early attempts at automatic text summarization in 1985 involved extracting important sentences based on word frequency [5]. This method aims to identify sentences from the source document verbatim that capture its essence. 2. Abstractive: The abstractive approach involves crafting new sentences by paraphrasing sections of the source document, aiming to condense the text more dynamically. Th… view at source ↗
Figures from the paper (1 more)
Figure 3
Figure 3. Figure 3: Text summarization approach. In this article, we guide the reader through the state of the art models in text summarization, exploring the diverse datasets used for pretraining and finetuning. Additionally, we delve into various metrics employed to evaluate the quality…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

120 extracted references · 67 canonical work pages

  1. [1]

    Text summarization sota experiment

    Github repository. Text summarization sota experiment. https://github.com/data-lang/ Text-Summarization-SOTA-Experiment.git , 2023

  2. [2]

    World report: Citable documents, 2022

    ScimagoJR. World report: Citable documents, 2022. https://www.scimagojr.com/worldreport.php, 2022

  3. [3]

    Report of all scientific articles on covid published from 2015 to 2024

    Scopus. Report of all scientific articles on covid published from 2015 to 2024. https://www.scopus.com/, 2024

  4. [4]

    Congbo Ma, Wei Emma Zhang, Mingyu Guo, Hu Wang, and Quan Z. Sheng. Multi-document summarization via deep learning techniques: A survey. arXiv preprint arXiv:, 2021

  5. [5]

    H. P. Luhn. The automatic creation of literature abstracts. IBM Journal of Research and Development, 2(2):159– 165, 1958

  6. [6]

    Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension

    Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. 2019. 18 Survey on Abstractive Text Summarization: Dataset, Models, and Metrics

  7. [7]

    Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter J. Liu. Pegasus: Pre-training with extracted gap- sentences for abstractive summarization. 2020

  8. [8]

    Peters, and Arman Cohan

    Iz Beltagy, Matthew E. Peters, and Arman Cohan. Longformer: The long-document transformer. arXiv preprint arXiv:, 2020

Show all 120 references
  1. [9]

    Big bird: Transformers for longer sequences

    Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed. Big bird: Transformers for longer sequences. arXiv preprint arXiv:, 2021

  2. [10]

    Longt5: Efficient text-to-text transformer for long sequences

    Mandy Guo, Joshua Ainslie, David Uthus, Santiago Ontañón, Jianmo Ni, Yun-Hsuan Sung, and Yinfei Yang. Longt5: Efficient text-to-text transformer for long sequences. In Association for Computational Linguistics, 2022

  3. [11]

    Gomez, Łukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. arXiv preprint arXiv:, 2017

  4. [12]

    Bert: Pre-training of deep bidirectional transformers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Langu...

  5. [13]

    Primera: Pyramid-based masked sentence pre-training for multi-document summarization

    Wen Xiao, Iz Beltagy, Giuseppe Carenini, and Arman Cohan. Primera: Pyramid-based masked sentence pre-training for multi-document summarization. arXiv preprint arXiv:, 2022

  6. [14]

    Antognini and B

    D. Antognini and B. Faltings. Learning to create sentence semantic relation graphs for multi-document summa- rization. In L. Wang, J. C. K. Cheung, G. Carenini, and F. Liu, editors, Proceedings of the 2nd Workshop on New Frontiers in Summarization, pages 32–41, Hong Kong, Chin...

  7. [15]

    M. T. Nayeem, T. A. Fuad, and Y . Chali. Abstractive unsupervised multi-document summarization using paraphrastic sentence fusion. In E. M. Bender, L. Derczynski, and P. Isabelle, editors, Proceedings of the 27th International Conference on Computational Linguistics, pages 119...

  8. [16]

    Yasunaga, R

    M. Yasunaga, R. Zhang, K. Meelu, A. Pareek, K. Srinivasan, and D. Radev. Graph-based neural multi-document summarization. In R. Levy and L. Specia, editors, Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017), pages 452–462, Vancouver, Ca...

  9. [17]

    Song, Y .-S

    Y .-Z. Song, Y .-S. Chen, and H.-H. Shuai. Improving multi-document summarization through referenced flexible extraction with credit-awareness. 2022

  10. [18]

    Efficiently summarizing text and graph encodings of multi-document clusters

    Ramakanth Pasunuru, Mengwen Liu, Mohit Bansal, Sujith Ravi, and Markus Dreyer. Efficiently summarizing text and graph encodings of multi-document clusters. In Association for Computational Linguistics, 2021

  11. [19]

    Learning to extract coherent summary via deep reinforcement learning

    Yuxiang Wu and Baotian Hu. Learning to extract coherent summary via deep reinforcement learning. In Thirty-Second AAAI Conference on Artificial Intelligence, 2018

  12. [20]

    Banditsum: Extractive summarization as a contextual bandit

    Yue Dong, Yikang Shen, Eric Crawford, Herke van Hoof, and Jackie Chi Kit Cheung. Banditsum: Extractive summarization as a contextual bandit. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3739–3748, Brussels, Belgium, 2018. Ass...

  13. [21]

    Neural document summa- rization by jointly learning to score and select sentences

    Qingyu Zhou, Nan Yang, Furu Wei, Shaohan Huang, Ming Zhou, and Tiejun Zhao. Neural document summa- rization by jointly learning to score and select sentences. In ACL 2018, pages 654–663, Melbourne, Australia,

  14. [22]

    Neural latent extractive document summarization

    Xingxing Zhang, Mirella Lapata, Furu Wei, and Ming Zhou. Neural latent extractive document summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 779–784, Brussels, Belgium, 2018. Association for Computational Linguistics

  15. [23]

    Cohen, and Mirella Lapata

    Shashi Narayan, Shay B. Cohen, and Mirella Lapata. Ranking sentences for extractive summarization with reinforcement learning. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Vol...

  16. [24]

    Williams

    Ronald J. Williams. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning, 8(3-4):229–256, 1992

  17. [25]

    Neural extractive text summarization with syntactic compression

    Jiacheng Xu and Greg Durrett. Neural extractive text summarization with syntactic compression. In EMNLP- IJCNLP 2019, pages 3292–3303, Hong Kong, China, 2019. Association for Computational Linguistics. 19 Survey on Abstractive Text Summarization: Dataset, Models, and Metrics

  18. [26]

    Strass: A light and effective method for extractive summarization based on sentence embeddings

    Leo Bouscarrat, Antoine Bonnefoy, Thomas Peel, and Cecile Pereira. Strass: A light and effective method for extractive summarization based on sentence embeddings. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: Student Research Works...

  19. [27]

    A novel extractive multi-document text summarization system using quantum-inspired genetic algorithm: Mtsqiga

    Mohammad Mojrian and Seyedabolghasem Mirroshandel. A novel extractive multi-document text summarization system using quantum-inspired genetic algorithm: Mtsqiga. Expert Systems with Applications, 171:114555, 01 2021

  20. [28]

    Alguliev, R.M

    R.M. Alguliev, R.M. Aliguliyev, and N.R. Isazade. Cdds: Constraint-driven document summarization models. Expert Systems with Applications, 40:458–465, 2013

  21. [29]

    Alguliyev, R

    R. Alguliyev, R. Aliguliyev, and N. Isazade. An unsupervised approach to generating generic summaries of documents. Applied Soft Computing, 34, 2015

  22. [30]

    Alguliyev, R

    R. Alguliyev, R. Aliguliyev, M. Hajirahimova, and C. Mehdiyev. Mcmr: Maximum coverage and minimum redundant text summarization model. Expert Systems with Applications, 38:14514–14522, Nov 2011

  23. [31]

    Y . Liu, X. Wang, J. Zhang, and H. Xu. Personalized pagerank based multi-document summarization. InIEEE International Workshop on Semantic Computing and Systems, pages 169–173, 2008

  24. [32]

    Alguliyev, R

    R. Alguliyev, R. Aliguliyev, and C. Mehdiyev. Sentence selection for generic document summarization using an adaptive differential evolution algorithm. Swarm and Evolutionary Computation, 1:213–222, Dec 2011

  25. [33]

    Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer

    Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. Deep contextualized word representations. arXiv preprint arXiv:, 2018

  26. [34]

    Abstractive text summarization using sequence-to-sequence rnns and beyond

    Ramesh Nallapati, Bowen Zhou, Caglar Gulcehre, and Bing Xiang. Abstractive text summarization using sequence-to-sequence rnns and beyond. In Proceedings of SIGNLL Conference on Computational Natural Language Learning (CoNLL), 2016

  27. [35]

    Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies

    Max Grusky, Mor Naaman, and Yoav Artzi. Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, volume 1 ...

  28. [36]

    Narayan, S

    S. Narayan, S. B. Cohen, and M. Lapata. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1797–1807, Brussels, Belgi...

  29. [37]

    An entity-driven framework for abstractive summarization

    Eva Sharma, Luyang Huang, Zhe Hu, and Lu Wang. An entity-driven framework for abstractive summarization. In Association for Computational Linguistics, 2019

  30. [38]

    Hierarchical transformers for multi-document summarization

    Yang Liu and Mirella Lapata. Hierarchical transformers for multi-document summarization. In Association for Computational Linguistics (ACL 2019), pages 5070–5081, Florence, Italy, 2019

  31. [39]

    Text summarization with pretrained encoders

    Yang Liu and Mirella Lapata. Text summarization with pretrained encoders. 2019

  32. [40]

    Liu, and Christopher D

    Abigail See, Peter J. Liu, and Christopher D. Manning. Get to the point: Summarization with pointer-generator networks. 2017

  33. [41]

    Cohan, F

    A. Cohan, F. Dernoncourt, D. S. Kim, T. Bui, S. Kim, W. Chang, and N. Goharian. A discourse-aware attention model for abstractive summarization of long documents. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguisti...

  34. [42]

    Soft layer-specific multi-task summarization with entailment and question generation

    Han Guo, Ramakanth Pasunuru, and Mohit Bansal. Soft layer-specific multi-task summarization with entailment and question generation. In Association for Computational Linguistics, 2018

  35. [43]

    Improving abstraction in text summarization

    Wojciech Kry´sci´nski, Romain Paulus, Caiming Xiong, and Richard Socher. Improving abstraction in text summarization. In Association for Computational Linguistics, 2018

  36. [44]

    Raffel, N

    C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y . Zhou, W. Li, and P. J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. 2023, 2023

  37. [45]

    Distillation knowledge applied on pegasus for summarization, 2019–2020

    Lorenzo Niccolai and Andrea Asperti. Distillation knowledge applied on pegasus for summarization, 2019–2020

  38. [46]

    Faithful to the original: Fact aware neural abstractive summarization

    Ziqiang Cao, Furu Wei, Wenjie Li, and Sujian Li. Faithful to the original: Fact aware neural abstractive summarization. In Proceedings of the Thirty Second AAAI Conference on Artificial Intelligence (AAAI-18), pages 4784–4791, New Orleans, Louisiana, USA, 2018. 20 Survey on Ab...

  39. [47]

    Liu, and Mohammad Saleh

    Ben Goodrich, Vinay Rao, Peter J. Liu, and Mohammad Saleh. Assessing the factual accuracy of generated text. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD 2019), pages 166–175, Anchorage, AK, USA, 2019

  40. [48]

    Tobias Falke, Leonardo F. R. Ribeiro, Prasetya Ajie Utama, Ido Dagan, and Iryna Gurevych. Ranking generated summaries by correctness: An interesting but challenging application for natural language inference. In Proceedings of the 57th Conference of the Association for Computa...

  41. [49]

    Neural text summarization: A critical evaluation

    Wojciech Kry´sci´nski, Nitish Shirish Keskar, Bryan McCann, Caiming Xiong, and Richard Socher. Neural text summarization: A critical evaluation. Salesforce Research, 2019

  42. [50]

    Long short-term memory

    Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural Computation, 9(8):1735–1780, 1997

  43. [51]

    Review summarization with pointer gen and bert, 2020

    Matthew Martin and Marjolein Pawlus. Review summarization with pointer gen and bert, 2020

  44. [52]

    A deep reinforced model for abstractive summarization

    Romain Paulus, Caiming Xiong, and Richard Socher. A deep reinforced model for abstractive summarization. arXiv preprint arXiv:1705.04304, 2017

  45. [53]

    Multi-reward reinforced summarization with saliency and entailment

    Ramakanth Pasunuru and Mohit Bansal. Multi-reward reinforced summarization with saliency and entailment. In Association for Computational Linguistics, 2018

  46. [54]

    Closed-book training to improve summarization encoder memory

    Yichen Jiang and Mohit Bansal. Closed-book training to improve summarization encoder memory. InAssociation for Computational Linguistics, 2018

  47. [55]

    Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B

    Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. Fine-tuning language models from human preferences. arXiv preprint arXiv:, 2020

  48. [56]

    Unified language model pre-training for natural language understanding and generation

    Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao- Wuen Hon. Unified language model pre-training for natural language understanding and generation. arXiv preprint arXiv:, 2019

  49. [57]

    Deep communicating agents for abstractive summarization

    Asli Celikyilmaz, Antoine Bosselut, Xiaodong He, and Yejin Choi. Deep communicating agents for abstractive summarization. In Association for Computational Linguistics, 2018

  50. [58]

    W. Li, X. Xiao, J. Liu, H. Wu, H. Wang, and J. Du. Leveraging graph to improve abstractive multi-document summarization. In D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault, editors, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pa...

  51. [59]

    Fast abstractive summarization with reinforce-selected sentence rewriting

    Yen-Chun Chen and Mohit Bansal. Fast abstractive summarization with reinforce-selected sentence rewriting. In Association for Computational Linguistics, 2018

  52. [60]

    Sebastian Gehrmann, Yuntian Deng, and Alexander M. Rush. Bottom-up abstractive summarization. 2018

  53. [61]

    A unified model for extractive and abstractive summarization using inconsistency loss

    Wan-Ting Hsu, Chieh-Kai Lin, Ming-Ying Lee, Kerui Min, Jing Tang, and Min Sun. A unified model for extractive and abstractive summarization using inconsistency loss. arXiv preprint arXiv:, 2018

  54. [62]

    Improving multi-document summarization through referenced flexible extraction with credit-awareness

    Yun-Zhu Song, Yi-Syuan Chen, and Hong-Han Shuai. Improving multi-document summarization through referenced flexible extraction with credit-awareness. arXiv preprint arXiv:, 2022

  55. [63]

    Noah A. Smith. Contextual word representations: Putting words into computers, 2020

  56. [64]

    Efficient estimation of word representations in vector space

    Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efficient estimation of word representations in vector space. arXiv preprint arXiv:, 2013

  57. [65]

    Jeffrey Pennington, Richard Socher, and Christopher D. Manning. Glove: Global vectors for word representation. In Association for Computational Linguistics, 2014

  58. [66]

    An introduction to convolutional neural networks

    Keiron O’Shea and Ryan Nash. An introduction to convolutional neural networks. ArXiv e-prints, 11 2015

  59. [67]

    Mike Schuster and Kuldip K. Paliwal. Bidirectional recurrent neural networks. IEEE Transactions on Signal Processing, 45(11):2673–2681, 1997

  60. [68]

    Huggingface

    HuggingFace. Huggingface. https://huggingface.co. Accessed: 2024-12-20

  61. [69]

    Survey of the state of the art in natural language generation: Core tasks, applications and evaluation

    Albert Gatt and Emiel Krahmer. Survey of the state of the art in natural language generation: Core tasks, applications and evaluation. arXiv preprint arXiv:1703.09902v4, 2018

  62. [70]

    Rouge: A package for automatic evaluation of summaries, 2004

    Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries, 2004

  63. [71]

    Learning to score system summaries for better content selection evaluation

    Maxime Peyrard, Teresa Botschen, and Iryna Gurevych. Learning to score system summaries for better content selection evaluation. In Association for Computational Linguistics, 2017. 21 Survey on Abstractive Text Summarization: Dataset, Models, and Metrics

  64. [72]

    Weinberger, and Yoav Artzi

    Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675v3, 2020

  65. [73]

    Meyer, and Steffen Eger

    Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. Moverscore: Text genera- tion evaluating with contextualized embeddings and earth mover distance. In Association for Computational Linguistics, 2019

  66. [74]

    Speeding up word mover’s distance and its variants via properties of distances between embeddings, 12 2019

    Matheus Werner and Eduardo Laber. Speeding up word mover’s distance and its variants via properties of distances between embeddings, 12 2019

  67. [75]

    Elizabeth Clark, Asli Celikyilmaz, and Noah A. Smith. Sentence mover’s similarity: Automatic evaluation for multi-sentence texts. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2748–2760, Florence, Italy, 2019. Association for...

  68. [76]

    Answers unite! unsupervised metrics for reinforced summarization models

    Thomas Scialom, Sylvain Lamprier, Benjamin Piwowarski, and Jacopo Staiano. Answers unite! unsupervised metrics for reinforced summarization models. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Confere...

  69. [77]

    Vasilyev, Vedant Dharnidharka, and John Bohannon

    Oleg V . Vasilyev, Vedant Dharnidharka, and John Bohannon. Fill in the blanc: human-free quality estimation of document summaries. CoRR, abs/2002.09836, 2020

  70. [78]

    Supert: towards new frontiers in unsupervised evaluation metrics for multi document summarization

    Yang Gao, Wei Zhao, and Steffen Eger. Supert: towards new frontiers in unsupervised evaluation metrics for multi document summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1347–1354, Online, 2020. Association for C...

  71. [79]

    Bleu: a method for automatic evaluation of machine translation

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In Association for Computational Linguistics, 2002

  72. [80]

    Chrf: Character n-gram f-score for automatic mt evaluation

    Maja Popovi´c. Chrf: Character n-gram f-score for automatic mt evaluation. In Association for Computational Linguistics, 2015

  73. [81]

    Meteor: An automatic metric for mt evaluation with high levels of correlation with human judgments

    Alon Lavie and Abhaya Agarwal. Meteor: An automatic metric for mt evaluation with high levels of correlation with human judgments. In Proceedings of the Second Workshop on Statistical Machine Translation , pages 228–231, Prague, Czech Republic, 2007. Association for Computatio...

  74. [82]

    Lawrence Zitnick, and Devi Parikh

    Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. Cider: Consensus-based image description evaluation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 4566–4575, 2015

  75. [83]

    Evaluating the factual consistency of abstractive text summarization

    Wojciech Kry´sci´nski, Bryan McCann, Caiming Xiong, and Richard Socher. Evaluating the factual consistency of abstractive text summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9332–9346, Online, 2020. Assoc...

  76. [84]

    Entity-level factual consistency of abstractive text summarization

    Feng Nan, Ramesh Nallapati, Zhiguo Wang, Cicero Nogueira dos Santos, Henghui Zhu, Dejiao Zhang, Kathleen McKeown, and Bing Xiang. Entity-level factual consistency of abstractive text summarization. In Proceedings of the 16th Conference of the European Chapter of the Associatio...

  77. [85]

    Fact-checking

    Fractalego. Fact-checking. https://huggingface.co/fractalego/fact-checking, 2021

  78. [86]

    Fabbri, Wojciech Kry´sci´nski, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev

    Alexander R. Fabbri, Wojciech Kry´sci´nski, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. Summeval: Re-evaluating summarization evaluation. In Association for Computational Linguistics, 2021

  79. [87]

    Generating representative headlines for news stories

    Xiaotao Gu, Yuning Mao, Jiawei Han, Jialu Liu, Hongkun Yu, You Wu, Cong Yu, Daniel Finnie, Jiaqi Zhai, and Nichol Zukoski. Generating representative headlines for news stories. arXiv preprint arXiv:2001.09386v4, 2020

  80. [88]

    K. M. Hermann, T. Kocisky, E. Grefenstette, L. Espeholt, W. Kay, M. Suleyman, and P. Blunsom. Teaching machines to read and comprehend. In Advances in Neural Information Processing Systems, pages 1693–1701, 2015

  81. [89]

    Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies

    Max Grusky, Mor Naaman, and Yoav Artzi. Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, volume 1 ...

  82. [90]

    A. M. Rush, S. Chopra, and J. Weston. A neural attention model for abstractive sentence summarization. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 379–389, Lisbon, Portugal, 2015. Association for Computational Linguistics

  83. [91]

    Graff, J

    D. Graff, J. Kong, K. Chen, and K. Maeda. English gigaword. Linguistic Data Consortium, 4(1):34, 2003. 22 Survey on Abstractive Text Summarization: Dataset, Models, and Metrics

  84. [92]

    Sharma, C

    E. Sharma, C. Li, and L. Wang. Bigpatent: A large scale dataset for abstractive and coherent summarization. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2204–2213, Florence, Italy, 2019. Association for Computational Linguistics

  85. [93]

    Koupaee and W

    M. Koupaee and W. Y . Wang. Wikihow: A large scale text summarization dataset. arXiv preprint arXiv:1810.09305, 2018

  86. [94]

    B. Kim, H. Kim, and G. Kim. Abstractive summarization of reddit posts with multi-level memory networks. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, volume 1 (Long and Short P...

  87. [95]

    Zhang and J

    R. Zhang and J. Tetreault. This email could save your life: Introducing the task of email subject line generation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 446–456, Florence, Italy, 2019. Association for Computational Li...

  88. [96]

    The enron corpus: A new dataset for email classification research

    Bryan Klimt and Yiming Yang. The enron corpus: A new dataset for email classification research. In Jean- François Boulicaut, Floriana Esposito, Fosca Giannotti, and Dino Pedreschi, editors, Machine Learning: ECML 2004, volume 3201 of Lecture Notes in Computer Science, Berlin, ...

  89. [97]

    Kornilova and V

    A. Kornilova and V . Eidelman. Billsum: A corpus for automatic summarization of us legislation. InProceedings of the 2nd Workshop on New Frontiers in Summarization, pages 48–56, Hong Kong, China, 2019. Association for Computational Linguistics

  90. [98]

    Booksum: A collection of datasets for long-form narrative summarization

    Wojciech Kry´sci´nski, Nazneen Rajani, Divyansh Agarwal, Caiming Xiong, and Dragomir Radev. Booksum: A collection of datasets for long-form narrative summarization. arXiv preprint arXiv:2105.08209v1, 2021

  91. [99]

    Efficient attentions for long document summarization

    Luyang Huang, Shuyang Cao, Nikolaus Parulian, Heng Ji, and Lu Wang. Efficient attentions for long document summarization. In Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hakkani-Tur, Iz Beltagy, Steven Bethard, Ryan Cotterell, Tanmoy Chakraborty, and Yichao Zhou...

  92. [100]

    Multi-lexsum: Real-world summaries of civil rights lawsuits at multiple granularities, 06 2022

    Zejiang Shen, Kyle Lo, Lauren Yu, Nathan Dahlberg, Margo Schlanger, and Doug Downey. Multi-lexsum: Real-world summaries of civil rights lawsuits at multiple granularities, 06 2022

  93. [101]

    Ms2: Multi-document summarization of medical studies

    Jay DeYoung, Iz Beltagy, Madeleine van Zuylen, Bailey Kuehl, and Lucy Lu Wang. Ms2: Multi-document summarization of medical studies. arXiv preprint arXiv:2104.06486v3, 2021

  94. [102]

    Multi-xscience: A large-scale dataset for extreme multi-document summarization of scientific articles

    Yao Lu, Yue Dong, and Laurent Charlin. Multi-xscience: A large-scale dataset for extreme multi-document summarization of scientific articles. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020

  95. [103]

    Fabbri, Irene Li, Tianwei She, Suyi Li, and Dragomir R

    Alexander R. Fabbri, Irene Li, Tianwei She, Suyi Li, and Dragomir R. Radev. Multi-news: a large-scale multi-document summarization dataset and abstractive hierarchical model. arXiv preprint arXiv:1906.01749v3, 2019

  96. [104]

    A large-scale multi-document summarization dataset from the wikipedia current events portal

    Demian Gholipour Ghalandari, Chris Hokamp, Nghia The Pham, John Glover, and Georgiana Ifrim. A large-scale multi-document summarization dataset from the wikipedia current events portal. In Proceedings of the 2020 Annual Meeting of the Association for Computational Linguistics ...

  97. [105]

    Neural network-based abstract generation for opinions and arguments

    Lu Wang and Wang Ling. Neural network-based abstract generation for opinions and arguments. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (HLT-NAACL 2016), pages 47–57, San Dieg...

  98. [106]

    Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer

    Peter J. Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer. Generating wikipedia by summarizing long sequences. In Proceedings of the 6th International Conference on Learning Representations (ICLR 2018), Vancouver, Canada, 2018

  99. [107]

    https://duc.nist.gov/duc2004/, 2004

    D u c 2 0 0 4: Documents, tasks, and measures. https://duc.nist.gov/duc2004/, 2004

  100. [108]

    Summarizing opinions: Aspect extraction meets sentiment prediction and they are both weakly supervised

    Stefanos Angelidis and Mirella Lapata. Summarizing opinions: Aspect extraction meets sentiment prediction and they are both weakly supervised. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2018

  101. [109]

    Nutri-bullets: Summarizing health studies by composing segments

    Darsh J Shah, Lili Yu, Tao Lei, and Regina Barzilay. Nutri-bullets: Summarizing health studies by composing segments. arXiv preprint arXiv:2103.11921v1, 2021

  102. [110]

    Aquamuse: Automatically generating datasets for query-based multi-document summarization

    Sayali Kulkarni, Sheide Chammas, Wan Zhu, Fei Sha, and Eugene Ie. Aquamuse: Automatically generating datasets for query-based multi-document summarization. arXiv preprint arXiv:2010.12694v1, 2020. 23 Survey on Abstractive Text Summarization: Dataset, Models, and Metrics

  103. [111]

    Gamewikisum: A novel large multi-document summarization dataset

    Diego Antognini and Boi Faltings. Gamewikisum: A novel large multi-document summarization dataset. In Proceedings of the 2020 Language Resources and Evaluation Conference (LREC), 2020

  104. [112]

    Large scale abstractive multi-review summarization (lsars) via aspect alignment

    Haojie Pan, Rongqin Yang, Xin Zhou, Rui Wang, Deng Cai, and Xiaozhong Liu. Large scale abstractive multi-review summarization (lsars) via aspect alignment. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, pages...

  105. [113]

    Opinosis: A graph based approach to abstractive summa- rization of highly redundant opinions

    Kavita Ganesan, ChengXiang Zhai, and Jiawei Han. Opinosis: A graph based approach to abstractive summa- rization of highly redundant opinions. In Proceedings of the 23rd International Conference on Computational Linguistics (COLING 2010), pages 340–348, Beijing, China, 2010

  106. [114]

    Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mari...

  107. [115]

    GitHub. Github. https://github.com/, 2024

  108. [116]

    Hallucinations in neural machine translation

    Katherine Lee, Orhan Firat, Ashish Agarwal, Clara Fannjiang, and David Sussillo. Hallucinations in neural machine translation. In NeurIPS 2018 Workshop on Interpretability and Robustness for Audio, Speech, and Language, 2018

  109. [117]

    On faithfulness and factuality in abstractive summarization

    Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1906–1919, Online, 2020

  110. [118]

    Object hallucination in image captioning

    Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. Object hallucination in image captioning. In Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii, editors, Proceedings of the 2018 Conference on Empirical Methods in Natural Langu...

  111. [119]

    Reducing quantity hallucinations in abstractive summarization

    Zheng Zhao, Shay B Cohen, and Bonnie Webber. Reducing quantity hallucinations in abstractive summarization. arXiv preprint arXiv:2009.13312, 2020. 24

  112. [2018]

    Association for Computational Linguistics

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.