REVIEW 2 major objections 5 minor 3 cited by
Neural Text Summarization: A Critical Evaluation
T0 review · 2 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read This paper argues that the stagnation in text summarization is built into the research setup: noisy datasets, weak automatic metrics, and models that exploit layout bias.
desk verdict The canonical critique of summarization benchmarks, with valuable new measurements; Table 5's correlation stats need error bars, but the qualitative findings hold up. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object under scrutiny is the research setup: automatically scraped corpora (CNN/DailyMail and Newsroom), ROUGE n-gram overlap as the evaluation protocol, and the neural models trained and judged on those. The argument is carried by three measurement instruments: crowd-sourced human studies on 100 randomly sampled CNN/DailyMail articles with five annotators each, which quantify content-selection disagreement and the front-loaded 'inverted pyramid' distribution of important sentences; correlation analysis between ROUGE variants and four human rating dimensions (relevance, consistency, fluency, coherence) across 13 published models; and a re-scoring experiment in which the Lead-3 baseline (the first three sentences) replaces the reference summary, exposing how much of model output is explained by layout bias. These instruments turn the critique into numbers rather than impressions.
What would settle it
A concrete falsifier: run the same ROUGE-human correlation protocol on a larger, multi-corpus sample (e.g., 1,000 articles spanning Newsroom, XSum, and non-news domains). If ROUGE Kendall correlations with human relevance ratings rise above about 0.6 for abstractive models, or if annotator agreement on important sentences reaches high consensus without question constraints, the paper's central claims weaken. Similarly, re-scoring a model that demonstrably does not use position (e.g., tested on shuffled paragraphs) and finding high ROUGE against references would contradict the layout-bias explanation.
Extended reading notes
Core claim
On the paper's own terms, the discovery is a three-pronged negative result about how the field measures and builds summarizers. First, human annotators given the same news article disagree sharply about which sentences matter: at a consensus threshold of three out of five votes, they agree on roughly 0.6 important sentences per article unconstrained, rising to about 1.4 when the task is constrained by questions. Second, ROUGE scores computed against one, five, or ten references correlate weakly (Kendall rank around 0.3 for abstractive models) with human ratings of relevance, consistency, fluency, and coherence, and the correlation grows weaker when more reference summaries are used. Third, re-scoring model outputs against the Lead-3 baseline as reference shows a steep rise in ROUGE, indicating that models copy the front-heavy layout of news articles; and ROUGE-1 overlap among different models is far higher than between models and reference summaries, suggesting the outputs converge on easy surface patterns. The paper also reports that 30% of 200 manually checked article-summary pairs from abstractive models contain factual inconsistencies, a dimension current evaluation ignores.
Load-bearing premise
The load-bearing premise is that the human studies, conducted on 100 randomly sampled CNN/DailyMail articles with five annotators each, are representative enough to generalize to all summarization datasets and models.
Editorial extensions
If this is right
- Benchmark comparisons on CNN/DailyMail that rely on ROUGE should be read as measuring partial overlap with front-loaded news prose, not summarization quality.
- New datasets need to be constrained (e.g., question-targeted) and manually curated; otherwise the task is underdetermined and models cannot be expected to learn genuine content selection.
- Evaluation protocols should include factual consistency as a separate dimension, since 30% of inspected abstractive summaries contained factual errors.
- ROUGE-based leaderboards should be supplemented or replaced by human evaluation that correlates better with actual quality.
- The layout bias of news corpora should be part of ablation studies; heuristics like Lead-3 should not be silently baked into preprocessing.
Reading between the lines
- A natural testable extension is to build constrained datasets (question-answer guided) at scale; the paper's constrained-vs-unconstrained comparison predicts such datasets will show clearer model differentiation.
- The paper implies that non-news domains (books, legal documents, conversations) which lack the inverted pyramid will expose models' reliance on position rather than content; benchmarking there could act as a stress test.
- The weak correlation with human judgment suggests ROUGE may reward summarizers that merely echo reference vocabulary; a metric that accepts paraphrase or checks facts would reorder today's leaderboard.
- If the factual-error rate of 30% generalizes, abstractive models deployed in high-stakes settings would likely require automatic fact-checking as a guardrail.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper critically evaluates the standard research setup in neural text summarization by examining three components: automatically collected datasets, ROUGE-based evaluation, and model behavior. It argues that current datasets leave the task underconstrained and contain noise, that ROUGE correlates only weakly with human judgments and ignores factual correctness, and that models overfit to the layout bias of news corpora and produce homogeneous outputs. The evidence includes new human annotation studies on 100 CNN/DailyMail articles, a human evaluation of 13 model outputs across four quality dimensions, manual inspection for factual errors in 200 article-summary pairs, regex-based noise detection in CNN/DM and Newsroom, and ROUGE comparisons against the Lead-3 baseline. The authors conclude that benchmark progress on CNN/DailyMail likely overstates real summarization ability and call for a more robust research setup.
Significance. If its claims hold, this paper is an important and timely critique that could redirect how the summarization community builds datasets, trains models, and reports progress. Its main strengths are the new human-annotation data, the constrained-versus-unconstrained summarization comparison, the inclusion of outputs from 13 existing systems provided by the original authors, and the detailed appendices documenting the annotation protocols. The paper also usefully extends prior ROUGE-criticism work (e.g., Graham 2015, Liu and Liu 2010) to the current neural benchmark setting. However, the strength of the central claims depends on the statistical support, which is currently incomplete for the key evaluation-metric analysis.
major comments (2)
- [Section 4.1, Table 5, Appendix A.2] The central claim that ROUGE is only weakly correlated with human judgment is underdetermined by the reported evidence. Table 5 gives point estimates only, without confidence intervals, p-values, or any indication of the unit of analysis. The manuscript never states whether the Pearson and Kendall values are computed over pooled summaries (13 models x 100 articles), per-model and then averaged, or over system-level scores. This matters because pooled correlations mix between-system and within-system signal and can be nonzero even when ROUGE cannot rank summaries within a system, while per-system correlations with N=100 are clustered by article and system, so the effective sample size for a claim about the evaluation protocol may be as small as the 13 (or 10 abstractive) systems. For 10 abstractive systems, a Kendall tau of about 0.3 is not obviously distinguishable from zero under an exact permutation test. The paper should report bootstrap confidence intervals or permutation-based p-values that respect the article- and system-level clustering, and should explicitly state the aggregation scheme. The related observation that correlations grow weaker with more reference summaries is also not supported without such intervals, since the differences between the values in the 1-reference, 5-reference, and 10-reference columns appear small relative to the likely sampling noise.
- [Section 3.3] The noise-detection analysis reports specific prevalence rates for flawed summaries in CNN/DM and Newsroom (e.g., 0.47% in CNN/DM training and 3.21% in Newsroom training) but provides no validation of the regular expressions and heuristics used to detect the flaws. The paper gives qualitative examples in Table 3 and states that manual inspection revealed consistent patterns, but it does not report the precision or recall of the heuristic detection against a manually labeled sample, nor an inter-annotator agreement for the manual inspection. Without this, the numeric rates are not reliable evidence for the claim that the datasets contain noise detrimental to training and evaluation. The authors should either validate the heuristics on a sample of manually labeled summaries or explicitly characterize the rates as lower-bound estimates from an unvalidated detector.
minor comments (5)
- [Section 3.1, Table 2] The agreement results in Table 2 are reported as point estimates without variance or confidence intervals; given the small sample of 100 articles, adding standard errors or bootstrap intervals would strengthen the underconstrained-task claim.
- [Section 4.1] The sentence "Correlations were computed between all pairs of Human-, ROUGE-scores, for all systems" is ambiguous and should specify the aggregation level and whether the reported coefficients are summary statistics over articles or over systems.
- [Section 4.2] The factual-consistency review reports that 60 of 200 manually inspected article-summary pairs contain consistency issues, but it does not state how many annotators performed the inspection, how the 200 pairs were selected, or what inter-annotator agreement was; these details should be added for the 30% figure to be interpretable.
- [Section 2.4 and References] The citation of Schulman et al. (2015) for the NP-hardness of global optimization with respect to ROUGE appears to be incorrect: the referenced paper, "Gradient estimation using stochastic computation graphs", is about a different topic and does not discuss ROUGE.
- [Section 5.1, Table 6] The right half of Table 6 measures ROUGE against the Lead-3 output as if it were the reference; this is a reasonable way to quantify overlap with the lead bias, but the text should remind readers that these scores are not directly comparable to the left-half scores and do not by themselves measure summary quality.
Circularity Check
No significant circularity found; the paper's critical claims rest on new human annotation studies, direct model-output comparisons, and external prior work, not on fitted parameters or self-citations that load-bear.
full rationale
This paper is an empirical critique rather than a derivation, so the standard circularity patterns do not apply. Claim 1 (underconstrained task) is supported by a Mechanical Turk agreement study comparing constrained and unconstrained summarization settings (Section 3.1, Appendix A.1), not by any definitional equivalence. Claim 2 (weak ROUGE correlation) is supported by newly collected human ratings on outputs of 13 named systems correlated directly with ROUGE scores (Section 4.1, Table 5, Appendix A.2); the correlation coefficient is the measurement, not a fitted input renamed as a prediction. Claim 3 (layout bias) is supported by human sentence-importance annotations (Figure 1) and by comparing model ROUGE scores against both target references and Lead-3 references (Table 6); the latter is an empirical comparison, and the prior work it invokes (Kedzie et al., 2018) is external and independent. The only self-citation, Kryściński et al. (2018), appears as one of thirteen evaluated models in Table 6 and in related-work mentions; its inclusion does not underwrite any of the paper's conclusions. No uniqueness theorem, ansatz, or renamed result is smuggled in via self-citation. Concerns about missing confidence intervals for Table 5 are statistical-evidence issues, not circularity. The derivation chain is thus self-contained against the paper's own empirical inputs, warranting a score of 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Human judgments on 100 randomly sampled CNN/DailyMail articles are representative of summarization quality across datasets and domains.
- domain assumption The four evaluation dimensions (relevance, consistency, fluency, coherence) comprehensively capture summary quality.
- domain assumption ROUGE scores against Lead-3 as a reference measure the degree of layout-bias reliance.
- domain assumption Standard correlation statistics on 100 articles and 13 models yield meaningful conclusions without significance testing.
Cite this review
Pith. "Pith review of Neural Text Summarization: A Critical Evaluation." pith.science (2026). https://pith.science/paper/XJDSB4RX
@misc{pith2026190808960,
author = {Pith},
title = {Pith review of: Neural Text Summarization: A Critical Evaluation},
year = {2026},
howpublished = {\url{https://pith.science/paper/XJDSB4RX}},
note = {Machine review of arXiv:1908.08960}
}
read the original abstract
Text summarization aims at compressing long documents into a shorter form that conveys the most important parts of the original document. Despite increased interest in the community and notable research effort, progress on benchmark datasets has stagnated. We critically evaluate key ingredients of the current research setup: datasets, evaluation metrics, and models, and highlight three primary shortcomings: 1) automatically collected datasets leave the task underconstrained and may contain noise detrimental to training and evaluation, 2) current evaluation protocol is weakly correlated with human judgment and does not account for important characteristics such as factual correctness, 3) models overfit to layout biases of current datasets and offer limited diversity in their outputs.
Figures
Forward citations
Cited by 3 Pith papers
-
ClassComet: Exploring and Designing AI-generated Danmaku in Educational Videos to Enhance Online Learning
AI-written danmaku, combining content and emotion types, can match human comment quality and significantly boost learner engagement and quiz gains in short educational videos.
-
Evaluate Summarization in Fine-Granularity: Auto Evaluation with LLM
SumAutoEval is an LLM-based entity-level summarization evaluator with four dimensions; its claimed human-correlation advantage is not consistently supported by the experiments.
-
An Evaluation of Large Language Models on Text Summarization Tasks Using Prompt Engineering Techniques
A broad benchmark of six open-weights LLMs shows prompt design and chunking affect summarization quality more than model size alone.
Reference graph
Works this paper leans on
-
[1]
URL: " 'urlintro :=
ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRINGS urlintro eprinturl eprintpr...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015. Neural machine translation by jointly learning to align and translate. In ICLR
2015
-
[4]
Fordyce, Philipp Koehn, Christof Monz, and Josh Schroeder
Chris Callison - Burch, Cameron S. Fordyce, Philipp Koehn, Christof Monz, and Josh Schroeder. 2007. (meta-) evaluation of machine translation. In WMT@ACL, pages 136--158. Association for Computational Linguistics
2007
-
[5]
Chris Callison - Burch, Miles Osborne, and Philipp Koehn. 2006. Re-evaluation the role of bleu in machine translation research. In EACL . The Association for Computer Linguistics
work page 2006
-
[6]
Danqi Chen, Jason Bolton, and Christopher D. Manning. 2016. A thorough examination of the cnn/daily mail reading comprehension task. In ACL (1) . The Association for Computer Linguistics
work page 2016
-
[7]
Yen - Chun Chen and Mohit Bansal. 2018. Fast abstractive summarization with reinforce-selected sentence rewriting. In ACL (1) , pages 675--686. Association for Computational Linguistics
work page 2018
-
[8]
Sumit Chopra, Michael Auli, and Alexander M. Rush. 2016. Abstractive sentence summarization with attentive recurrent neural networks. In NAACL HLT 2016, The 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, San Diego California, USA, June 12-17, 2016
work page 2016
Show all 84 references
-
[9]
Eric Chu and Peter J. Liu. 2018. Unsupervised neural multi-document abstractive summarization. CoRR, abs/1810.05739
2018 arXiv
-
[10]
Arman Cohan, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Seokhwan Kim, Walter Chang, and Nazli Goharian. 2018. A discourse-aware attention model for abstractive summarization of long documents. In Proceedings of the 2018 Conference of the North American Chapter of the Associa...
2018
-
[11]
Carlos A Colmenares, Marina Litvak, Amin Mantrach, and Fabrizio Silvestri. 2015. Heads: Headline generation as sequence prediction using an abstract feature-rich space. In HLT-NAACL, pages 133--142
2015
-
[12]
Yue Dong, Yikang Shen, Eric Crawford, Herke van Hoof, and Jackie Chi Kit Cheung. 2018. Banditsum: Extractive summarization as a contextual bandit. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - Novembe...
2018
-
[13]
Bonnie Dorr, David Zajic, and Richard Schwartz. 2003. Hedge trimmer: A parse-and-trim approach to headline generation. In HLT-NAACL
2003
-
[14]
Katja Filippova and Yasemin Altun. 2013. Overcoming the lack of parallel data in sentence compression. In Proceedings of EMNLP, pages 1481--1491. Citeseer
2013
-
[15]
Kavita Ganesan. 2018. ROUGE 2.0: Updated and improved measures for evaluation of summarization tasks. CoRR, abs/1803.01937
2018 arXiv
-
[16]
Sebastian Gehrmann, Yuntian Deng, and Alexander M. Rush. 2018. Bottom-up abstractive summarization. In EMNLP , pages 4098--4109. Association for Computational Linguistics
2018
-
[17]
Dimitra Gkatzia and Saad Mahamood. 2015. A snapshot of NLG evaluation practices 2005 - 2014. In ENLG 2015 - Proceedings of the 15th European Workshop on Natural Language Generation, 10-11 September 2015, University of Brighton, Brighton, UK , pages 57--60
2015
-
[18]
Max Glockner, Vered Shwartz, and Yoav Goldberg. 2018. Breaking NLI systems with sentences that require simple lexical inferences. In ACL (2) , pages 650--655. Association for Computational Linguistics
2018
-
[19]
Yash Goyal, Tejas Khot, Douglas Summers - Stay, Dhruv Batra, and Devi Parikh. 2017. Making the V in VQA matter: Elevating the role of image understanding in visual question answering. In CVPR , pages 6325--6334. IEEE Computer Society
2017
-
[20]
David Graff and C Cieri. 2003. English gigaword, linguistic data consortium
2003
-
[21]
Yvette Graham. 2015. Re-evaluating automatic summarization with BLEU and 192 shades of ROUGE . In EMNLP , pages 128--137. The Association for Computational Linguistics
2015
-
[22]
Max Grusky, Mor Naaman, and Yoav Artzi. 2018. Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAA...
2018
-
[23]
Han Guo, Ramakanth Pasunuru, and Mohit Bansal. 2018. Soft layer-specific multi-task summarization with entailment and question generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Melbourne, Australia, July 15-20, 20...
2018
-
[24]
Bowman, and Noah A
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel R. Bowman, and Noah A. Smith. 2018. Annotation artifacts in natural language inference data. In NAACL-HLT (2) , pages 107--112. Association for Computational Linguistics
2018
-
[25]
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015. Teaching machines to read and comprehend. In NIPS
2015
-
[26]
Kai Hong and Ani Nenkova. 2014. Improving the estimation of word importance for news multi-document summarization. In Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2014, April 26-30, 2014, Gothenburg, Sweden
2014
-
[27]
Wan Ting Hsu, Chieh - Kai Lin, Ming - Ying Lee, Kerui Min, Jing Tang, and Min Sun. 2018. A unified model for extractive and abstractive summarization using inconsistency loss. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018,...
2018
-
[28]
Yichen Jiang and Mohit Bansal. 2018. Closed-book training to improve summarization encoder memory. In EMNLP , pages 4067--4077. Association for Computational Linguistics
2018
-
[29]
Divyansh Kaushik and Zachary C. Lipton. 2018. How much reading does reading comprehension require? A critical investigation of popular benchmarks. In EMNLP , pages 5010--5015. Association for Computational Linguistics
2018
-
[30]
McKeown, and Hal Daum \' e III
Chris Kedzie, Kathleen R. McKeown, and Hal Daum \' e III. 2018. Content selection in deep learning models of summarization. In EMNLP , pages 1818--1828. Association for Computational Linguistics
2018
-
[31]
Byeongchang Kim, Hyunwoo Kim, and Gunhee Kim. 2018. Abstractive summarization of reddit posts with multi-level memory networks. CoRR, abs/1811.00783
2018 arXiv
-
[32]
Mahnaz Koupaee and William Yang Wang. 2018. Wikihow: A large scale text summarization dataset. CoRR, abs/1810.09305
2018 arXiv
-
[33]
Wojciech Kry \'s ci \'n ski, Romain Paulus, Caiming Xiong, and Richard Socher. 2018. Improving abstraction in text summarization. In EMNLP , pages 1808--1817. Association for Computational Linguistics
2018
-
[34]
Moontae Lee, Xiaodong He, Wen - tau Yih, Jianfeng Gao, Li Deng, and Paul Smolensky. 2016. Reasoning in vector space: An exploratory study of question answering. In ICLR
2016
-
[35]
Junyi Jessy Li, Kapil Thadani, and Amanda Stent. 2016. The role of discourse units in near-extractive summarization. In Proceedings of the SIGDIAL 2016 Conference, The 17th Annual Meeting of the Special Interest Group on Discourse and Dialogue, 13-15 September 2016, Los Angele...
2016
-
[36]
Wei Li, Xinyan Xiao, Yajuan Lyu, and Yuanzhuo Wang. 2018. Improving neural abstractive document summarization with structural regularization. In EMNLP , pages 4078--4087. Association for Computational Linguistics
2018
-
[37]
Chin-Yew Lin. 2004. http://research.microsoft.com/ cyl/download/papers/WAS2004.pdf Rouge: A package for automatic evaluation of summaries . In Proc. ACL workshop on Text Summarization Branches Out, page 10
2004
-
[38]
Lipton and Jacob Steinhardt
Zachary C. Lipton and Jacob Steinhardt. 2019. Troubling trends in machine learning scholarship. ACM Queue , 17(1):80
2019
-
[39]
Feifan Liu and Yang Liu. 2010. Exploring correlation between ROUGE and human evaluation on meeting summaries. IEEE Trans. Audio, Speech & Language Processing , 18(1):187--196
2010
-
[40]
Jingyun Liu, Jackie Chi Kit Cheung, and Annie Louis. 2019. What comes next? extractive summarization by next-sentence prediction. CoRR, abs/1901.03859
2019 arXiv
-
[41]
Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer
Peter J. Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer. 2018. Generating wikipedia by summarizing long sequences. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 20...
2018
-
[42]
Ramesh Nallapati, Feifei Zhai, and Bowen Zhou. 2017. Summarunner: A recurrent neural network based sequence model for extractive summarization of documents. In AAAI
2017
-
[43]
Ramesh Nallapati, Bowen Zhou, C a g lar G \"u l c ehre, Bing Xiang, et al. 2016 a . Abstractive text summarization using sequence-to-sequence rnns and beyond. Proceedings of SIGNLL Conference on Computational Natural Language Learning
2016
-
[44]
Ramesh Nallapati, Bowen Zhou, and Mingbo Ma. 2016 b . Classify or select: Neural architectures for extractive document summarization. CoRR, abs/1611.04244
2016 arXiv
-
[45]
Cohen, and Mirella Lapata
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018 a . Don't give me the details, just the summary! T opic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium
2018
-
[46]
Cohen, and Mirella Lapata
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018 b . Ranking sentences for extractive summarization with reinforcement learning. In NAACL-HLT , pages 1747--1759. Association for Computational Linguistics
2018
-
[47]
Shashi Narayan, Nikos Papasarantopoulos, Mirella Lapata, and Shay B. Cohen. 2017. Neural extractive summarization with side information. CoRR, abs/1704.04530
2017 arXiv
-
[48]
Passonneau
Ani Nenkova and Rebecca J. Passonneau. 2004. Evaluating content selection in summarization: The pyramid method. In Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics, HLT-NAACL 2004, Boston, Massachusetts, USA, M...
2004
-
[49]
Joel Larocca Neto, Alex A Freitas, and Celso AA Kaestner. 2002. Automatic text summarization using a machine learning approach. In Brazilian Symposium on Artificial Intelligence, pages 205--215. Springer
2002
-
[50]
Jun - Ping Ng and Viktoria Abrecht. 2015. Better summarization evaluation with word embeddings for ROUGE . CoRR, abs/1508.06034
2015 arXiv
-
[51]
Benjamin Nye and Ani Nenkova. 2015. Identification and characterization of newsworthy verbs in world news. In NAACL HLT 2015, The 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Denver, Colorado, USA,...
2015
-
[52]
Paul Over and James Yen. 2001. An introduction to duc-2001: Intrinsic evaluation of generic news text summarization systems
2001
-
[53]
Paul Over and James Yen. 2002. An introduction to duc-2002: Intrinsic evaluation of generic news text summarization systems
2002
-
[54]
Paul Over and James Yen. 2003. An introduction to duc-2003: Intrinsic evaluation of generic news text summarization systems
2003
-
[55]
Rankel, Hoa Trang Dang, and John M
Karolina Owczarzak, Peter A. Rankel, Hoa Trang Dang, and John M. Conroy. 2012. Assessing the effect of inconsistent assessors on summarization evaluation. In ACL (2) , pages 359--362. The Association for Computer Linguistics
2012
-
[56]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei - Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In ACL , pages 311--318. ACL
2002
-
[57]
Passonneau, Emily Chen, Weiwei Guo, and Dolores Perin
Rebecca J. Passonneau, Emily Chen, Weiwei Guo, and Dolores Perin. 2013. Automated pyramid scoring of summaries using distributional semantics. In ACL (2) , pages 143--147. The Association for Computer Linguistics
2013
-
[58]
Ramakanth Pasunuru and Mohit Bansal. 2018. http://arxiv.org/abs/1804.06451 Multi-reward reinforced summarization with saliency and entailment . CoRR, abs/1804.06451
2018 arXiv
-
[59]
Romain Paulus, Caiming Xiong, and Richard Socher. 2017. A deep reinforced model for abstractive summarization. In ICLR
2017
-
[60]
Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018. Hypothesis only baselines in natural language inference. In *SEM@NAACL-HLT, pages 180--191. Association for Computational Linguistics
2018
-
[61]
Matt Post. 2018. A call for clarity in reporting BLEU scores. In WMT , pages 186--191. Association for Computational Linguistics
2018
-
[62]
PurdueOWL. 2019. Journalism and journalistic writing: The inverted pyramid structure. Accessed: 2019-05-15
2019
-
[63]
Ehud Reiter. 2018. A structured review of the validity of BLEU . Computational Linguistics, 44(3)
2018
-
[64]
Ehud Reiter and Anja Belz. 2009. An investigation into the validity of some metrics for automatically evaluating natural language generation systems. Computational Linguistics, 35(4):529--558
2009
-
[65]
Alexander M Rush, Sumit Chopra, and Jason Weston. 2015. A neural attention model for abstractive sentence summarization. Proceedings of EMNLP
2015
-
[66]
Evan Sandhaus. 2008. The new york times annotated corpus. Linguistic Data Consortium, Philadelphia, 6(12):e26752
2008
-
[67]
John Schulman, Nicolas Heess, Theophane Weber, and Pieter Abbeel. 2015. Gradient estimation using stochastic computation graphs. In NIPS
2015
-
[68]
Raphael Schumann. 2018. Unsupervised abstractive sentence summarization using length controlled variational autoencoder. CoRR, abs/1809.05233
2018 arXiv
-
[69]
Sculley, Jasper Snoek, Alexander B
D. Sculley, Jasper Snoek, Alexander B. Wiltschko, and Ali Rahimi. 2018. Winner's curse? on pace, progress, and empirical rigor. In ICLR (Workshop) . OpenReview.net
2018
-
[70]
Liu, and Christopher D
Abigail See, Peter J. Liu, and Christopher D. Manning. 2017. Get to the point: Summarization with pointer-generator networks. In ACL
2017
-
[71]
Wong, and Fang Chen
Elaheh ShafieiBavani, Mohammad Ebrahimi, Raymond K. Wong, and Fang Chen. 2018. A graph-theoretic summary evaluation for rouge. In EMNLP , pages 762--767. Association for Computational Linguistics
2018
-
[72]
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014. Sequence to sequence learning with neural networks. In NIPS
2014
-
[73]
Sho Takase, Jun Suzuki, Naoaki Okazaki, Tsutomu Hirao, and Masaaki Nagata. 2016. Neural headline generation on abstract meaning representation. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, EMNLP 2016, Austin, Texas, USA, November 1-4, 2016
2016
-
[74]
Jiwei Tan, Xiaojun Wan, and Jianguo Xiao. 2017. Abstractive document summarization with a graph-based attentional neural model. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Vancouver, Canada, July 30 - August 4, Volume 1...
2017
-
[75]
Liling Tan, Jon Dehdari, and Josef van Genabith. 2015. An awkward disparity between BLEU / RIBES scores and human judgements in machine translation. In WAT , pages 74--81. Workshop on Asian Translation
2015
-
[76]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 20...
2017
-
[77]
Oriol Vinyals, Meire Fortunato, and Navdeep Jaitly. 2015. Pointer networks. In NIPS
2015
-
[78]
Yuxiang Wu and Baotian Hu. 2018. Learning to extract coherent summary via deep reinforcement learning. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI-18), the 30th innovative Applications of Artificial Intelligence (IAAI-18), and the 8th ...
2018
-
[79]
Yongqin Xian, Bernt Schiele, and Zeynep Akata. 2017. Zero-shot learning - the good, the bad and the ugly. In CVPR , pages 3077--3086. IEEE Computer Society
2017
-
[80]
Jiacheng Xu and Greg Durrett. 2019. Neural extractive text summarization with syntactic compression. CoRR, abs/1902.00863
2019 arXiv
-
[81]
Yinfei Yang and Ani Nenkova. 2014. Detecting information-dense texts in multiple news domains. In Proceedings of the Twenty-Eighth AAAI Conference on Artificial Intelligence, July 27 -31, 2014, Qu \' e bec City, Qu \' e bec, Canada
2014
-
[82]
Fangfang Zhang, Jin - ge Yao, and Rui Yan. 2018. On the abstractiveness of neural document summarization. In EMNLP , pages 785--790. Association for Computational Linguistics
2018
-
[83]
Liang Zhou, Chin - Yew Lin, Dragos Stefan Munteanu, and Eduard H. Hovy. 2006. Paraeval: Using paraphrases to evaluate summaries automatically. In Human Language Technology Conference of the North American Chapter of the Association of Computational Linguistics, Proceedings, Ju...
2006
-
[84]
Qingyu Zhou, Nan Yang, Furu Wei, Shaohan Huang, Ming Zhou, and Tiejun Zhao. 2018. Neural document summarization by jointly learning to score and select sentences. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Melbourne, A...
2018
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.