REVIEW 4 major objections 5 minor 3 cited by
Evaluating the Evaluators: Are readability metrics good measures of readability?
T0 review · 4 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read Traditional readability formulas mostly fail to match human judgments of plain-language summaries, while language models agree much better.
desk verdict The negative finding on FKGL is real but the paper overstates the LM advantage and needs confidence intervals before it should drive practice. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is a single gold-standard benchmark: 60 summaries of 10 scientific papers rated for reading ease by crowd workers on a 1-5 scale, averaged per summary (August et al. 2024). The paper measures each candidate evaluator by its Pearson and Kendall-Tau correlation with those averages: 8 traditional formulas (FKGL, FRE, DCRS, ARI, CLI, GFI, Spache, Linsear Write) versus 5 LMs prompted to use their own judgment and explain their score. The LM reasoning text is then mined with YAKE keyword extraction to show the models appeal to explanation of terms and required background knowledge. A second mechanism is the dataset-level comparison: mean scores across 10 summarization datas
What would settle it
Gather a fresh human-evaluated readability corpus for plain-language summaries covering, say, 30+ papers in several domains, and recompute the correlations. The paper's claim would be falsified if FKGL's Pearson correlation reaches or exceeds the best LM's, or if the best LM's correlation drops below 0.3; a finding that PLOS and CELLS score high on the new corpus would also undercut the dataset conclusions.
Extended reading notes
Core claim
The paper's central claim is that readability, for plain language summaries, is not the property that traditional readability formulas measure. Formulas built on syllable counts, sentence lengths, and word lists label a summary of acute respiratory distress syndrome as college-level even when human readers find it clear, because they cannot see that 'a very serious lung disease' defines the term. On the only available human-judgment dataset for PLS (60 summaries of 10 scientific papers, rated 1-5), FKGL correlates 0.16 (Pearson), while all five tested LMs correlate above 0.45, led by Llama 3.3 70B at 0.56; the best traditional metrics, DCRS and CLI, reach only about 0.37. The paper extends t
Load-bearing premise
The comparison stands entirely on one gold standard: 60 human-readability ratings of summaries of 10 English science papers, collected for a different study; if those ratings are noisy or unrepresentative of plain-language readability, every correlation and dataset ranking inherits the error.
Editorial extensions
If this is right
- Papers and shared tasks that use FKGL as the primary readability check for plain-language summaries should stop: its correlation with human readability is near chance (0.16).
- DCRS and CLI are the strongest of the traditional formulas but still modest; the paper recommends pairing them with LM evaluators rather than relying on either alone.
- Off-the-shelf LMs of 7B-70B are workable readability judges; their written reasons can be inspected to see whether they reward definitions of technical terms and penalize missing background.
- The PLOS and CELLS datasets, widely used in PLS shared tasks, are better regarded as general scientific summarization data, not plain-language data; CDSR and SciNews fit the PLS label better.
- Readability metrics for PLS should be redesigned around the human notion—explanations, context, background—rather than lexical complexity.
Reading between the lines
- If the gold standard is representative, the same blind spot likely affects readability claims in health communication, legal documents, and other accessibility work that leans on grade-level formulas, because those formulas systematically reward short acronyms and penalize long explanatory words.
- The 0.56 ceiling leaves room: prompt tuning, calibrated scales, or fine-tuned judge models may push LM-human agreement higher, and the near-tie among small and large LMs suggests capability is not the main constraint.
- Because 60 summaries are nested in only 10 source papers, the dataset-level rankings (especially PLOS and CELLS) should be read as preliminary; a broader corpus could shift those means.
- If LM judges replace formulas, their known biases and inconsistency need auditing before they become the new default; the paper itself flags this.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper surveys evaluation practice in Plain Language Summarization (PLS), finding that FKGL is the most commonly used readability metric in ACL PLS papers. It then compares eight traditional readability metrics and five LLM-based judges against human readability judgments from the August et al. (2024) dataset of 60 summaries over 10 scientific papers. The central empirical claims are that six of eight traditional metrics correlate below 0.3 with human judgments (FKGL at Pearson r=0.16), that LLM evaluators correlate higher (best Llama 3.3 70B at r=0.56), and that applying the best LLM judge to ten summarization datasets yields different dataset-level conclusions than FKGL, including the recommendation that PLOS and CELLS are not plain-language datasets. The authors release analysis code and survey data, and make concrete recommendations for future PLS evaluation.
Significance. If the empirical claims hold, the paper makes a useful and timely contribution: it is the first direct comparison of standard readability formulas with human readability judgments in the PLS setting, and it offers actionable guidance to a community that currently relies on FKGL. The central comparison is not circular: the LLM scores are zero-shot judgments on a held-out human-judgment set, not fitted to those labels, and the traditional metric scores come from an independent package. The survey of ACL PLS evaluation practices and the release of code and data are concrete strengths. However, the empirical foundation is narrow—a single 60-summary, 10-source-paper human dataset—and the paper's confidence intervals and significance tests are either absent or undercut the headline contrasts. The stress-test concern about circularity does not land, but the concern about statistical support does.
major comments (4)
- [§3.2, Tables 2a/2b, Appendix A] All headline correlations rest on n=60 summaries nested in 10 source papers, yet the paper reports no confidence intervals and treats the summaries as independent. At n=60, FKGL's r=0.16 is not significantly different from zero (approximate p≈0.22), and clustering by paper would widen the uncertainty further. This uncertainty propagates into every comparison in Tables 2a/2b and into the §4.4 dataset conclusions. The Limitations section acknowledges domain and language limits, but not this sampling limitation. The authors should report cluster-robust confidence intervals, mixed-effects or leave-one-paper-out analyses, and adjust the strength of the claims accordingly.
- [Appendix C, Table 9] The Williams tests reported by the authors themselves undercut the abstract's 'LMs are better judges' claim. Llama 3.3 70B is not significantly better than DCRS (p=0.14) or CLI (p=0.13), and Llama 3.1 8B is not significantly better than FKGL (p=0.06). The only robust pairwise improvements are against FKGL for the stronger models. The paper should present these p-values in the main text and qualify the claim, or add enough data/evidence to support the stronger reading.
- [§4.4, Tables 4 and 5] The recommendation that PLOS and CELLS 'may not be well-suited for PLS' and the Cohen's Kappa=0.17 disagreement analysis implicitly treat the Llama 3.3 70B scores as ground truth on datasets for which no human readability judgments are available. The expert/kid sanity checks are useful, but they do not establish that the LM's 1–5 scale is calibrated across these new datasets. Furthermore, the binary thresholds (LM score ≥3 vs FKGL score <12) are chosen without justification and directly determine Kappa and the dataset-level conclusions. The authors should report sensitivity to threshold choices or validate on a human-rated sample from these datasets.
- [Appendix A and §3.2] The human gold standard is the only available PLS human judgment dataset, but it is also narrow: 10 source papers sampled from r/science, 6 summaries per paper (2 expert, 4 GPT-3 generated), and binarized Cohen's Kappa of 0.6. The paper's general conclusion that traditional metrics are poor measures of readability for PLS assumes this set is representative of PLS outputs more broadly. I would like to see a sensitivity analysis that removes one source paper at a time, and a discussion of how the expert/GPT-3 composition and the r/science sampling may affect the correlations. This is not fatally circular, but it is a load-bearing limitation.
minor comments (5)
- [Table 9 and Table 2b] The row header 'Llama 3.1 70B' in Table 9 conflicts with 'Llama 3.3 70B' used in Table 2b and throughout the text.
- [Table 4 and §4.4] The dataset is referred to as 'SKJ' in Table 4 but as 'SJK' (Science Journal for Kids) elsewhere.
- [Table 3] The caption appears to mislabel its own panels: '3b contains an example summary' should likely refer to panel (a), and the description of panel (b) is duplicated.
- [§3.2] Typo: 'Flesh-Kincaid' should be 'Flesch-Kincaid'.
- [References] The Gunning Fog Index is cited to Isnaeni (2017), but the original source is Gunning (1952), which is listed in the bibliography.
Circularity Check
Core comparison is an external benchmark, but the headline LM correlation is selected from prompts tuned on the same 60 human-scored summaries, giving a mild fitted-input bias rather than construction-level circularity.
-
fitted input called prediction
[Section 3.3, Appendix B (Table 8 vs Table 2b)]
"We experiment with 3 prompts and report the prompts in appendix B. We report the Pearson and Kendall-Tau correlations of the scores provided by each LM with the human judgments. ... The Own Reasoning Prompt performs the best when averaged across all models."
The paper's headline result, Llama 3.3 70B achieving Pearson r=0.56 with human judgments, is the correlation of the specific prompt that was selected as best on the same 60 human-scored summaries used to compute that correlation. The prompt choice is thus fitted to the gold-standard labels, and the reported value is a selected maximum over three prompt variants (and implicitly over five models) rather than an unbiased out-of-sample estimate. This is a mild form of fitted-input-called-prediction: the 'best-performing model' number is not independent of the human labels it is compared against. It does not make the correlation tautological or reduce it by construction, because the LM scores themselves are not trained on the human labels, so the circularity is minor rather than central.
full rationale
The central comparison in the paper is an external benchmark. Traditional readability scores come from a standard package (py-readability-metrics), and the LM scores are generated by prompting models that were not trained or fine-tuned on the 60 human judgments from August et al. (2024). The human gold standard is an external dataset, not a self-citation, and no uniqueness theorem is imported from the authors' prior work. The survey of PLS evaluation practices (RQ1) is a literature count, not a derivation. The dataset-level conclusions in Section 4.4 are extrapolations of the best LM's judgments to other corpora; they are not circular, though they inherit the statistical limitations of the small gold standard. The only mild circularity is the selection of the best prompt on the same 60-summary test set before reporting the headline correlation. Because the paper discloses all prompt results in Appendix B and does not hide the selection, this is a transparency/selection-bias issue rather than a construction-level equivalence. No equation in the paper is defined in terms of the target result, and no fitted parameter is renamed as a prediction. Therefore the overall circularity is low.
Assumptions & free parameters
free parameters (2)
- FKGL high-readability threshold =
12 (grade level)
- LM high-readability threshold =
3 on a 1 to 5 scale
assumptions (5)
- domain assumption Human judgments collected by August et al. (2024), averaged per summary, are a valid gold standard for PLS readability.
- domain assumption The 60 summaries can be treated as 60 independent observations for correlation and significance testing.
- standard math Pearson correlation between averaged human 1 to 5 scores and metric scores is an appropriate measure of evaluation quality.
- domain assumption The ACL Anthology query (55 hits, 18 PLS-relevant papers) fairly represents PLS evaluation practice.
- domain assumption LM scores from the chosen prompt and decoding are stable enough to rank datasets.
Cite this review
Pith. "Pith review of Evaluating the Evaluators: Are readability metrics good measures of readability?." pith.science (2026). https://pith.science/paper/2VSVJRBM
@misc{pith2026250819221,
author = {Pith},
title = {Pith review of: Evaluating the Evaluators: Are readability metrics good measures of readability?},
year = {2026},
howpublished = {\url{https://pith.science/paper/2VSVJRBM}},
note = {Machine review of arXiv:2508.19221}
}
read the original abstract
Plain Language Summarization (PLS) aims to distill complex documents into accessible summaries for non-expert audiences. In this paper, we conduct a thorough survey of PLS literature, and identify that the current standard practice for readability evaluation is to use traditional readability metrics, such as Flesch-Kincaid Grade Level (FKGL). However, despite proven utility in other fields, these metrics have not been compared to human readability judgments in PLS. We evaluate 8 readability metrics and show that most correlate poorly with human judgments, including the most popular metric, FKGL. We then show that Language Models (LMs) are better judges of readability, with the best-performing model achieving a Pearson correlation of 0.56 with human judgments. Extending our analysis to PLS datasets, which contain summaries aimed at non-expert audiences, we find that LMs better capture deeper measures of readability, such as required background knowledge, and lead to different conclusions than the traditional metrics. Based on these findings, we offer recommendations for best practices in the evaluation of plain language summaries. We release our analysis code and survey data.
Figures
Forward citations
Cited by 3 Pith papers
-
Disco-RAG: Discourse-Aware Retrieval-Augmented Generation
Disco-RAG improves RAG by building intra-chunk discourse trees and inter-chunk rhetorical graphs that feed into a planning blueprint, delivering state-of-the-art results on question answering and long-document summari...
-
OmniPresent: Generating Coherent Presentation Suites from Scientific Papers
A multi-agent HTML pipeline with shared knowledge and cross-artifact verify-and-repair generates coherent poster/slides/video/page suites from papers and beats specialized baselines on OmniPreBench.
-
Machine Learning Research Has Outpaced Its Communication Norms and NeurIPS Should Act
NeurIPS papers have grown much harder to read since 1987, with acronym use tripling and sensational language rising, and the authors propose seven concrete writing standards for the conference to pilot.
Reference graph
Works this paper leans on
-
[1]
Tal August, Kyle Lo, Noah A. Smith, and Katharina Reinecke. 2024. https://api.semanticscholar.org/CorpusID:268297229 Know your audience: The benefits and pitfalls of generating plain language summaries beyond the "general" audience . Proceedings of the CHI Conference on Human Factors in Computing Systems
work page 2024
-
[2]
Tal August, Katharina Reinecke, and Noah A. Smith. 2022. https://api.semanticscholar.org/CorpusID:248780294 Generating scientific definitions with controllable complexity . In Annual Meeting of the Association for Computational Linguistics
work page 2022
-
[3]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gr...
work page 2020
-
[4]
Isabel Cachola, Kyle Lo, Arman Cohan, and Daniel Weld. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.428 TLDR : Extreme summarization of scientific documents . In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4766--4777, Online. Association for Computational Linguistics
-
[5]
Ricardo Campos, Vítor Mangaravite, Arian Pasquali, Alípio Jorge, Célia Nunes, and Adam Jatowt. 2020. https://doi.org/10.1016/j.ins.2019.09.013 Yake! keyword extraction from single documents using multiple local features . Information Sciences, 509:257--289
-
[6]
Afonso Cavaco Carla Pires and Marina Vigário. 2017. https://doi.org/10.1080/09296174.2017.1311448 Towards the definition of linguistic metrics for evaluating text readability . Journal of Quantitative Linguistics, 24(4):319--349
arXiv 2017
-
[7]
Muthu Kumar Chandrasekaran, Guy Feigenblat, Eduard Hovy, Abhilasha Ravichander, Michal Shmueli-Scheuer, and Anita de Waard. 2020. https://doi.org/10.18653/v1/2020.sdp-1.24 Overview and insights from the shared tasks at scholarly document processing 2020: CL - S ci S umm, L ay S umm and L ong S umm . In Proceedings of the First Workshop on Scholarly Docume...
-
[8]
Arman Cohan, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Seokhwan Kim, W. Chang, and Nazli Goharian. 2018. https://api.semanticscholar.org/CorpusID:4894594 A discourse-aware attention model for abstractive summarization of long documents . In North American Chapter of the Association for Computational Linguistics
work page 2018
Show all 68 references
-
[9]
Meri Coleman and Ta Lin Liau. 1975. https://api.semanticscholar.org/CorpusID:144250124 A computer readability formula designed for machine scoring. Journal of Applied Psychology, 60:283--284
1975
-
[10]
Scott Andrew Crossley, Aron Heintz, Joon Suh Choi, Jordan Batchelor, Mehrnoush Karimi, and Agnes Malatinszky. 2021. https://api.semanticscholar.org/CorpusID:247321796 The commonlit ease of readability (clear) corpus . In Educational Data Mining
2021
-
[11]
Edgar Dale and Jeanne S Chall. 1948. A formula for predicting readability: Instructions. Educational research bulletin, pages 37--54
1948
-
[12]
William H DuBay. 2004. The principles of readability. Impact Information
2004
-
[13]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zha...
2024 arXiv
-
[14]
A. R. Fabbri, Wojciech Kryscinski, Bryan McCann, Richard Socher, and Dragomir R. Radev. 2020. https://api.semanticscholar.org/CorpusID:220768873 Summeval: Re-evaluating summarization evaluation . Transactions of the Association for Computational Linguistics, 9:391--409
2020
-
[15]
simplification of flesch reading ease formula
Rudolf Flesch. 1952. "simplification of flesch reading ease formula". Journal of Applied Psychology
1952
-
[16]
Rudolf Franz Flesch. 1943. https://api.semanticscholar.org/CorpusID:191201423 Marks of readable style: a study in adult education . In Teachers College Contributions to Education
1943
-
[17]
Lorenzo Jaime Flores, Heyuan Huang, Kejian Shi, Sophie Chheang, and Arman Cohan. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.322 Medical text simplification: Optimizing for readability with unlikelihood training and reranked beam search decoding . In Findings of the ...
2023 doi
-
[18]
Tomas Goldsack, Zheheng Luo, Qianqian Xie, Carolina Scarton, Matthew Shardlow, Sophia Ananiadou, and Chenghua Lin. 2023. Overview of the biolaysumm 2023 shared task on lay summarization of biomedical research articles. In Proceedings of the 22st Workshop on Biomedical Language...
2023
-
[19]
Tomas Goldsack, Carolina Scarton, Matthew Shardlow, and Chenghua Lin. 2024. Overview of the biolaysumm 2024 shared task on the lay summarization of biomedical research articles. In The 23rd Workshop on Biomedical Natural Language Processing and BioNLP Shared Tasks, Bangkok, Th...
2024
-
[20]
Tomas Goldsack, Zhihao Zhang, Chenghua Lin, and Carolina Scarton. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.724 Making science simple: Corpora for the lay summarisation of scientific literature . In Proceedings of the 2022 Conference on Empirical Methods in Natural Lan...
2022 doi
-
[21]
Tanya Goyal, Junyi Jessy Li, and Greg Durrett. 2022. https://api.semanticscholar.org/CorpusID:252532176 News summarization and evaluation in the era of GPT-3 . ArXiv, abs/2209.12356
2022 arXiv
-
[22]
Yvette Graham and Timothy Baldwin. 2014. https://doi.org/10.3115/v1/D14-1020 Testing for significance of increased correlation with human judgment . In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing ( EMNLP ) , pages 172--176, Doha, Qata...
2014 doi
-
[23]
Robert Gunning. 1952. https://books.google.com/books?id=ofI0AAAAMAAJ The Technique of Clear Writing . McGraw-Hill
1952
-
[24]
Cohen, and Lucy Lu Wang
Yue Guo, Tal August, Gondy Leroy, Trevor A. Cohen, and Lucy Lu Wang. 2023. https://api.semanticscholar.org/CorpusID:258841161 APPLS : Evaluating evaluation metrics for plain language summarization . In Conference on Empirical Methods in Natural Language Processing
2023
-
[25]
Yue Guo, Wei Qiu, Gondy Leroy, Sheng Wang, and Trevor A. Cohen. 2022. https://api.semanticscholar.org/CorpusID:253397732 Retrieval augmentation of large language models for lay language generation . Journal of biomedical informatics, page 104580
2022
-
[26]
Yue Guo, Wei Qiu, Yizhong Wang, and Trevor Cohen. 2021. Automated lay language summarization of biomedical scientific reviews. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 160--168
2021
-
[27]
Yu Han, Aaron Ceross, and Jeroen Bergmann. 2024. https://api.semanticscholar.org/CorpusID:274023239 The use of readability metrics in legal text: A systematic literature review . ArXiv, abs/2411.09497
2024 arXiv
-
[28]
Nur Rachma Isnaeni. 2017. https://journal3.uin-alauddin.ac.id/index.php/elite/article/view/3356 Readability of english written materials . Elite : English and Literature Journal, 1(1):179--191
2017
-
[29]
Yuelyu Ji, Zhuochun Li, Rui Meng, Sonish Sivarajkumar, Yanshan Wang, Zeshui Yu, Hui Ji, Yushui Han, Hanyu Zeng, and Daqing He. 2024. https://doi.org/10.18653/v1/2024.bionlp-1.75 RAG - RLRC - L ay S um at B io L ay S umm: Integrating retrieval-augmented generation and readabili...
2024 doi
-
[30]
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024. Mixtral of experts. arXiv preprint arXiv:2401.04088
2024 arXiv
-
[31]
Albert Qiaochu Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L'elio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, T...
2023 arXiv
-
[32]
Daniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie, Jungo Kasai, Yejin Choi, Noah A Smith, and Daniel S Weld. 2022. https://arxiv.org/abs/2101.06561 GENIE: Toward Reproducible and Standardized Human Evaluation for Text Generation . In Conference on Empirical M...
2022 arXiv
-
[33]
Qintong Li, Leyang Cui, Lingpeng Kong, and Wei Bi. 2025. https://aclanthology.org/2025.coling-main.688/ Exploring the reliability of large language models as customized evaluators for diverse NLP tasks . In Proceedings of the 31st International Conference on Computational Ling...
2025
-
[34]
Chin-Yew Lin. 2004. https://api.semanticscholar.org/CorpusID:964287 ROUGE : A package for automatic evaluation of summaries . In Annual Meeting of the Association for Computational Linguistics
2004
-
[35]
Dongqi Liu, Yifan Wang, Jia Loy, and Vera Demberg. 2024. https://aclanthology.org/2024.lrec-main.1258/ S ci N ews: From scholarly complexities to public narratives -- a dataset for scientific news report generation . In Proceedings of the 2024 Joint International Conference on...
2024
-
[36]
Yang Liu, Dan Iter, Yichong Xu, Shuo Wang, Ruochen Xu, and Chenguang Zhu. 2023 a . https://api.semanticscholar.org/CorpusID:257804696 G-eval: Nlg evaluation using gpt-4 with better human alignment . In Conference on Empirical Methods in Natural Language Processing
2023
-
[37]
Yiqi Liu, Nafise Sadat Moosavi, and Chenghua Lin. 2023 b . https://api.semanticscholar.org/CorpusID:265220860 LLM s as narcissistic evaluators: When ego inflates evaluation scores . In Annual Meeting of the Association for Computational Linguistics
2023
-
[38]
Yixin Liu, Alexander Fabbri, Yilun Zhao, Pengfei Liu, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, and Dragomir Radev. 2023 c . https://doi.org/10.18653/v1/2023.emnlp-main.1018 Towards interpretable and efficient automatic reference-based summarization evaluation . In Proceedin...
2023 doi
-
[39]
Fabbri, Pengfei Liu, Yilun Zhao, Linyong Nan, Ruilin Han, Simeng Han, Shafiq R
Yixin Liu, Alexander R. Fabbri, Pengfei Liu, Yilun Zhao, Linyong Nan, Ruilin Han, Simeng Han, Shafiq R. Joty, Chien-Sheng Wu, Caiming Xiong, and Dragomir R. Radev. 2022. https://api.semanticscholar.org/CorpusID:254685611 Revisiting the gold standard: Grounding summarization ev...
2022
-
[40]
Zheheng Luo, Qianqian Xie, and Sophia Ananiadou. 2022. https://api.semanticscholar.org/CorpusID:252780390 Readability controllable biomedical document summarization . ArXiv, abs/2210.04705
2022 arXiv
-
[41]
Laura Manor and Junyi Jessy Li. 2019. https://doi.org/10.18653/v1/W19-2201 Plain E nglish summarization of contracts . In Proceedings of the Natural Legal Language Processing Workshop 2019, pages 1--11, Minneapolis, Minnesota. Association for Computational Linguistics
2019 doi
-
[42]
M. L. McHugh. 2012. https://api.semanticscholar.org/CorpusID:5421278 Interrater reliability: the kappa statistic . Biochemia Medica, 22:276 -- 282
2012
-
[43]
Rostislav Nedelchev, Jens Lehmann, and Ricardo Usbeck. 2020. https://doi.org/10.18653/v1/2020.coling-main.599 Language model transformers as evaluators for open-domain dialogues . In Proceedings of the 28th International Conference on Computational Linguistics, pages 6797--680...
2020 doi
-
[44]
John O'hayre. 1966. Gobbledygook has gotta go. US Department of the Interior, Bureau of Land Management
1966
-
[45]
Timothy Rush
R. Timothy Rush. 1985. http://www.jstor.org/stable/20199072 Assessing readability: Formulas and alternatives . The Reading Teacher, 39(3):274--283
1985
-
[46]
Roy Schwartz, Jesse Dodge, Noah A Smith, and Oren Etzioni. 2020. https://dl.acm.org/doi/pdf/10.1145/3381831 Green ai . Communications of the ACM, 63(12):54--63
2020 doi
-
[47]
Chenhui Shen, Liying Cheng, Yang You, and Lidong Bing. 2023. https://api.semanticscholar.org/CorpusID:258833685 Large language models are not yet human-level evaluators for abstractive summarization . In Conference on Empirical Methods in Natural Language Processing
2023
-
[48]
Johannes Sibeko and Menno van Zaanen. 2022. https://doi.org/10.55492/dhasa.v3i01.3864 An analysis of readability metrics on english exam . Journal of the Digital Humanities Association of Southern Africa, 3(01)
2022 doi
-
[49]
Edgar A Smith and RJ Senter. 1967. Automated readability index, volume 66. Aerospace Medical Research Laboratories, Aerospace Medical Division, Air Force Systems Command
1967
-
[50]
Hwanjun Song, Hang Su, Igor Shalyminov, Jason Cai, and Saab Mansour. 2024. https://doi.org/10.18653/v1/2024.acl-long.51 F ine S ur E : Fine-grained summarization evaluation using LLM s . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics...
2024 doi
-
[51]
George Spache. 1953. A new readability formula for primary-grade reading materials. The Elementary School Journal, 53(7):410--413
1953
-
[52]
Brown, Adam Santoro, Aditya Gupta, Adri\` a Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adri\` a Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W. Kocurek, ...
2023 arXiv
-
[53]
Sanja S tajner, Richard Evans, Constantin Orasan, and Ruslan Mitkov. 2012. What can readability measures really tell us about text complexity. In Proceedings of workshop on natural language processing for improving textual accessibility, pages 14--22. Citeseer
2012
-
[54]
Loukritia Stefanou, Tatiana Passali, and Grigorios Tsoumakas. 2024. Auth at biolaysumm 2024: Bringing scientific content to kids. In Proceedings of the ACL 2024 BioNLP Workshop, Bangkok, Thailand. A paper presented at the BioLaySumm 2024 shared task on lay summarization of bio...
2024
-
[55]
Rickard Stureborg, Dimitris Alikaniotis, and Yoshi Suhara. 2024. https://api.semanticscholar.org/CorpusID:269588132 Large language models are inconsistent and biased evaluators . ArXiv, abs/2405.01724
2024 arXiv
-
[56]
Teerapaun Tanprasert and David Kauchak. 2021. https://api.semanticscholar.org/CorpusID:236486164 Flesch-kincaid is not a text simplification evaluation metric . Proceedings of the 1st Workshop on Natural Language Generation, Evaluation, and Metrics (GEM 2021)
2021
-
[57]
Gemma Team. 2024. https://doi.org/10.34740/KAGGLE/M/3301 Gemma
2024 doi
-
[58]
Thorndike
Edward L. Thorndike. 1936. The Elementary School Journal, http://www.jstor.org/stable/995916 36(6):470--472
1936
-
[59]
Pranav Narayanan Venkit, Tatiana Chakravorti, Vipul Gupta, Heidi Biggs, Mukund Srinath, Koustava Goswami, Sarah Rajtmajer, and Shomir Wilson. 2024. https://api.semanticscholar.org/CorpusID:269042872 An audit on the perspectives and challenges of hallucinations in nlp . In Conf...
2024
-
[60]
Murdick, Devvret Rishi, Jerry Sheehan, Zhihong Shen, Brandon Stilson, Alex D
Lucy Lu Wang, Kyle Lo, Yoganand Chandrasekhar, Russell Reas, Jiangjiang Yang, Doug Burdick, Darrin Eide, Kathryn Funk, Yannis Katsis, Rodney Michael Kinney, Yunyao Li, Ziyang Liu, William Merrill, Paul Mooney, Dewey A. Murdick, Devvret Rishi, Jerry Sheehan, Zhihong Shen, Brand...
2020
-
[61]
Peiyi Wang, Lei Li, Liang Chen, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. 2023. https://api.semanticscholar.org/CorpusID:258960339 Large language models are not fair evaluators . ArXiv, abs/2305.17926
2023 arXiv
-
[62]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824--24837
2022
-
[63]
Sohee Yang, Elena Gribovskaya, Nora Kassner, Mor Geva, and Sebastian Riedel. 2024. https://api.semanticscholar.org/CorpusID:268032051 Do large language models latently perform multi-hop reasoning? In Annual Meeting of the Association for Computational Linguistics
2024
-
[64]
Aljohani, and Raheel Nawaz
Farooq Zaman, Matthew Shardlow, Saeed-Ul Hassan, Naif R. Aljohani, and Raheel Nawaz. 2020. https://api.semanticscholar.org/CorpusID:224867112 HTSS : A novel hybrid text summarisation and simplification architecture . Inf. Process. Manag., 57:102351
2020
-
[65]
Weinberger, and Yoav Artzi
Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020. https://openreview.net/forum?id=SkeHuCVFDr BERTS core: Evaluating text generation with bert . In International Conference on Learning Representations
2020
-
[66]
Xiaoyu Zhang, Yishan Li, Jiayin Wang, Bowen Sun, Weizhi Ma, Peijie Sun, and Min Zhang. 2024. https://doi.org/10.1145/3640457.3688075 Large language models as evaluators for recommendation explanations . In Proceedings of the 18th ACM Conference on Recommender Systems, RecSys '...
2024
-
[67]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[68]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.