Pith. sign in

REVIEW 4 major objections 5 minor 3 cited by

Evaluating the Evaluators: Are readability metrics good measures of readability?

T0 review · 4 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Traditional readability formulas mostly fail to match human judgments of plain-language summaries, while language models agree much better.

desk verdict The negative finding on FKGL is real but the paper overstates the LM advantage and needs confidence intervals before it should drive practice. read the letter →

arxiv 2508.19221 v1 pith:2VSVJRBM submitted 2025-08-26 cs.CL

classification cs.CL
keywords plainlanguagesummarizationreadabilitymetricsFlesch-KincaidGradeLevellanguage-modelevaluatorshumanjudgmentsscientificdatasetsevaluationmethodologyLLM-as-judge
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether the formulas used to grade readability in plain-language summarization research actually measure what readers experience. Surveying 18 papers in the ACL Anthology, it finds Flesch-Kincaid Grade Level (FKGL) is the standard readability check. Measured against 60 human-rated plain-language summaries, 6 of 8 traditional formulas correlate below 0.3 with human judgments; FKGL scores only 0.16. Language models asked to rate readability with their own judgment do better, with the best model reaching 0.56, and they justify scores by pointing to explanations and background knowledge rather than word length. The paper concludes that FKGL should be dropped for PLS evaluation and that two widely used 'plain language' datasets, PLOS and CELLS, score like expert-targeted abstracts under the better judge.

What carries the argument

The load-bearing object is a single gold-standard benchmark: 60 summaries of 10 scientific papers rated for reading ease by crowd workers on a 1-5 scale, averaged per summary (August et al. 2024). The paper measures each candidate evaluator by its Pearson and Kendall-Tau correlation with those averages: 8 traditional formulas (FKGL, FRE, DCRS, ARI, CLI, GFI, Spache, Linsear Write) versus 5 LMs prompted to use their own judgment and explain their score. The LM reasoning text is then mined with YAKE keyword extraction to show the models appeal to explanation of terms and required background knowledge. A second mechanism is the dataset-level comparison: mean scores across 10 summarization datas

What would settle it

Gather a fresh human-evaluated readability corpus for plain-language summaries covering, say, 30+ papers in several domains, and recompute the correlations. The paper's claim would be falsified if FKGL's Pearson correlation reaches or exceeds the best LM's, or if the best LM's correlation drops below 0.3; a finding that PLOS and CELLS score high on the new corpus would also undercut the dataset conclusions.

Watch

Extended reading notes

Core claim

The paper's central claim is that readability, for plain language summaries, is not the property that traditional readability formulas measure. Formulas built on syllable counts, sentence lengths, and word lists label a summary of acute respiratory distress syndrome as college-level even when human readers find it clear, because they cannot see that 'a very serious lung disease' defines the term. On the only available human-judgment dataset for PLS (60 summaries of 10 scientific papers, rated 1-5), FKGL correlates 0.16 (Pearson), while all five tested LMs correlate above 0.45, led by Llama 3.3 70B at 0.56; the best traditional metrics, DCRS and CLI, reach only about 0.37. The paper extends t

Load-bearing premise

The comparison stands entirely on one gold standard: 60 human-readability ratings of summaries of 10 English science papers, collected for a different study; if those ratings are noisy or unrepresentative of plain-language readability, every correlation and dataset ranking inherits the error.

Editorial extensions

If this is right

  • Papers and shared tasks that use FKGL as the primary readability check for plain-language summaries should stop: its correlation with human readability is near chance (0.16).
  • DCRS and CLI are the strongest of the traditional formulas but still modest; the paper recommends pairing them with LM evaluators rather than relying on either alone.
  • Off-the-shelf LMs of 7B-70B are workable readability judges; their written reasons can be inspected to see whether they reward definitions of technical terms and penalize missing background.
  • The PLOS and CELLS datasets, widely used in PLS shared tasks, are better regarded as general scientific summarization data, not plain-language data; CDSR and SciNews fit the PLS label better.
  • Readability metrics for PLS should be redesigned around the human notion—explanations, context, background—rather than lexical complexity.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the gold standard is representative, the same blind spot likely affects readability claims in health communication, legal documents, and other accessibility work that leans on grade-level formulas, because those formulas systematically reward short acronyms and penalize long explanatory words.
  • The 0.56 ceiling leaves room: prompt tuning, calibrated scales, or fine-tuned judge models may push LM-human agreement higher, and the near-tie among small and large LMs suggests capability is not the main constraint.
  • Because 60 summaries are nested in only 10 source papers, the dataset-level rankings (especially PLOS and CELLS) should be read as preliminary; a broader corpus could shift those means.
  • If LM judges replace formulas, their known biases and inconsistency need auditing before they become the new default; the paper itself flags this.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper surveys evaluation practice in Plain Language Summarization (PLS), finding that FKGL is the most commonly used readability metric in ACL PLS papers. It then compares eight traditional readability metrics and five LLM-based judges against human readability judgments from the August et al. (2024) dataset of 60 summaries over 10 scientific papers. The central empirical claims are that six of eight traditional metrics correlate below 0.3 with human judgments (FKGL at Pearson r=0.16), that LLM evaluators correlate higher (best Llama 3.3 70B at r=0.56), and that applying the best LLM judge to ten summarization datasets yields different dataset-level conclusions than FKGL, including the recommendation that PLOS and CELLS are not plain-language datasets. The authors release analysis code and survey data, and make concrete recommendations for future PLS evaluation.

Significance. If the empirical claims hold, the paper makes a useful and timely contribution: it is the first direct comparison of standard readability formulas with human readability judgments in the PLS setting, and it offers actionable guidance to a community that currently relies on FKGL. The central comparison is not circular: the LLM scores are zero-shot judgments on a held-out human-judgment set, not fitted to those labels, and the traditional metric scores come from an independent package. The survey of ACL PLS evaluation practices and the release of code and data are concrete strengths. However, the empirical foundation is narrow—a single 60-summary, 10-source-paper human dataset—and the paper's confidence intervals and significance tests are either absent or undercut the headline contrasts. The stress-test concern about circularity does not land, but the concern about statistical support does.

major comments (4)
  1. [§3.2, Tables 2a/2b, Appendix A] All headline correlations rest on n=60 summaries nested in 10 source papers, yet the paper reports no confidence intervals and treats the summaries as independent. At n=60, FKGL's r=0.16 is not significantly different from zero (approximate p≈0.22), and clustering by paper would widen the uncertainty further. This uncertainty propagates into every comparison in Tables 2a/2b and into the §4.4 dataset conclusions. The Limitations section acknowledges domain and language limits, but not this sampling limitation. The authors should report cluster-robust confidence intervals, mixed-effects or leave-one-paper-out analyses, and adjust the strength of the claims accordingly.
  2. [Appendix C, Table 9] The Williams tests reported by the authors themselves undercut the abstract's 'LMs are better judges' claim. Llama 3.3 70B is not significantly better than DCRS (p=0.14) or CLI (p=0.13), and Llama 3.1 8B is not significantly better than FKGL (p=0.06). The only robust pairwise improvements are against FKGL for the stronger models. The paper should present these p-values in the main text and qualify the claim, or add enough data/evidence to support the stronger reading.
  3. [§4.4, Tables 4 and 5] The recommendation that PLOS and CELLS 'may not be well-suited for PLS' and the Cohen's Kappa=0.17 disagreement analysis implicitly treat the Llama 3.3 70B scores as ground truth on datasets for which no human readability judgments are available. The expert/kid sanity checks are useful, but they do not establish that the LM's 1–5 scale is calibrated across these new datasets. Furthermore, the binary thresholds (LM score ≥3 vs FKGL score <12) are chosen without justification and directly determine Kappa and the dataset-level conclusions. The authors should report sensitivity to threshold choices or validate on a human-rated sample from these datasets.
  4. [Appendix A and §3.2] The human gold standard is the only available PLS human judgment dataset, but it is also narrow: 10 source papers sampled from r/science, 6 summaries per paper (2 expert, 4 GPT-3 generated), and binarized Cohen's Kappa of 0.6. The paper's general conclusion that traditional metrics are poor measures of readability for PLS assumes this set is representative of PLS outputs more broadly. I would like to see a sensitivity analysis that removes one source paper at a time, and a discussion of how the expert/GPT-3 composition and the r/science sampling may affect the correlations. This is not fatally circular, but it is a load-bearing limitation.
minor comments (5)
  1. [Table 9 and Table 2b] The row header 'Llama 3.1 70B' in Table 9 conflicts with 'Llama 3.3 70B' used in Table 2b and throughout the text.
  2. [Table 4 and §4.4] The dataset is referred to as 'SKJ' in Table 4 but as 'SJK' (Science Journal for Kids) elsewhere.
  3. [Table 3] The caption appears to mislabel its own panels: '3b contains an example summary' should likely refer to panel (a), and the description of panel (b) is duplicated.
  4. [§3.2] Typo: 'Flesh-Kincaid' should be 'Flesch-Kincaid'.
  5. [References] The Gunning Fog Index is cited to Isnaeni (2017), but the original source is Gunning (1952), which is listed in the bibliography.

Circularity Check

1 steps flagged · score 2.0 of 10

Core comparison is an external benchmark, but the headline LM correlation is selected from prompts tuned on the same 60 human-scored summaries, giving a mild fitted-input bias rather than construction-level circularity.

  1. fitted input called prediction [Section 3.3, Appendix B (Table 8 vs Table 2b)]
    "We experiment with 3 prompts and report the prompts in appendix B. We report the Pearson and Kendall-Tau correlations of the scores provided by each LM with the human judgments. ... The Own Reasoning Prompt performs the best when averaged across all models."

    The paper's headline result, Llama 3.3 70B achieving Pearson r=0.56 with human judgments, is the correlation of the specific prompt that was selected as best on the same 60 human-scored summaries used to compute that correlation. The prompt choice is thus fitted to the gold-standard labels, and the reported value is a selected maximum over three prompt variants (and implicitly over five models) rather than an unbiased out-of-sample estimate. This is a mild form of fitted-input-called-prediction: the 'best-performing model' number is not independent of the human labels it is compared against. It does not make the correlation tautological or reduce it by construction, because the LM scores themselves are not trained on the human labels, so the circularity is minor rather than central.

full rationale

The central comparison in the paper is an external benchmark. Traditional readability scores come from a standard package (py-readability-metrics), and the LM scores are generated by prompting models that were not trained or fine-tuned on the 60 human judgments from August et al. (2024). The human gold standard is an external dataset, not a self-citation, and no uniqueness theorem is imported from the authors' prior work. The survey of PLS evaluation practices (RQ1) is a literature count, not a derivation. The dataset-level conclusions in Section 4.4 are extrapolations of the best LM's judgments to other corpora; they are not circular, though they inherit the statistical limitations of the small gold standard. The only mild circularity is the selection of the best prompt on the same 60-summary test set before reporting the headline correlation. Because the paper discloses all prompt results in Appendix B and does not hide the selection, this is a transparency/selection-bias issue rather than a construction-level equivalence. No equation in the paper is defined in terms of the target result, and no fitted parameter is renamed as a prediction. Therefore the overall circularity is low.

Assumptions & free parameters 2 free parameters · 5 assumptions · 0 invented entities

The paper fits no parameters in its main computations: the 8 classical metrics are standard formulas computed by an external package, and the LM scores are held-out judgments from pretrained models, which keeps the headline correlations external. The hand-set choices are the binary readability thresholds used for the Kappa analysis, and the conclusions lean on domain assumptions about the human gold standard, the independence of the 60 summaries, and the representativeness of a 55-paper ACL Anthology survey. No new theoretical entities are introduced.

free parameters (2)
  • FKGL high-readability threshold = 12 (grade level)
    Chosen by hand in Section 4.4 to binarize datasets into high/low readability for the Cohen's Kappa agreement test; the reported Kappa of 0.17 between FKGL and the LM evaluator depends on this cut.
  • LM high-readability threshold = 3 on a 1 to 5 scale
    Chosen by hand in Section 4.4 to match the binarization of human scores in Appendix A; the same cut is applied to Llama 3.3 70B scores to compute the 0.17 Kappa.
assumptions (5)
  • domain assumption Human judgments collected by August et al. (2024), averaged per summary, are a valid gold standard for PLS readability.
    Invoked in Section 3.2; the entire correlation analysis and all dataset-level conclusions rest on this single 60-summary dataset.
  • domain assumption The 60 summaries can be treated as 60 independent observations for correlation and significance testing.
    Summaries are nested within 10 papers (Section 3.2, Appendix A); within-paper topic difficulty may induce clustering, and the paper does not use a mixed-effects or cluster-robust analysis.
  • standard math Pearson correlation between averaged human 1 to 5 scores and metric scores is an appropriate measure of evaluation quality.
    Used throughout Sections 3.2 and 4.2; Pearson assumes interval scale and linearity, a modeling choice for Likert-derived averages.
  • domain assumption The ACL Anthology query (55 hits, 18 PLS-relevant papers) fairly represents PLS evaluation practice.
    RQ1 survey in Section 3.1; small corpus, single venue family, manual relevance filtering.
  • domain assumption LM scores from the chosen prompt and decoding are stable enough to rank datasets.
    Section 4.4 uses a single prompt and a single model run per summary; no temperature or seed variance is reported.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Evaluating the Evaluators: Are readability metrics good measures of readability?." pith.science (2026). https://pith.science/paper/2VSVJRBM

@misc{pith2026250819221,
  author       = {Pith},
  title        = {Pith review of: Evaluating the Evaluators: Are readability metrics good measures of readability?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2VSVJRBM}},
  note         = {Machine review of arXiv:2508.19221}
}
read the original abstract

Plain Language Summarization (PLS) aims to distill complex documents into accessible summaries for non-expert audiences. In this paper, we conduct a thorough survey of PLS literature, and identify that the current standard practice for readability evaluation is to use traditional readability metrics, such as Flesch-Kincaid Grade Level (FKGL). However, despite proven utility in other fields, these metrics have not been compared to human readability judgments in PLS. We evaluate 8 readability metrics and show that most correlate poorly with human judgments, including the most popular metric, FKGL. We then show that Language Models (LMs) are better judges of readability, with the best-performing model achieving a Pearson correlation of 0.56 with human judgments. Extending our analysis to PLS datasets, which contain summaries aimed at non-expert audiences, we find that LMs better capture deeper measures of readability, such as required background knowledge, and lead to different conclusions than the traditional metrics. Based on these findings, we offer recommendations for best practices in the evaluation of plain language summaries. We release our analysis code and survey data.

Figures

Figures reproduced from arXiv: 2508.19221 by the authors.

Figure 1
Figure 1. Evaluation metrics used by papers published in the [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. The best performing prompt of the 3 we tested. We [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Histogram of LM readability scores and the mean scores ( [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Keywords mentioned in the reasoning of the LM [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Disco-RAG: Discourse-Aware Retrieval-Augmented Generation

    cs.CL 2026-01 unverdicted novelty 7.0 of 10

    Disco-RAG improves RAG by building intra-chunk discourse trees and inter-chunk rhetorical graphs that feed into a planning blueprint, delivering state-of-the-art results on question answering and long-document summari...

  2. OmniPresent: Generating Coherent Presentation Suites from Scientific Papers

    cs.SE 2026-07 conditional novelty 6.5 of 10

    A multi-agent HTML pipeline with shared knowledge and cross-artifact verify-and-repair generates coherent poster/slides/video/page suites from papers and beats specialized baselines on OmniPreBench.

  3. Machine Learning Research Has Outpaced Its Communication Norms and NeurIPS Should Act

    cs.LG 2026-05 conditional novelty 5.0 of 10

    NeurIPS papers have grown much harder to read since 1987, with acronym use tripling and sensational language rising, and the authors propose seven concrete writing standards for the conference to pilot.

Reference graph

Works this paper leans on

68 extracted references · 49 canonical work pages · cited by 3 Pith papers

  1. [1]

    Smith, and Katharina Reinecke

    Tal August, Kyle Lo, Noah A. Smith, and Katharina Reinecke. 2024. https://api.semanticscholar.org/CorpusID:268297229 Know your audience: The benefits and pitfalls of generating plain language summaries beyond the "general" audience . Proceedings of the CHI Conference on Human Factors in Computing Systems

  2. [2]

    Tal August, Katharina Reinecke, and Noah A. Smith. 2022. https://api.semanticscholar.org/CorpusID:248780294 Generating scientific definitions with controllable complexity . In Annual Meeting of the Association for Computational Linguistics

  3. [3]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gr...

  4. [4]

    Isabel Cachola, Kyle Lo, Arman Cohan, and Daniel Weld. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.428 TLDR : Extreme summarization of scientific documents . In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4766--4777, Online. Association for Computational Linguistics

  5. [5]

    Ricardo Campos, Vítor Mangaravite, Arian Pasquali, Alípio Jorge, Célia Nunes, and Adam Jatowt. 2020. https://doi.org/10.1016/j.ins.2019.09.013 Yake! keyword extraction from single documents using multiple local features . Information Sciences, 509:257--289

  6. [6]

    Afonso Cavaco Carla Pires and Marina Vigário. 2017. https://doi.org/10.1080/09296174.2017.1311448 Towards the definition of linguistic metrics for evaluating text readability . Journal of Quantitative Linguistics, 24(4):319--349

  7. [7]

    Muthu Kumar Chandrasekaran, Guy Feigenblat, Eduard Hovy, Abhilasha Ravichander, Michal Shmueli-Scheuer, and Anita de Waard. 2020. https://doi.org/10.18653/v1/2020.sdp-1.24 Overview and insights from the shared tasks at scholarly document processing 2020: CL - S ci S umm, L ay S umm and L ong S umm . In Proceedings of the First Workshop on Scholarly Docume...

  8. [8]

    Chang, and Nazli Goharian

    Arman Cohan, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Seokhwan Kim, W. Chang, and Nazli Goharian. 2018. https://api.semanticscholar.org/CorpusID:4894594 A discourse-aware attention model for abstractive summarization of long documents . In North American Chapter of the Association for Computational Linguistics

Show all 68 references
  1. [9]

    Meri Coleman and Ta Lin Liau. 1975. https://api.semanticscholar.org/CorpusID:144250124 A computer readability formula designed for machine scoring. Journal of Applied Psychology, 60:283--284

  2. [10]

    Scott Andrew Crossley, Aron Heintz, Joon Suh Choi, Jordan Batchelor, Mehrnoush Karimi, and Agnes Malatinszky. 2021. https://api.semanticscholar.org/CorpusID:247321796 The commonlit ease of readability (clear) corpus . In Educational Data Mining

  3. [11]

    Edgar Dale and Jeanne S Chall. 1948. A formula for predicting readability: Instructions. Educational research bulletin, pages 37--54

  4. [12]

    William H DuBay. 2004. The principles of readability. Impact Information

  5. [13]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zha...

  6. [14]

    A. R. Fabbri, Wojciech Kryscinski, Bryan McCann, Richard Socher, and Dragomir R. Radev. 2020. https://api.semanticscholar.org/CorpusID:220768873 Summeval: Re-evaluating summarization evaluation . Transactions of the Association for Computational Linguistics, 9:391--409

  7. [15]

    simplification of flesch reading ease formula

    Rudolf Flesch. 1952. "simplification of flesch reading ease formula". Journal of Applied Psychology

  8. [16]

    Rudolf Franz Flesch. 1943. https://api.semanticscholar.org/CorpusID:191201423 Marks of readable style: a study in adult education . In Teachers College Contributions to Education

  9. [17]

    Lorenzo Jaime Flores, Heyuan Huang, Kejian Shi, Sophie Chheang, and Arman Cohan. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.322 Medical text simplification: Optimizing for readability with unlikelihood training and reranked beam search decoding . In Findings of the ...

  10. [18]

    Tomas Goldsack, Zheheng Luo, Qianqian Xie, Carolina Scarton, Matthew Shardlow, Sophia Ananiadou, and Chenghua Lin. 2023. Overview of the biolaysumm 2023 shared task on lay summarization of biomedical research articles. In Proceedings of the 22st Workshop on Biomedical Language...

  11. [19]

    Tomas Goldsack, Carolina Scarton, Matthew Shardlow, and Chenghua Lin. 2024. Overview of the biolaysumm 2024 shared task on the lay summarization of biomedical research articles. In The 23rd Workshop on Biomedical Natural Language Processing and BioNLP Shared Tasks, Bangkok, Th...

  12. [20]

    Tomas Goldsack, Zhihao Zhang, Chenghua Lin, and Carolina Scarton. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.724 Making science simple: Corpora for the lay summarisation of scientific literature . In Proceedings of the 2022 Conference on Empirical Methods in Natural Lan...

  13. [21]

    Tanya Goyal, Junyi Jessy Li, and Greg Durrett. 2022. https://api.semanticscholar.org/CorpusID:252532176 News summarization and evaluation in the era of GPT-3 . ArXiv, abs/2209.12356

  14. [22]

    Yvette Graham and Timothy Baldwin. 2014. https://doi.org/10.3115/v1/D14-1020 Testing for significance of increased correlation with human judgment . In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing ( EMNLP ) , pages 172--176, Doha, Qata...

  15. [23]

    Robert Gunning. 1952. https://books.google.com/books?id=ofI0AAAAMAAJ The Technique of Clear Writing . McGraw-Hill

  16. [24]

    Cohen, and Lucy Lu Wang

    Yue Guo, Tal August, Gondy Leroy, Trevor A. Cohen, and Lucy Lu Wang. 2023. https://api.semanticscholar.org/CorpusID:258841161 APPLS : Evaluating evaluation metrics for plain language summarization . In Conference on Empirical Methods in Natural Language Processing

  17. [25]

    Yue Guo, Wei Qiu, Gondy Leroy, Sheng Wang, and Trevor A. Cohen. 2022. https://api.semanticscholar.org/CorpusID:253397732 Retrieval augmentation of large language models for lay language generation . Journal of biomedical informatics, page 104580

  18. [26]

    Yue Guo, Wei Qiu, Yizhong Wang, and Trevor Cohen. 2021. Automated lay language summarization of biomedical scientific reviews. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 160--168

  19. [27]

    Yu Han, Aaron Ceross, and Jeroen Bergmann. 2024. https://api.semanticscholar.org/CorpusID:274023239 The use of readability metrics in legal text: A systematic literature review . ArXiv, abs/2411.09497

  20. [28]

    Nur Rachma Isnaeni. 2017. https://journal3.uin-alauddin.ac.id/index.php/elite/article/view/3356 Readability of english written materials . Elite : English and Literature Journal, 1(1):179--191

  21. [29]

    Yuelyu Ji, Zhuochun Li, Rui Meng, Sonish Sivarajkumar, Yanshan Wang, Zeshui Yu, Hui Ji, Yushui Han, Hanyu Zeng, and Daqing He. 2024. https://doi.org/10.18653/v1/2024.bionlp-1.75 RAG - RLRC - L ay S um at B io L ay S umm: Integrating retrieval-augmented generation and readabili...

  22. [30]

    Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024. Mixtral of experts. arXiv preprint arXiv:2401.04088

  23. [31]

    Albert Qiaochu Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L'elio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, T...

  24. [32]

    Daniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie, Jungo Kasai, Yejin Choi, Noah A Smith, and Daniel S Weld. 2022. https://arxiv.org/abs/2101.06561 GENIE: Toward Reproducible and Standardized Human Evaluation for Text Generation . In Conference on Empirical M...

  25. [33]

    Qintong Li, Leyang Cui, Lingpeng Kong, and Wei Bi. 2025. https://aclanthology.org/2025.coling-main.688/ Exploring the reliability of large language models as customized evaluators for diverse NLP tasks . In Proceedings of the 31st International Conference on Computational Ling...

  26. [34]

    Chin-Yew Lin. 2004. https://api.semanticscholar.org/CorpusID:964287 ROUGE : A package for automatic evaluation of summaries . In Annual Meeting of the Association for Computational Linguistics

  27. [35]

    Dongqi Liu, Yifan Wang, Jia Loy, and Vera Demberg. 2024. https://aclanthology.org/2024.lrec-main.1258/ S ci N ews: From scholarly complexities to public narratives -- a dataset for scientific news report generation . In Proceedings of the 2024 Joint International Conference on...

  28. [36]

    Yang Liu, Dan Iter, Yichong Xu, Shuo Wang, Ruochen Xu, and Chenguang Zhu. 2023 a . https://api.semanticscholar.org/CorpusID:257804696 G-eval: Nlg evaluation using gpt-4 with better human alignment . In Conference on Empirical Methods in Natural Language Processing

  29. [37]

    Yiqi Liu, Nafise Sadat Moosavi, and Chenghua Lin. 2023 b . https://api.semanticscholar.org/CorpusID:265220860 LLM s as narcissistic evaluators: When ego inflates evaluation scores . In Annual Meeting of the Association for Computational Linguistics

  30. [38]

    Yixin Liu, Alexander Fabbri, Yilun Zhao, Pengfei Liu, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, and Dragomir Radev. 2023 c . https://doi.org/10.18653/v1/2023.emnlp-main.1018 Towards interpretable and efficient automatic reference-based summarization evaluation . In Proceedin...

  31. [39]

    Fabbri, Pengfei Liu, Yilun Zhao, Linyong Nan, Ruilin Han, Simeng Han, Shafiq R

    Yixin Liu, Alexander R. Fabbri, Pengfei Liu, Yilun Zhao, Linyong Nan, Ruilin Han, Simeng Han, Shafiq R. Joty, Chien-Sheng Wu, Caiming Xiong, and Dragomir R. Radev. 2022. https://api.semanticscholar.org/CorpusID:254685611 Revisiting the gold standard: Grounding summarization ev...

  32. [40]

    Zheheng Luo, Qianqian Xie, and Sophia Ananiadou. 2022. https://api.semanticscholar.org/CorpusID:252780390 Readability controllable biomedical document summarization . ArXiv, abs/2210.04705

  33. [41]

    Laura Manor and Junyi Jessy Li. 2019. https://doi.org/10.18653/v1/W19-2201 Plain E nglish summarization of contracts . In Proceedings of the Natural Legal Language Processing Workshop 2019, pages 1--11, Minneapolis, Minnesota. Association for Computational Linguistics

  34. [42]

    M. L. McHugh. 2012. https://api.semanticscholar.org/CorpusID:5421278 Interrater reliability: the kappa statistic . Biochemia Medica, 22:276 -- 282

  35. [43]

    Rostislav Nedelchev, Jens Lehmann, and Ricardo Usbeck. 2020. https://doi.org/10.18653/v1/2020.coling-main.599 Language model transformers as evaluators for open-domain dialogues . In Proceedings of the 28th International Conference on Computational Linguistics, pages 6797--680...

  36. [44]

    John O'hayre. 1966. Gobbledygook has gotta go. US Department of the Interior, Bureau of Land Management

  37. [45]

    Timothy Rush

    R. Timothy Rush. 1985. http://www.jstor.org/stable/20199072 Assessing readability: Formulas and alternatives . The Reading Teacher, 39(3):274--283

  38. [46]

    Roy Schwartz, Jesse Dodge, Noah A Smith, and Oren Etzioni. 2020. https://dl.acm.org/doi/pdf/10.1145/3381831 Green ai . Communications of the ACM, 63(12):54--63

  39. [47]

    Chenhui Shen, Liying Cheng, Yang You, and Lidong Bing. 2023. https://api.semanticscholar.org/CorpusID:258833685 Large language models are not yet human-level evaluators for abstractive summarization . In Conference on Empirical Methods in Natural Language Processing

  40. [48]

    Johannes Sibeko and Menno van Zaanen. 2022. https://doi.org/10.55492/dhasa.v3i01.3864 An analysis of readability metrics on english exam . Journal of the Digital Humanities Association of Southern Africa, 3(01)

  41. [49]

    Edgar A Smith and RJ Senter. 1967. Automated readability index, volume 66. Aerospace Medical Research Laboratories, Aerospace Medical Division, Air Force Systems Command

  42. [50]

    Hwanjun Song, Hang Su, Igor Shalyminov, Jason Cai, and Saab Mansour. 2024. https://doi.org/10.18653/v1/2024.acl-long.51 F ine S ur E : Fine-grained summarization evaluation using LLM s . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics...

  43. [51]

    George Spache. 1953. A new readability formula for primary-grade reading materials. The Elementary School Journal, 53(7):410--413

  44. [52]

    Brown, Adam Santoro, Aditya Gupta, Adri\` a Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W

    Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adri\` a Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W. Kocurek, ...

  45. [53]

    Sanja S tajner, Richard Evans, Constantin Orasan, and Ruslan Mitkov. 2012. What can readability measures really tell us about text complexity. In Proceedings of workshop on natural language processing for improving textual accessibility, pages 14--22. Citeseer

  46. [54]

    Loukritia Stefanou, Tatiana Passali, and Grigorios Tsoumakas. 2024. Auth at biolaysumm 2024: Bringing scientific content to kids. In Proceedings of the ACL 2024 BioNLP Workshop, Bangkok, Thailand. A paper presented at the BioLaySumm 2024 shared task on lay summarization of bio...

  47. [55]

    Rickard Stureborg, Dimitris Alikaniotis, and Yoshi Suhara. 2024. https://api.semanticscholar.org/CorpusID:269588132 Large language models are inconsistent and biased evaluators . ArXiv, abs/2405.01724

  48. [56]

    Teerapaun Tanprasert and David Kauchak. 2021. https://api.semanticscholar.org/CorpusID:236486164 Flesch-kincaid is not a text simplification evaluation metric . Proceedings of the 1st Workshop on Natural Language Generation, Evaluation, and Metrics (GEM 2021)

  49. [57]

    Gemma Team. 2024. https://doi.org/10.34740/KAGGLE/M/3301 Gemma

  50. [58]

    Thorndike

    Edward L. Thorndike. 1936. The Elementary School Journal, http://www.jstor.org/stable/995916 36(6):470--472

  51. [59]

    Pranav Narayanan Venkit, Tatiana Chakravorti, Vipul Gupta, Heidi Biggs, Mukund Srinath, Koustava Goswami, Sarah Rajtmajer, and Shomir Wilson. 2024. https://api.semanticscholar.org/CorpusID:269042872 An audit on the perspectives and challenges of hallucinations in nlp . In Conf...

  52. [60]

    Murdick, Devvret Rishi, Jerry Sheehan, Zhihong Shen, Brandon Stilson, Alex D

    Lucy Lu Wang, Kyle Lo, Yoganand Chandrasekhar, Russell Reas, Jiangjiang Yang, Doug Burdick, Darrin Eide, Kathryn Funk, Yannis Katsis, Rodney Michael Kinney, Yunyao Li, Ziyang Liu, William Merrill, Paul Mooney, Dewey A. Murdick, Devvret Rishi, Jerry Sheehan, Zhihong Shen, Brand...

  53. [61]

    Peiyi Wang, Lei Li, Liang Chen, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. 2023. https://api.semanticscholar.org/CorpusID:258960339 Large language models are not fair evaluators . ArXiv, abs/2305.17926

  54. [62]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824--24837

  55. [63]

    Sohee Yang, Elena Gribovskaya, Nora Kassner, Mor Geva, and Sebastian Riedel. 2024. https://api.semanticscholar.org/CorpusID:268032051 Do large language models latently perform multi-hop reasoning? In Annual Meeting of the Association for Computational Linguistics

  56. [64]

    Aljohani, and Raheel Nawaz

    Farooq Zaman, Matthew Shardlow, Saeed-Ul Hassan, Naif R. Aljohani, and Raheel Nawaz. 2020. https://api.semanticscholar.org/CorpusID:224867112 HTSS : A novel hybrid text summarisation and simplification architecture . Inf. Process. Manag., 57:102351

  57. [65]

    Weinberger, and Yoav Artzi

    Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020. https://openreview.net/forum?id=SkeHuCVFDr BERTS core: Evaluating text generation with bert . In International Conference on Learning Representations

  58. [66]

    Xiaoyu Zhang, Yishan Li, Jiayin Wang, Bowen Sun, Weizhi Ma, Peijie Sun, and Min Zhang. 2024. https://doi.org/10.1145/3640457.3688075 Large language models as evaluators for recommendation explanations . In Proceedings of the 18th ACM Conference on Recommender Systems, RecSys '...

  59. [67]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...

  60. [68]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.