{"id":"8fc8938e-081d-4481-9d9c-d171b6509db5","arxiv_id":"2508.19221","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Classic readability formulas correlate weakly with human readability judgments for plain-language summaries (FKGL r=0.16), the best LLM judge reaches r=0.56, and the two evaluator families rank datasets nearly opposite.","lead":"The most popular readability formula used to evaluate plain-language summaries barely agrees with human readers, and language models do noticeably better at judging what is readable. That flips routine evaluation practice in the field and changes which datasets and methods are trusted as readable.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Single 60-summary, 10-paper human gold standard drives all headline correlations; no confidence intervals or cluster-robust tests, and FKGL's r=0.16 is not significantly different from zero as reported.","rationale":"The reader's weakest assumption is the right one: the entire empirical core is a correlation against a single 60-summary human-judgment set. I sharpen it into a statistical-correctness concern that is checkable from the paper's own data. At n=60, FKGL r=0.16 is within noise of zero, and the paper itself reports that the best LM's advantage over DCRS/CLI is not significant (Appendix C). The lack of CIs and the nesting of summaries in 10 source papers mean the paper overstates the precision of its headline numbers. This is not an objection to the negative direction of the result—it may well be true that FKGL is a weak measure for PLS—but the paper's quantitative claims, and the §4.4 dataset recommendations built on them, should be read as provisional until the uncertainty is quantified. The cluster bootstrap is a direct, low-cost check using the released code and data. If the CIs are narrow and the ranking survives, the central claim is strengthened; if they are wide, CONDITIONAL remains the appropriate verdict, and the abstract's 'LMs are better judges' phrasing should be qualified. I therefore keep the reader's CONDITIONAL verdict unchanged.","tokens_in":23444,"tokens_out":9630,"duration_ms":105029,"concrete_test":"Recompute the Table 2 correlations with a cluster bootstrap over the 10 source papers (resample papers, keep their 6 summaries), reporting 95% CIs for FKGL, DCRS, CLI, and Llama 3.3 70B plus cluster-robust p-values for the LM-vs-traditional differences. If the FKGL CI includes 0, or the Llama-vs-DCRS/CLI CI includes 0, the headline ranking in §4.2–4.3 and the §4.4 dataset conclusions are not statistically supported at current precision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"All headline numbers in Tables 2a/2b and the dataset-level conclusions in §4.4 are computed against one gold standard: 60 summaries from 10 papers (6 per paper; Appendix A), with averaged annotator scores. The paper reports no confidence intervals, no cluster-robust inference, and treats the 60 summaries as independent observations despite nesting by source paper. At the reported n=60, FKGL's r=0.16 corresponds to p≈0.22, so the flagship 'most popular metric correlates only 0.16' is statistically indistinguishable from zero; clustering by paper would only widen the uncertainty. The claim that LMs are better judges is also weaker than the abstract implies: Appendix C's own Williams tests show Llama 3.3 70B is not significantly better than DCRS (p=0.14) or CLI (p=0.13), and Llama 3.1 8B is not significant over FKGL (p=0.06). Because every metric correlation and every LM-vs-traditional comparison uses the same 60-summary dataset, these uncertainties propagate directly into the recommendation to drop FKGL and into the §4.4 judgment that PLOS and CELLS are not plain-language datasets. The paper's limitations paragraph acknowledges domain and language limits, but it does not quantify the sampling error of the central correlations.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper surveys evaluation practice in Plain Language Summarization (PLS), finding that FKGL is the most commonly used readability metric in ACL PLS papers. It then compares eight traditional readability metrics and five LLM-based judges against human readability judgments from the August et al. (2024) dataset of 60 summaries over 10 scientific papers. The central empirical claims are that six of eight traditional metrics correlate below 0.3 with human judgments (FKGL at Pearson r=0.16), that LLM evaluators correlate higher (best Llama 3.3 70B at r=0.56), and that applying the best LLM judge to ten summarization datasets yields different dataset-level conclusions than FKGL, including the recommendation that PLOS and CELLS are not plain-language datasets. The authors release analysis code and survey data, and make concrete recommendations for future PLS evaluation.","tokens_in":23656,"tokens_out":4118,"duration_ms":45901,"significance":"If the empirical claims hold, the paper makes a useful and timely contribution: it is the first direct comparison of standard readability formulas with human readability judgments in the PLS setting, and it offers actionable guidance to a community that currently relies on FKGL. The central comparison is not circular: the LLM scores are zero-shot judgments on a held-out human-judgment set, not fitted to those labels, and the traditional metric scores come from an independent package. The survey of ACL PLS evaluation practices and the release of code and data are concrete strengths. However, the empirical foundation is narrow—a single 60-summary, 10-source-paper human dataset—and the paper's confidence intervals and significance tests are either absent or undercut the headline contrasts. The stress-test concern about circularity does not land, but the concern about statistical support does.","major_comments":[{"comment":"All headline correlations rest on n=60 summaries nested in 10 source papers, yet the paper reports no confidence intervals and treats the summaries as independent. At n=60, FKGL's r=0.16 is not significantly different from zero (approximate p≈0.22), and clustering by paper would widen the uncertainty further. This uncertainty propagates into every comparison in Tables 2a/2b and into the §4.4 dataset conclusions. The Limitations section acknowledges domain and language limits, but not this sampling limitation. The authors should report cluster-robust confidence intervals, mixed-effects or leave-one-paper-out analyses, and adjust the strength of the claims accordingly.","section":"§3.2, Tables 2a/2b, Appendix A"},{"comment":"The Williams tests reported by the authors themselves undercut the abstract's 'LMs are better judges' claim. Llama 3.3 70B is not significantly better than DCRS (p=0.14) or CLI (p=0.13), and Llama 3.1 8B is not significantly better than FKGL (p=0.06). The only robust pairwise improvements are against FKGL for the stronger models. The paper should present these p-values in the main text and qualify the claim, or add enough data/evidence to support the stronger reading.","section":"Appendix C, Table 9"},{"comment":"The recommendation that PLOS and CELLS 'may not be well-suited for PLS' and the Cohen's Kappa=0.17 disagreement analysis implicitly treat the Llama 3.3 70B scores as ground truth on datasets for which no human readability judgments are available. The expert/kid sanity checks are useful, but they do not establish that the LM's 1–5 scale is calibrated across these new datasets. Furthermore, the binary thresholds (LM score ≥3 vs FKGL score <12) are chosen without justification and directly determine Kappa and the dataset-level conclusions. The authors should report sensitivity to threshold choices or validate on a human-rated sample from these datasets.","section":"§4.4, Tables 4 and 5"},{"comment":"The human gold standard is the only available PLS human judgment dataset, but it is also narrow: 10 source papers sampled from r/science, 6 summaries per paper (2 expert, 4 GPT-3 generated), and binarized Cohen's Kappa of 0.6. The paper's general conclusion that traditional metrics are poor measures of readability for PLS assumes this set is representative of PLS outputs more broadly. I would like to see a sensitivity analysis that removes one source paper at a time, and a discussion of how the expert/GPT-3 composition and the r/science sampling may affect the correlations. This is not fatally circular, but it is a load-bearing limitation.","section":"Appendix A and §3.2"}],"minor_comments":[{"comment":"The row header 'Llama 3.1 70B' in Table 9 conflicts with 'Llama 3.3 70B' used in Table 2b and throughout the text.","section":"Table 9 and Table 2b"},{"comment":"The dataset is referred to as 'SKJ' in Table 4 but as 'SJK' (Science Journal for Kids) elsewhere.","section":"Table 4 and §4.4"},{"comment":"The caption appears to mislabel its own panels: '3b contains an example summary' should likely refer to panel (a), and the description of panel (b) is duplicated.","section":"Table 3"},{"comment":"Typo: 'Flesh-Kincaid' should be 'Flesch-Kincaid'.","section":"§3.2"},{"comment":"The Gunning Fog Index is cited to Isnaeni (2017), but the original source is Gunning (1952), which is listed in the bibliography.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"This is a borderline major revision. The paper's empirical base is thin but honestly presented, and the code/data release is a real asset. The primary blockers are statistical: missing confidence intervals, unaddressed nesting, and significance tests that already undermine the strongest claims. If the authors add uncertainty quantification, soften the 'better judges' claim to match the Williams tests, and mark the PLOS/CELLS conclusions as provisional or validate them on human ratings, the paper could be acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Hey [colleague],\n\nHere's my take on arXiv:2508.19221. The thing to know: it's the first direct comparison of eight classical readability formulas against human readability judgments for plain language summaries, and the central negative result—FKGL sits at r=0.16—is probably right. But the confidence you can place in the exact numbers is more limited than the paper's tone suggests, because everything rests on one gold standard of 60 summaries from 10 papers, with no confidence intervals or cluster-robust tests.\n\nWhat's genuinely good: the literature survey is useful and shows how dominant FKGL is in PLS evaluation, including shared tasks. The paper ships code and data. The example in Table 3 nicely illustrates why FKGL fails on a summary humans rate as readable. The dataset re-ranking is thought-provoking, and the keyword analysis gives some insight into what LMs attend to. The authors are honest in the limitations section.\n\nNow the soft spots. The gold standard is a single dataset collected for a different purpose, with moderate binarized annotator agreement (kappa 0.6). At n=60, FKGL's r=0.16 is not significantly different from zero—the paper doesn't report confidence intervals anywhere. The 60 summaries are nested in 10 papers, so treating them as independent inflates power. More importantly, the abstract claims LMs are better judges, but their own Williams tests in Appendix C show the best LM (Llama 3.3 70B) is not significantly better than DCRS (p=0.14) or CLI (p=0.13). The best prompt was selected on the same 60 summaries used for the headline numbers, so r=0.56 is a selected statistic. The dataset-level recommendations about PLOS and CELLS follow from the data but inherit the same fragility.\n\nNone of this kills the paper. The core claim that FKGL is a poor measure for PLS is plausible and consistent with prior work. But the magnitude of the correlations and the LM advantage are uncertain. The paper needs revision before publication: confidence intervals, cluster-robust tests, and ideally an independent human-judgment set. The abstract should be toned down.\n\nWho's this for? Anyone working on PLS evaluation, shared task organizers, and dataset builders. It deserves a serious referee. I'd send it out, but I'd expect a major revision.","headline":"The negative finding on FKGL is real but the paper overstates the LM advantage and needs confidence intervals before it should drive practice.","tokens_in":24272,"tokens_out":3519,"would_cite":true,"duration_ms":30675,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Traditional readability formulas mostly fail to match human judgments of plain-language summaries, while language models agree much better.","keywords":["plain language summarization","readability metrics","Flesch-Kincaid Grade Level","language-model evaluators","human readability judgments","scientific summarization datasets","evaluation methodology","LLM-as-judge"],"falsifier":"Gather a fresh human-evaluated readability corpus for plain-language summaries covering, say, 30+ papers in several domains, and recompute the correlations. The paper's claim would be falsified if FKGL's Pearson correlation reaches or exceeds the best LM's, or if the best LM's correlation drops below 0.3; a finding that PLOS and CELLS score high on the new corpus would also undercut the dataset conclusions.","tokens_in":23230,"feed_emoji":"📖","tokens_out":6345,"duration_ms":59931,"temperature":0.7,"pith_summary":"This paper asks whether the formulas used to grade readability in plain-language summarization research actually measure what readers experience. Surveying 18 papers in the ACL Anthology, it finds Flesch-Kincaid Grade Level (FKGL) is the standard readability check. Measured against 60 human-rated plain-language summaries, 6 of 8 traditional formulas correlate below 0.3 with human judgments; FKGL scores only 0.16. Language models asked to rate readability with their own judgment do better, with the best model reaching 0.56, and they justify scores by pointing to explanations and background knowledge rather than word length. The paper concludes that FKGL should be dropped for PLS evaluation and that two widely used 'plain language' datasets, PLOS and CELLS, score like expert-targeted abstracts under the better judge.","feed_headline":"FKGL scores 0.16; most readability formulas fail humans","feed_subtitle":"Language models rate plain-language summaries closer to human readers, and shake up which datasets count as readable.","key_machinery":"The load-bearing object is a single gold-standard benchmark: 60 summaries of 10 scientific papers rated for reading ease by crowd workers on a 1-5 scale, averaged per summary (August et al. 2024). The paper measures each candidate evaluator by its Pearson and Kendall-Tau correlation with those averages: 8 traditional formulas (FKGL, FRE, DCRS, ARI, CLI, GFI, Spache, Linsear Write) versus 5 LMs prompted to use their own judgment and explain their score. The LM reasoning text is then mined with YAKE keyword extraction to show the models appeal to explanation of terms and required background knowledge. A second mechanism is the dataset-level comparison: mean scores across 10 summarization datas","core_discovery":"The paper's central claim is that readability, for plain language summaries, is not the property that traditional readability formulas measure. Formulas built on syllable counts, sentence lengths, and word lists label a summary of acute respiratory distress syndrome as college-level even when human readers find it clear, because they cannot see that 'a very serious lung disease' defines the term. On the only available human-judgment dataset for PLS (60 summaries of 10 scientific papers, rated 1-5), FKGL correlates 0.16 (Pearson), while all five tested LMs correlate above 0.45, led by Llama 3.3 70B at 0.56; the best traditional metrics, DCRS and CLI, reach only about 0.37. The paper extends t","pith_inferences":["If the gold standard is representative, the same blind spot likely affects readability claims in health communication, legal documents, and other accessibility work that leans on grade-level formulas, because those formulas systematically reward short acronyms and penalize long explanatory words.","The 0.56 ceiling leaves room: prompt tuning, calibrated scales, or fine-tuned judge models may push LM-human agreement higher, and the near-tie among small and large LMs suggests capability is not the main constraint.","Because 60 summaries are nested in only 10 source papers, the dataset-level rankings (especially PLOS and CELLS) should be read as preliminary; a broader corpus could shift those means.","If LM judges replace formulas, their known biases and inconsistency need auditing before they become the new default; the paper itself flags this."],"forward_implications":["Papers and shared tasks that use FKGL as the primary readability check for plain-language summaries should stop: its correlation with human readability is near chance (0.16).","DCRS and CLI are the strongest of the traditional formulas but still modest; the paper recommends pairing them with LM evaluators rather than relying on either alone.","Off-the-shelf LMs of 7B-70B are workable readability judges; their written reasons can be inspected to see whether they reward definitions of technical terms and penalize missing background.","The PLOS and CELLS datasets, widely used in PLS shared tasks, are better regarded as general scientific summarization data, not plain-language data; CDSR and SciNews fit the PLS label better.","Readability metrics for PLS should be redesigned around the human notion—explanations, context, background—rather than lexical complexity."],"supporting_citations":[{"why":"Supplies the only human-judgment gold standard: 60 summaries of 10 papers rated for reading ease, used for all correlations.","marker":"August et al. (2024)"},{"why":"Defines FKGL, the most popular readability metric whose poor correlation (0.16) drives the critique.","marker":"Flesch (1952)"},{"why":"Defines DCRS, one of the two best-performing traditional metrics in the comparison.","marker":"Dale and Chall (1948)"},{"why":"Defines CLI, the other best-performing traditional metric recommended as a complement to LMs.","marker":"Coleman and Liau (1975)"},{"why":"Provides the PLOS and eLife datasets whose LM readability scores support the dataset conclusions.","marker":"Goldsack et al. (2022)"},{"why":"Provides the CELLS dataset that the LM judge ranks as low-readability, not plain language.","marker":"Guo et al. (2022)"},{"why":"Provides the CDSR dataset that the LM judge ranks as highly readable for general audiences.","marker":"Guo et al. (2021)"},{"why":"Provides the SciNews dataset, also ranked highly readable by the LM judge.","marker":"Liu et al. (2024)"},{"why":"Supplies the Williams test used to establish that LM-human correlations exceed most traditional metrics significantly.","marker":"Graham and Baldwin (2014)"},{"why":"Provides the Llama 3 family, including the best-performing evaluator Llama 3.3 70B.","marker":"Dubey et al. (2024)"}],"fun_headline_variants":["FKGL scores 0.16: readability formulas fail human judgment","Language models beat formulas on human readability judgments","Readability metrics misread plain-language summaries, LMs do better","FKGL's 0.16 correlation: why traditional metrics miss readability","LMs judge plain summaries like humans; formulas don't"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The comparison stands entirely on one gold standard: 60 human-readability ratings of summaries of 10 English science papers, collected for a different study; if those ratings are noisy or unrepresentative of plain-language readability, every correlation and dataset ranking inherits the error.","fun_headline_variants_meta":{"raw":{"variants":["FKGL scores 0.16: readability formulas fail human judgment","Language models beat formulas on human readability judgments","Readability metrics misread plain-language summaries, LMs do better","FKGL's 0.16 correlation: why traditional metrics miss readability","LMs judge plain summaries like humans; formulas don't"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000466,"raw_usage":{"total_tokens":2160,"prompt_tokens":741,"completion_tokens":1419,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":485,"completion_tokens_details":{"reasoning_tokens":1334}},"tokens_in":485,"tokens_out":1419,"duration_ms":12624,"temperature":1.0,"reasoning_tokens":1334,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T15:48:33.292524+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Gather a fresh human-evaluated readability corpus for plain-language summaries covering, say, 30+ papers in several domains, and recompute the correlations. The paper's claim would be falsified if FKGL's Pearson correlation reaches or exceeds the best LM's, or if the best LM's correlation drops below 0.3; a finding that PLOS and CELLS score high on the new corpus would also undercut the dataset conclusions.","supporting_citations":[{"cited_title":"Smith, and Katharina Reinecke","cited_arxiv_id":null,"evidence_quote":"Supplies the only human-judgment gold standard: 60 summaries of 10 papers rated for reading ease, used for all correlations."},{"cited_title":"simplification of flesch reading ease formula","cited_arxiv_id":null,"evidence_quote":"Defines FKGL, the most popular readability metric whose poor correlation (0.16) drives the critique."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines DCRS, one of the two best-performing traditional metrics in the comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines CLI, the other best-performing traditional metric recommended as a complement to LMs."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the CELLS dataset that the LM judge ranks as low-readability, not plain language."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the CDSR dataset that the LM judge ranks as highly readable for general audiences."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the SciNews dataset, also ranked highly readable by the LM judge."}],"review_version":1}