{"id":"d3cb08f9-b090-4f4d-bbbf-b975f8821b74","arxiv_id":"2411.09497","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of readability metrics in legal text identifies 16 metrics across 34 studies, with Flesch-Kincaid Grade Level dominant and informed consent forms the most-studied domain.","lead":"This paper systematically reviews how readability metrics are used in legal and regulatory texts, screening over 3,500 papers and keeping 34 relevant studies. It finds that Flesch-Kincaid Grade Level is the most common metric and that informed consent forms account for most applications, while many legal domains remain unstudied.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central counts are not internally consistent: Table 2 includes non-metric studies and an uncounted RCE row, and Table 3's F-KGL citation list disagrees with Table 2, so the headline metric enumeration and frequencies are unverifiable.","rationale":"The reader's weakest assumption correctly identifies inclusion criteria as the load-bearing issue. My read strengthens that concern by locating concrete inconsistencies in the manuscript's own tables, not just an unmeasured screening judgment. The paper's contribution is a frequency/landscape description; if the underlying coding of methods and domains is wrong, the contribution loses its quantitative value. The broad conclusion that F-KGL is the most common metric and that ICF dominates is plausible and likely robust, but the reported precision ('sixteen', '73.5%') is not defensible from the current data. A conditional verdict is therefore appropriate, and my read does not move the reader's verdict; it merely supplies sharper evidence for the same condition. The proposed re-extraction audit would settle whether the errors are cosmetic or substantive.","tokens_in":13843,"tokens_out":5947,"duration_ms":51764,"concrete_test":"Independent re-extraction audit: take the 34 full texts in Table 2 and code the method(s) in each using a pre-registered rule ('a readability metric is a named quantitative formula or test producing a score from text features; qualitative complexity heuristics are excluded'). Compute (a) the unique metric list, (b) per-metric paper counts, and (c) domain percentages. Then check whether the 16-metric count, F-KGL's top rank, and the 73.5% ICF figure survive with and without rows 5 and 34 and with RCE included. If the recoded F-KGL count differs by more than one paper, or the unique metric count changes, or the ICF fraction moves outside 70-78%, the paper's central descriptive claims require correction.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is not merely hard to verify; the reported evidence contradicts it. Table 2's inclusion set contains rows that do not use readability metrics: No. 5 ('Tax law improvement', method 'consideration of rules in general terms...') and No. 34 ('Common contexts of meaning', method 'multilingual meaning problem') are legal-complexity discussions, not metric applications. Table 2 also lists RCE as a method in No. 6 and defines RCE in the Abbreviations, yet Table 3 ('Overview of linguistic measurement methods') omits RCE, so the headline 'sixteen different metrics' is not reconcilable with the review's own extraction. The F-KGL frequency data are also inconsistent: Table 3's F-KGL row cites [31], but Table 2 row 4 (ref [31]) lists Cloze Procedure, Cetinkaya and Uzun's formula, FRES, and Ateşman, not F-KGL, while Table 2 rows 22 and 28 (refs [46] and [9]) list F-KGL but are absent from Table 3's F-KGL list. Because the headline statistics are frequencies over the included set and methods, any of these errors changes the counts. The absence of an explicit coding rubric and of inter-rater agreement statistics leaves no way to know which table is authoritative.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This systematic literature review identifies and characterizes studies that apply readability or linguistic complexity metrics to legal and regulatory texts. Following a PRISMA-style search of Scopus, Web of Science, and IEEE Xplore, the authors screened 3,566 records and retained 34 studies. The paper's central descriptive claims are that sixteen different metrics appear in this literature, that the Flesch-Kincaid Grade Level is the most frequently used metric, and that 73.5% of studies concern informed consent forms. The review also reports the domain distribution, describes the most common metrics, and argues that there is no consensus on a standard readability metric for legal texts. The paper includes a quality assessment based on a modified SURE checklist and a discussion of limitations and future directions.","tokens_in":14110,"tokens_out":6663,"duration_ms":51486,"significance":"If the descriptive claims can be verified, the review fills a genuine gap by mapping which readability metrics are actually used for legal and regulatory documents and by showing that the field is concentrated on informed consent forms. The PRISMA-structured search, explicit inclusion criteria, dual-reviewer screening, and third-reviewer conflict resolution are appropriate methodological elements for a systematic review. The paper also usefully distinguishes F-KGL from FRES and discusses domain-specific limitations of common formulas. However, the paper's value as a reference work depends on the correctness and internal consistency of its tables and counts; at present, several load-bearing inconsistencies make the headline statistics unverifiable. The review is therefore potentially useful but needs substantial correction before it can be relied upon.","major_comments":[{"comment":"The inclusion criteria in §2 require studies to 'contain readability or linguistics measurements or methodology,' but Table 2 includes rows whose reported methods are not readability metrics. Row 5 ('Tax law improvement', method 'Consideration of rules in general terms, covering predictability, proportionality, consistency, compliance, administration, coordination and expression, etc.') and row 34 ('Common contexts of meaning', method 'Multilingual meaning problem') appear to be discussions of legal complexity or interpretation, not applications of a readability or linguistic measurement. Because the headline statistics ('sixteen different metrics', F-KGL as most frequent, 73.5% ICF) are frequencies over the included set, the inclusion of these studies can change the counts. Please re-apply the stated criteria to each row and either exclude these studies or explicitly justify their inclusion.","section":"§3, Table 2"},{"comment":"The claim of 'sixteen different metrics' is not reconcilable with the extraction tables. Table 2 row 6 lists RCE as a method and the Abbreviations section defines RCE, yet Table 3 omits RCE entirely. Table 3 also omits methods listed in Table 2, including 'Certain key word number percentage' (row 2), 'Cetinkaya and Uzun's formula' (row 4), 'Ateşman' (rows 4 and 22), and 'FRE' as listed in row 19. Conversely, Table 3 includes 'Grammatical intricacy and lexical density' and 'Cloze Procedure' as metric categories. Please provide a single extraction table that maps every included study to a well-defined metric taxonomy and derive the 'sixteen metrics' count from that table.","section":"Table 2 vs. Table 3"},{"comment":"The F-KGL frequency data are internally inconsistent. Table 3's F-KGL row cites [31], but Table 2 row 4 (reference [31]) reports Cloze Procedure, Cetinkaya and Uzun's formula, FRES, and Ateşman, not F-KGL. The same row omits [34] (Table 2 row 7), [9] (Table 2 row 28), and [46] (Table 2 row 22), all of which list F-KGL in Table 2. Because the central claim that F-KGL is the most frequently used metric is a count over these citations, the F-KGL frequency must be recomputed from a corrected, internally consistent version of Table 2 and Table 3.","section":"Table 3, F-KGL row"},{"comment":"The domain counts in Table 4 do not sum to the 34 included studies: the table reports 25 ICF + 4 Tax + 2 Legal + 1 Finance + 1 Medical Regulations = 33, and the percentages sum to 96.9%. Table 2 lists three studies in the legal/other legal domain (rows 20, 33, and 34), not two, and uses both 'legal' and 'Other legal' category labels. Please reconcile the domain coding, the category labels, and the totals, and recompute the percentages from the corrected counts.","section":"Table 4"},{"comment":"The quality assessment is not verifiable as presented. Section 2 states that 'The quality score of each paper is shown in Appendix 2,' but no appendix is present in the manuscript, and the statement that all included studies score full marks on the first three items is unsupported by any per-study data. In addition, while the paper reports that two independent reviewers screened and extracted data, it reports no measure of inter-rater agreement (for example, Cohen's kappa) for screening, extraction, or quality scoring. Without these materials, the reliability of the inclusion decisions and the quality claims cannot be assessed.","section":"§2, Quality Assessment"}],"minor_comments":[{"comment":"The manuscript references a PRISMA flowchart in Figure 1, but no figure content appears in the text provided; please ensure the figure is actually included in the submitted version.","section":"Figure 1"},{"comment":"The Abbreviations section lists 'Informed Consent Forms = ICF' twice, and 'Flesh readability ease score' should read 'Flesch readability ease score.'","section":"Abbreviations"},{"comment":"The Conclusions section refers to 'F-KCL' rather than F-KGL; please correct the acronym.","section":"Conclusions"},{"comment":"There are numerous language issues, such as 'According to researcher,' 'Legal law field language, which includes complex sentences archaic or apply large amount of words,' and 'what are been deemed as simple and useful regulation is always debatable.' A careful language edit is needed.","section":"§1.1"},{"comment":"Table 3 contains encoding issues and inconsistent naming, including 'Ate¸sman' and 'Dale-Chale,' and it is not always clear whether 'FRE' (Table 2 row 19) and 'GFOG' (Table 2 row 13) are intended to be the same as FRES and Gunning Fog; please standardize the metric names.","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"The central descriptive claims are plausible in direction, but the manuscript currently contains several load-bearing inconsistencies that prevent verification of the headline statistics. The missing quality-assessment appendix and the absence of inter-rater agreement measures are also non-trivial for a systematic review. These issues are fixable by re-extracting the data into a single consistent table and recomputing all counts; therefore I recommend major revision rather than rejection. The self-citation of reference [9] is not itself a problem, but the corrected counts should be checked to ensure the conclusions do not depend on it."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is the first systematic review I know of that is scoped specifically to readability metrics in legal and regulatory texts, and that alone makes it a useful map. The search strategy is transparent, the tabulated inventory of methods and domains is practical, and the discussion of why classic readability formulas are a poor fit for legal language is honest and well-grounded. If you need a quick sense of how informed consent forms dominate this literature and how Flesch-Kincaid is the default metric, the paper delivers it.\n\nBut the headline numbers don't hold up under inspection. Table 2 includes at least two studies that are not using readability metrics at all: No. 5 is a tax-law simplification strategy discussion, and No. 34 is a semiotic analysis of meaning. That alone puts the \"34 studies\" and \"sixteen metrics\" figures in doubt. The internal consistency is worse: RCE appears in Table 2 (row 6) and in the abbreviations list, but is omitted from Table 3's method overview, so the metric count is off. And Table 3's F-KGL reference list disagrees with Table 2: refs [46] and [9] apply F-KGL according to Table 2 but are absent from Table 3's F-KGL row. If the central claim is a frequency distribution, these are not cosmetic errors.\n\nThe reportability gaps are milder but still relevant: the PRISMA flowchart is referenced but not included in the provided text, the quality assessment scores are mentioned but not shown, and no inter-rater agreement is reported. All of this is fixable with better reporting.\n\nI would send this to peer review with major revisions. The topic deserves a serious referee, and the qualitative synthesis—especially the point that consent forms dominate while other legal domains are neglected—is likely correct. But as it stands, the numeric claims are not independently auditable, and the paper should not be cited for its counts until the tables are reconciled.\n\nFor you: worth a skim if you work in legal NLP or readability. I wouldn't build anything on its aggregate numbers yet.","headline":"A useful first map of a scattered field, but the numbers in the tables are not internally consistent enough to be cited until the review is revised.","tokens_in":14594,"tokens_out":3891,"would_cite":false,"duration_ms":33857,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A systematic review finds 16 readability metrics used in legal texts, with Flesch-Kincaid Grade Level the most common.","keywords":["readability metrics","legal text","systematic review","Flesch-Kincaid Grade Level","informed consent forms","regulatory complexity","plain language","linguistic complexity"],"falsifier":"Re-run the three-database search with the same keyword combination and inclusion criteria, then have two independent teams apply the criteria to the full texts; if the reproduced set does not include the same 34 studies, or if judges disagree about entries such as \"tax law improvement\" or \"common contexts of meaning\" counting as readability metrics, then the 16-metric count and the claim that F-KGL is most frequent would need revision.","tokens_in":13648,"feed_emoji":"⚖️","tokens_out":8174,"duration_ms":74185,"temperature":0.7,"pith_summary":"This systematic review sets out to establish, for the first time, which readability metrics researchers actually apply to legal and regulatory texts, and where they apply them. Screening 3,566 records from three databases, the authors narrow the field to 34 studies and find 16 distinct metrics. Their central descriptive result is that the Flesch-Kincaid Grade Level is the most frequently used metric, followed by the Flesch Reading Ease score and SMOG, and that 73.5% of the studies examine informed consent forms. The paper concludes that legal readability measurement is concentrated in a narrow medical niche and lacks any consensus standard, which matters because overly complex laws are hard for the public to engage with and because NLP tools for legal text need a common readability baseline.","feed_headline":"Flesch-Kincaid tops 16 metrics used on legal texts","feed_subtitle":"Review of 34 studies finds consent forms dominate and no single readability standard exists for law.","key_machinery":"The load-bearing object is the inventory itself: a table of 34 studies, each tagged with the method it applied and the legal subdomain it studied, assembled through a structured literature-screening protocol with two independent reviewers and a third for disagreements. The named formulas that carry the counting argument include the Flesch-Kincaid Grade Level (F-KGL), which converts average words per sentence and syllables per word into a U.S. school-grade score, the Flesch Reading Ease score (FRES), a 0-100 scale with higher meaning easier, and SMOG, which counts polysyllabic words in sampled sentences. These formulas are the units being counted, and the inventory is what transforms them into the frequency claims.","core_discovery":"The paper's claim is descriptive: the published literature on readability measurement for legal and regulatory texts, as of February 2023, consists of 34 eligible studies using 16 different metrics. The Flesch-Kincaid Grade Level appears in the largest share of those studies, with the Flesch Reading Ease score and SMOG the next most common. 25 of the 34 studies, 73.5%, fall in the medical domain, almost all of them about informed consent forms; tax, general legal, financial, and medical-regulatory texts account for the rest. The authors further claim that this distribution shows no field-wide consensus on which metric fits legal text, and that most studies adopt a metric for convenience rather than demonstrated suitability.","pith_inferences":["The concentration in consent-form studies is likely a response to ethical-review pressure to prove participants understand what they sign; extending the same measurement habit to financial or tax disclosure could quickly broaden the evidence base.","A directly testable extension is to apply the same battery of 16 metrics to matched corpora of statutes, regulations, and consent forms; if the paper's domain imbalance is meaningful, metric agreement and grade-level outcomes should diverge across those text types.","The paper's discussion of Dale-Chall-style word lists suggests a concrete next step: build a legal-domain word list and compare it head-to-head against SMOG and F-KGL on the same documents."],"forward_implications":["If the counts are right, any attempt to standardize legal readability assessment starts from a field with no agreed metric; the most common choice is F-KGL, but most studies use different formulas and their scores are not directly comparable.","Because 73.5% of the evidence sits in informed consent forms, conclusions about legal readability are really conclusions about medical consent documents; other legal domains are empirically open territory.","Researchers applying NLP or machine learning to legal text have no common readability baseline to compare models or corpora, so progress on legal-text simplification will be hard to measure against a shared scale.","The prevalence of education-oriented formulas such as F-KGL and SMOG implies that domain-specific legal vocabulary is being scored by proxies like syllable count, which the paper argues misses semantics, repetition, and document structure."],"supporting_citations":[{"why":"Supplies the structured screening protocol that organizes the search, deduplication, and inclusion steps.","marker":"[27]"},{"why":"Provides the quality-appraisal checklist used to score each included paper.","marker":"[28]"},{"why":"Is one of the 34 included studies, contributing the financial domain and the LIX metric.","marker":"[10]"},{"why":"Contributes the tax-law domain and helps anchor the Flesch-Kincaid frequency count.","marker":"[12]"},{"why":"Represents the informed-consent-form stream that dominates the 73.5% figure.","marker":"[8]"},{"why":"Is one of the F-KGL-using consent-form studies that feeds the most-frequent-metric claim.","marker":"[36]"},{"why":"Contributes the medical-regulations domain and links readability to Halstead-style complexity measures.","marker":"[9]"},{"why":"Contributes the Russian legal-text hybrid model and the adapted-metric category.","marker":"[44]"},{"why":"Supplies the finance-regulation adaptation of Halstead measures that the discussion draws on for future metric designs.","marker":"[71]"}],"fun_headline_variants":["Legal readability metrics: 16 tools, no consensus","Flesch-Kincaid tops 16 metrics in legal readability review","Consent forms skew legal readability metric research","Systematic review maps readability metrics for legal texts","Law lacks readable-text standard: 16 metrics, 34 studies"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central counts rest on the reviewers' judgment that each of the 34 included papers really applies a readability metric to legal text, and the review does not report an agreement measure for that screening judgment.","fun_headline_variants_meta":{"raw":{"variants":["Legal readability metrics: 16 tools, no consensus","Flesch-Kincaid tops 16 metrics in legal readability review","Consent forms skew legal readability metric research","Systematic review maps readability metrics for legal texts","Law lacks readable-text standard: 16 metrics, 34 studies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000323,"raw_usage":{"total_tokens":1792,"prompt_tokens":901,"completion_tokens":891,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":517,"completion_tokens_details":{"reasoning_tokens":811}},"tokens_in":517,"tokens_out":891,"duration_ms":7671,"temperature":1.0,"reasoning_tokens":811,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T20:34:23.044849+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the three-database search with the same keyword combination and inclusion criteria, then have two independent teams apply the criteria to the full texts; if the reproduced set does not include the same 34 studies, or if judges disagree about entries such as \"tax law improvement\" or \"common contexts of meaning\" counting as readability metrics, then the 16-metric count and the claim that F-KGL is most frequent would need revision.","supporting_citations":[{"cited_title":"Preferred reporting items for systematic reviews and meta-analyses: the prisma statement,","cited_arxiv_id":null,"evidence_quote":"Supplies the structured screening protocol that organizes the search, deduplication, and inclusion steps."},{"cited_title":"Questions to assist with the critical appraisal of qualitative studies,","cited_arxiv_id":null,"evidence_quote":"Provides the quality-appraisal checklist used to score each included paper."},{"cited_title":"Voluntary disclosure and complexity of reporting in egypt: the roles of profitability and earnings management,","cited_arxiv_id":null,"evidence_quote":"Is one of the 34 included studies, contributing the financial domain and the LIX metric."},{"cited_title":"The readability of australia’s taxation laws and supple- mentary materials: an empirical investigation,","cited_arxiv_id":null,"evidence_quote":"Contributes the tax-law domain and helps anchor the Flesch-Kincaid frequency count."},{"cited_title":"Readability standards for informed-consent forms as compared with actual readability,","cited_arxiv_id":null,"evidence_quote":"Represents the informed-consent-form stream that dominates the 73.5% figure."},{"cited_title":"Readability of foot and ankle consent forms in queensland,","cited_arxiv_id":null,"evidence_quote":"Is one of the F-KGL-using consent-form studies that feeds the most-frequent-metric claim."},{"cited_title":"A hybrid model of complexity estimation: Evidence from russian legal texts,","cited_arxiv_id":null,"evidence_quote":"Contributes the Russian legal-text hybrid model and the adapted-metric category."},{"cited_title":"Measuring regulatory complexity,","cited_arxiv_id":null,"evidence_quote":"Supplies the finance-regulation adaptation of Halstead measures that the discussion draws on for future metric designs."}],"review_version":1}