{"id":"c94e2a7d-d0ca-462b-9b56-729cc7cbf4ed","arxiv_id":"2505.01980","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Readers of simplified texts answered 3.9% more multiple-choice questions correctly than readers of originals, a gain that persisted when the text was not available while answering.","lead":"A randomized study with 4,563 participants found that people who read LLM-simplified versions of complex texts answered roughly 4% more comprehension questions correctly overall, with a 14.6% gain on medical abstracts. The result is evidence that on-demand AI rewriting could make expert information easier for general readers to use.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Simplification is confounded with elaboration: the simplified texts add definitions and context absent from the originals (Table 3), so the 3.9% accuracy gain cannot be attributed to simpler language rather than added information without a control arm that separates the two.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern I find. The paper's headline effect compares two conditions that differ on two dimensions: linguistic simplification and added explanatory content. Since the study was explicitly designed around a 'minimally lossy' approach and the examples in Table 3 show content additions, the effect cannot be cleanly attributed to simplification. This is more central than the secondary statistical concerns: the overall participant-level test is large and likely robust to reasonable clustering, and the subjectivity of the autoeval mainly affects the ancillary 'minimally lossy' claim. The elaboration confound, by contrast, determines what the study has actually shown. If the confound holds, the correct conclusion is that a particular LLM rewriting pipeline—which simplifies and elaborates—improves comprehension, not that simplification per se does. The proposed factorial experiment would settle the mechanism. Since the reader already marked the verdict CONDITIONAL and identified this weakness, no verdict change is needed.","tokens_in":27906,"tokens_out":10934,"duration_ms":128019,"concrete_test":"Run a 2×2 factorial follow-up on a random subset of the 31 texts, crossing text version (original vs LLM-simplified wording) with content (no added definitions vs the definitions/glosses that appeared in the original simplified outputs), using the same MCQs and a new sample of roughly 1000 participants per cell. The added-content factor should be constructed so that the original-wording-plus-definitions cell is matched to the simplified-plus-definitions cell. If the simplified-wording main effect is near zero once the added-content factor is included, the headline claim must be re-framed as 'LLM-generated elaboration improves comprehension', not 'text simplification'. Verify manipulation by human annotation that the no-addition cells introduce no atomic claims beyond the originals.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that LLM-based text simplification improves comprehension. The randomized comparison, however, does not isolate simplification: the 'simplified' condition changes wording and adds propositional content. In Table 3 and the Supplementary Data, the simplified texts contain glosses such as 'hepatocytes, which are liver cells', 'semaglutide, a medicine', 'NAc, a part of the brain involved in reward and motivation', 'lipocytes (fat-storing cells)', and 'fibrosis (scarring)'—none of which is present in the corresponding original. These are factual elaborations, not rewording. Because every participant in the simplified arm received both rewording and elaboration, the observed 3.9% gain (and the 14.6% PubMed gain) could be driven by supplying the meanings of technical terms rather than by making the language simpler. The 'minimally lossy' fidelity autoeval does not close this gap: its weights are subjective, it was not validated against human fidelity judgments, and the examples above show that the additions pass through it, so either it fails to detect information gain or the designers deliberately tolerated it. The paper's Discussion lists limitations but does not acknowledge this confound. A reader who wants to conclude 'simplification causally improves comprehension' needs evidence that a version with only lexical/syntactic simplification, holding content constant, produces the same benefit.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper describes a system for LLM-based text simplification developed through iterative prompt refinement with Gemini-based readability and fidelity autoevaluations, and evaluates it in a randomized study with 4,563 participants across six subject areas. Participants read original or simplified texts in open-book or closed-book conditions and answered multiple-choice comprehension questions; the authors report a 3.9% absolute improvement in MCQ accuracy for simplified texts (95% CI 1.6–6.3, p<0.05), the largest gain for PubMed, and significant improvements in self-reported confidence and perceived ease. The manuscript concludes that minimally lossy LLM text simplification can enhance comprehension of expert information.","tokens_in":28102,"tokens_out":7335,"duration_ms":66781,"significance":"The study's strengths include the large sample, random assignment, blinded MCQ writing, the open/closed-book replication, and the reporting of confidence intervals. If the effect were cleanly attributable to simplification, this would be an important contribution to the text simplification and information accessibility literature. However, as presented, the intervention conflates rewording with the addition of factual content, and the 'minimally lossy' claim is not supported by the autoeval design, so the current support for the central claim is incomplete.","major_comments":[{"comment":"The simplified texts add factual content absent from the originals; for example, Table 3 includes 'hepatocytes, which are liver cells', 'semaglutide, a medicine', and 'NAc, a part of the brain involved in reward and motivation'. Because participants in the simplified arm received both reworded text and these elaborations, the observed 3.9% overall accuracy gain (and the 14.6% PubMed gain) cannot be attributed to simplification per se. The authors should add a control arm that holds propositional content constant, or re-frame the intervention as simplification-plus-elaboration and temper the causal language in the abstract and Discussion accordingly.","section":"Table 3; Supplementary Data"},{"comment":"The system's stated objective is to simplify while 'avoiding either adding or losing information' (Figure 1), yet the fidelity autoeval only assigns error weights to 'unfactual' (weight 4) and 'off topic' (weight 1) information gains; factual on-topic additions such as the glosses in Table 3 are not penalized. This makes the 'minimally lossy' claim internally inconsistent with the evaluation procedure. The authors should either penalize all information gain in the autoeval, validate the weights against human fidelity judgments, or revise the 'minimally lossy' characterization.","section":"Figure 1; Methods: Autoeval system"},{"comment":"The accuracy analysis uses per-participant proportion correct, but responses are also nested within 31 texts and six topic areas, and participants who read the same text are correlated. Ignoring this clustering can yield overconfident standard errors and p-values for the overall and per-topic effects. The authors should fit a mixed-effects model with random intercepts for text and participant (and possibly topic) to confirm the 3.9% effect and the per-topic estimates.","section":"Methods: Statistical analysis; Results"},{"comment":"Per-topic effects are compared across six topic areas without multiple-comparison correction, while the text describes '4 or 5' of these as significant. At alpha=0.05, six tests imply a non-trivial false-positive risk; the authors should report adjusted p-values or clearly designate per-topic results as exploratory, reserving confirmatory status for the pre-specified overall comparison.","section":"Results"}],"minor_comments":[{"comment":"Consider adding a balance table of participant characteristics by experimental arm, since the targeted recruitment quotas make it easy to confirm that randomization produced comparable groups.","section":"Table 2"},{"comment":"Throughout the results, exact p-values and test statistics are omitted in favor of 'p<0.05'; reporting the regression coefficients, standard errors, and exact p-values would improve reproducibility.","section":"Results"},{"comment":"The 'simplified NASA Task Load Index' is a single bipolar Likert item rather than the full NASA-TLX; consider calling it a single-item task-load rating to avoid overstating the measure.","section":"Methods"},{"comment":"The sentence 'the same 4 of 5 topics (as the confidence results)' is ambiguous because there are six topic areas; specify which topics were significant.","section":"Discussion"},{"comment":"In the typeset version, the second row of Table 3 appears split across cells, making the original and simplified texts difficult to compare; please fix the table formatting.","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"The paper would be strengthened by a follow-up experiment or a re-analysis that isolates lexical/syntactic simplification from elaboration. Given the journal's audience, the confound in the opening claim may draw criticism; the revision should address it head-on. The authors may also consider whether the 'minimally lossy' terminology is defensible."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a read before you trust the headline effect. This is the largest randomized evaluation of LLM-based text simplification I've seen: 4,563 participants, 31 texts, six domains, MCQs written blind to the simplified outputs, and a clean open-/closed-book split. The headline result—3.9% absolute comprehension gain, p<0.05, robust to closed-book, with a 14.6% gain on PubMed—is a real empirical contribution, practically relevant for accessibility features. The confidence and perceived-ease results are coherent secondary findings. The automated prompt-refinement pipeline is a nice piece of engineering, even if the autoevals are the weakest link.\n\nThe stress-test concern holds up. The 'simplified' texts add propositional content, not just simpler wording: 'hepatocytes, which are liver cells', 'semaglutide, a medicine', 'NAc, a part of the brain involved in reward and motivation', 'fibrosis (scarring)'. These are factual elaborations. Because every participant in the simplified arm got rewording plus elaboration, you can't attribute the gain to simplification per se. The fact that the largest gains are in PubMed—where additions are most frequent—makes the confound look load-bearing. The fidelity autoeval is supposed to enforce 'minimally lossy', but these examples pass through, so either it's not sensitive to this kind of information gain or the designers intentionally tolerated it. The weights are subjectively set and not validated against human fidelity judgments. The paper's Discussion lists limitations but doesn't acknowledge this.\n\nSmaller issues: no clustering adjustment for questions nested within participants and texts, no multiple-comparison correction across six topic areas, no balance table, no code or data release. All fixable.\n\nMy take: the empirical result is likely real, but the paper oversells the mechanism. A revision should either add an arm that isolates rewording from elaboration, or honestly reframe the claim as 'LLM-generated accessible rewrites with added context improve comprehension'—which is still a useful finding. This deserves a serious referee and a revise-and-resubmit, not a desk reject.","headline":"Large RCT with a real signal, but the 'simplification' effect is confounded with added definitions; accept with heavy revision.","tokens_in":28804,"tokens_out":3228,"would_cite":true,"duration_ms":30345,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"In a randomized trial with 4,563 participants, reading LLM-simplified versions of technical texts raised comprehension multiple-choice accuracy by 3.9 percentage points over reading the originals, with the largest gain—14.6 points—on…","keywords":["text simplification","large language models","reading comprehension","randomized controlled trial","readability","cognitive load","biomedical abstracts","prompt refinement"],"falsifier":"Rerun the comprehension test with a paraphrase-only arm in which the LLM is forbidden to add any definition, gloss, or context sentence beyond the original's content. If the paraphrase-only arm shows no accuracy advantage over the original text, the paper's attribution of the benefit to simplification is falsified, and the effect is instead due to added information.","tokens_in":27642,"feed_emoji":"📖","tokens_out":5480,"duration_ms":54883,"temperature":0.7,"pith_summary":"This paper tries to establish that a large language model can rewrite expert-level texts into versions that non-expert readers genuinely understand better, not just find shorter or friendlier. In a randomized study of 4,563 adults across six subject areas, people who read the simplified versions answered 3.9% more comprehension questions correctly than people who read the originals (p<0.05), and the gain persisted when the text was not visible while answering. The benefit was largest for biomedical abstracts (14.6% absolute) and smaller but positive in law, finance, and aerospace/computer science, while simplified readers also reported higher confidence and lower perceived effort. If the effect is real, a low-cost rewrite step could raise how much of the web's expert content is usable by the general public.","feed_headline":"Simplified texts lift reader comprehension by 3.9 percent","feed_subtitle":"4,563-person randomized trial: LLM rewrites beat originals on accuracy and perceived ease, with biggest gains in medicine.","key_machinery":"The load-bearing mechanism is an automated self-refinement loop. A fast LLM produces simplified rewrites from a prompt with few-shot examples; two autoeval models score each rewrite, one for readability on a 1-10 scale and one for fidelity by decomposing the original into atomic claims and weighting completeness and entailment errors; then a refinement model revises the prompt to maximize readability minus weighted error. The loop ran 824 iterations. This is what turns \"write this more simply\" into a minimally-lossy rewrite, and the study's randomized design is what tests whether that rewrite changes reader outcomes.","core_discovery":"The paper's central discovery is that an LLM-based \"minimally lossy\" simplification pipeline produces text that improves comprehension in a randomized controlled setting. Averaged over 31 texts and 49,582 answers, simplified-text readers scored 48.2% versus 44.3% on the original (3.9% absolute gain, 95% CI 1.6 to 6.3). The gains were robust to closed-book conditions, where accuracy dropped by about 9% in both arms, and the largest effect appeared in PubMed texts (14.6%). The authors interpret this as evidence that simplification can carry the informational content of expert writing while making it accessible, and they read the persistence in closed-book conditions as a sign that comprehension and short-term retention, not just lookup, improve.","pith_inferences":["The main rival mechanism is information addition: the simplified texts in Table 3 insert definitions and context (e.g., \"hepatocytes, which are liver cells\") that the originals lack. If that extra context, not simpler syntax or vocabulary, drives the accuracy gain, then pure lexical and syntactic simplification systems would show a smaller effect.","The closed-book persistence could be due to the inserted definitions serving as retrieval cues at test time rather than deeper encoding; a delayed recall test would separate those possibilities.","The per-question pattern of larger gains where original accuracy is low may partly reflect ceiling effects on easy items; matching items on difficulty before comparing conditions would sharpen the claim.","The confidence and ease gains suggest a separate product use: simplification may increase willingness to attempt difficult material even when measurable learning is unchanged."],"forward_implications":["If the result generalizes, publishing LLM-simplified versions alongside technical documents could lift non-expert comprehension by a few percentage points on average and much more for dense biomedical material.","Because the benefit survives closed-book testing, simplified text appears to support memory and understanding, not just faster on-page reference.","The largest gains on questions where original accuracy was lowest suggest simplification helps most where readers are most likely to fail, though this pattern is a trend in per-question scatter plots rather than a separate hypothesis test.","Self-reported ease and confidence improve even in domains where accuracy gains were small, so simplification may also reduce the effort barrier to engaging with expert text.","The prompt-refinement and autoeval pipeline is presented as generalizable to other text-generation tasks that need both quality and fidelity constraints."],"supporting_citations":[{"why":"Supplies the closest prior human reading-comprehension evaluation of simplification, with 112 participants, which this study expands by an order of magnitude.","marker":"[25]"},{"why":"Prior small-N psycholinguistic comparison of simplified versus original text comprehension that this randomized design extends.","marker":"[27]"},{"why":"Prior comprehensibility study of simplified German text, used as a comparison point for target-population evaluation.","marker":"[28]"},{"why":"Prior controlled study of lexical simplification for readers with dyslexia, serving as a baseline for effects on diverse readers.","marker":"[29]"},{"why":"Randomized school experiment on LLM assistance and reading comprehension, a recent comparison point for LLM-based help.","marker":"[26]"},{"why":"Describes the LLM family the simplification model is built on.","marker":"[8]"},{"why":"Documents limitations of readability formulas, motivating the paper's learned readability autoeval.","marker":"[6,7]"},{"why":"Prior work adding definitions through retrieval augmentation, related to the elaboration observed in simplified texts.","marker":"[14]"},{"why":"Defines the NASA-TLX scale the study adapts for its cognitive-load question.","marker":"[30]"}],"fun_headline_variants":["LLM rewrites boost comprehension by 3.9%","4,563-reader trial: LLM simplifications lift comprehension 3.9%","LLM simplification helps readers: 3.9% gain, even closed book","LLM simplifications: +3.9% overall, 14.6% for medicine"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study assumes the measured comprehension gain comes from making the language simpler, but the simplified texts also add definitions and explanatory context absent from the originals, and the design never isolates rewording from added information.","fun_headline_variants_meta":{"raw":{"variants":["LLM rewrites boost comprehension by 3.9%","4,563-reader trial: LLM simplifications lift comprehension 3.9%","LLM simplification helps readers: 3.9% gain, even closed book","LLM simplifications: +3.9% overall, 14.6% for medicine"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000984,"raw_usage":{"total_tokens":4230,"prompt_tokens":1057,"completion_tokens":3173,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":673,"completion_tokens_details":{"reasoning_tokens":3085}},"tokens_in":673,"tokens_out":3173,"duration_ms":21467,"temperature":1.0,"reasoning_tokens":3085,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:04:17.294208+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the comprehension test with a paraphrase-only arm in which the LLM is forbidden to add any definition, gloss, or context sentence beyond the original's content. If the paraphrase-only arm shows no accuracy advantage over the original text, the paper's attribution of the benefit to simplification is falsified, and the effect is instead due to added information.","supporting_citations":[],"review_version":1}