{"id":"ded34c8d-d1e3-4ab8-a8f7-138e8ea982ee","arxiv_id":"2501.16865","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A three-role LLM loop improves readability-formula scores of automatically generated science journalism without fine-tuning the base models.","lead":"JRE-L lets three language models act as journalist, reader, and editor, rewriting scientific abstracts into popular-science articles through repeated feedback rounds. On three datasets the generated articles beat GPT-4 on readability formulas, but human raters found no statistically significant readability advantage over GPT-4.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim that JRE-L outperforms GPT-4 on accessibility rests on surface readability formulas; the paper's own human evaluation shows no significant difference, making formula validity the load-bearing premise.","rationale":"The paper is otherwise coherent: the framework is clearly described, the ablations support the contribution of each module, and the automatic readability improvements are consistent across datasets. However, the central claim that small open-source LLMs produce articles 'more accessible' than GPT-4 is only as strong as the metrics used to measure accessibility. The reader's weakest-assumption analysis correctly identifies that the formulas are surface-level and not validated against actual comprehension. My stress-test pass did not find a more load-bearing flaw: the human-evaluation null result is the decisive internal evidence that the abstract overstates the finding. The post hoc selection of iteration 3 is a secondary concern, since Figure 4 shows the scores stabilize by then, but it would be worth disclosing in any revision. The proposed comprehension study is the direct check that would settle the matter; until it is run, the conditional verdict remains appropriate.","tokens_in":18719,"tokens_out":2483,"duration_ms":25002,"concrete_test":"Conduct a preregistered comprehension study with at least 30 general-audience participants who lack domain background. For a matched sample of 20–50 papers from SCITech, eLife, and PLOS, present JRE-L and GPT-4 outputs blind, then measure (a) accuracy on multiple-choice questions about key facts in the original paper, (b) self-reported ease of understanding, and (c) reading time. Also compute the CLI/FKGL/DCRS deltas for the same outputs and test whether those deltas predict comprehension gains. If JRE-L does not significantly beat GPT-4 on comprehension, or if formula improvements do not correlate with comprehension, the central claim loses its footing.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract claims JRE-L generates articles 'more accessible' than GPT-4, and the only statistically significant evidence for this is Table 1, where JRE-L achieves lower CLI, FKGL, and DCRS scores. These formulas count characters, syllables, sentence length, and presence of words on a common-word list; they do not measure whether a non-specialist comprehends the article. The paper's own human evaluation (Table 2, Section 4.3) shows no statistically significant difference between JRE-L and GPT-4 on any dimension, including the directly labeled 'Readability' dimension, and the Limitations section concedes that 'statistical measures may miss semantic information.' Shorter sentences and substituting technical terms with everyday synonyms can improve formula scores even when the content becomes vague, under-specified, or harder to follow in context. Because the framework's iterative reader/editor loop plausibly selects for these surface properties, the headline quantitative advantage over GPT-4 may be an artifact of the chosen metrics rather than a real gain in accessibility to the general audience.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes JRE-L, a framework in which three open-source LLMs (a 7B journalist, a 1.8B reader, and a 7B editor) collaborate in an iterative write-read-feedback-revise loop to turn scientific paper abstracts into popular-science articles. The journalist drafts an article, the smaller reader LLM takes notes to expose comprehension failures, and the editor evaluates the notes and gives revision advice, which the journalist incorporates over multiple iterations. The authors evaluate on three existing corpora (SCITech, eLife, PLOS) using the Coleman-Liau Index, Flesch-Kincaid Grade Level, and Dale-Chall Readability Score, plus a four-annotator human study covering readability, information conveyance, authenticity, and interestingness. They report lower readability-formula scores for JRE-L than for all baselines, including GPT-4, and report human ratings close to those of GPT-4. They also present ablations, iteration trends, and case studies, and release their code.","tokens_in":18898,"tokens_out":2891,"duration_ms":25239,"significance":"If the central claim holds, the paper is a useful empirical contribution: it shows that a small, fully open-source ensemble of LLMs can match or beat a much larger closed-source model on readability metrics for popular-science writing, and it provides a reproducible recipe (roles, prompts, iteration) plus extensive ablations. The framework is simple and the code release enables follow-up. The paper also documents failed design attempts, which is informative. The main quantitative advantage, however, is only as strong as the validity of the readability formulas, and the human evaluation does not independently confirm an advantage over GPT-4; this limits the significance of the headline 'more accessible' claim to the narrow metric level.","major_comments":[{"comment":"The abstract claims JRE-L generates articles 'more accessible' than GPT-4, but the only statistically significant evidence for this is the automatic readability-formula scores (CLI, FKGL, DCRS) in Table 1. These formulas count characters, syllables, sentence length, and word frequency (Section 4.1) and do not measure comprehension by a general audience. The human evaluation in Table 2 shows no statistically significant difference between JRE-L and GPT-4 on the 'Readability' dimension (3.95 vs 3.80 within field; 3.65 vs 3.40 outside field). The Limitations section itself concedes that 'these statistical measures may miss semantic information.' Since the framework's iterative loop plausibly optimizes exactly the surface features the formulas reward, the headline advantage over GPT-4 may be an artifact of the metrics rather than a genuine gain in accessibility. I ask the authors to either (a) provide a human or model-based comprehension measure that validates the formula difference, or (b) revise the abstract's wording to state that JRE-L improves readability-formula scores.","section":"Section 4.2 (Table 1) and Section 4.3 (Table 2)"},{"comment":"The final reported output is selected post hoc: the authors state, 'We iterate five times and empirically select the output from the third iteration as the final result.' This selection is made on the test set, and Table 1 reports the scores of this chosen iteration without accounting for the selection process. The statistical significance marks (†, ††) therefore compare the best-of-five-iterations output of JRE-L against single-shot outputs of baselines, which is an unfair comparison and inflates the apparent advantage. The paper should either fix the iteration count and the chosen index in advance (e.g., by a validation split), apply a multiple-comparison correction for the five iterations, or report the full trajectory for all baselines with a pre-specified stopping rule.","section":"Appendix C (Hyperparameters)"}],"minor_comments":[{"comment":"The caption spells 'Iteraction' instead of 'Iteration.'","section":"Figure 4 caption"},{"comment":"The label 'Sugesstions for article revision' contains a typo; it should be 'Suggestions for article revision.'","section":"Figure 2"},{"comment":"The text says 'As depicted in Table 7' when referring to ablation results, but the ablation results appear in Table 3; the detailed results are in Appendix F, Table 9. Please correct the cross-reference.","section":"Section 5.2"},{"comment":"The human evaluation uses only four participants and reports Krippendorff's alpha of 0.52, which the authors themselves describe as 'slightly lower' than prior work's 0.57. With such low inter-annotator agreement and a small sample, the absence of a significant difference between JRE-L and GPT-4 on readability should be interpreted cautiously; this point should be acknowledged in the main text, not only implicitly through the table.","section":"Section 4.3"},{"comment":"The statement that 'all approaches in our comparison study share the same set of hyperparameters setting' is misleading because the JRE-L framework has an iteration count and a selected iteration index, while single-LLM baselines do not. Please clarify which hyperparameters are shared and which are specific to the iterative framework.","section":"Appendix C"},{"comment":"The 'Failed Attempts' appendix is useful, but it would be even more helpful to report the quantitative evidence for why the alternative formulations (reflection, reading comprehension, single-prompt) underperformed, rather than only describing them qualitatively.","section":"Appendix B"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable empirical study for an applied NLP venue, and the code release is a strength. The main reservation concerns the validity of the readability formulas as a measure of 'accessibility' and the post hoc iteration selection; both issues are fixable within the manuscript's scope (e.g., by adding a comprehension-based validation, pre-registering the iteration stop, or softening the abstract's wording). I would therefore not reject the paper outright, but these load-bearing points must be addressed before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take on arXiv:2501.16865. The core idea is better than the framing. Using a deliberately smaller, weaker LLM as a reader whose note-taking quality degrades when the text is harder is a clever way to get a cheap readability signal. The journalist-reader-editor loop is a new combination, and the ablations show that both the reader and the editor modules contribute beyond single-model prompting. The experiments are reasonably thorough: three datasets, multiple baselines, and statistically significant improvements on the automatic metrics. I also give the authors credit for documenting failed attempts (Appendix B) and for stating limitations plainly. The code link is a plus, though there's no commit hash or reproduction script.\n\nThe soft spot is the one the stress-test note flags, and it's load-bearing. The abstract says JRE-L generates articles \"more accessible\" than GPT-4. That claim is supported only by CLI, FKGL, and DCRS scores — formulas that count characters, syllables, sentence length, and presence on a common-word list. They do not measure comprehension. The paper's own human evaluation shows no statistically significant difference from GPT-4 on any dimension, including the dimension literally labeled 'Readability.' That's a direct contradiction of the abstract's strongest phrasing. The iteration count and the choice to report the third iteration out of five are also post hoc; there's no held-out selection criterion.\n\nNone of this makes the method worthless. The framework does beat single small models and prior multi-agent baselines on the formulas, and the trend over iterations is clean. But the claim that this translates to real accessibility gains for general readers is not established. The formulas may select for shorter sentences and simpler synonyms even when the content becomes vague or harder to follow in context — the stress-test note makes that point well, and the case study actually shows the reader's notes getting more detailed, which is suggestive but not proof of comprehension.\n\nWho should read this: people working on automatic science journalism, LLM agents for writing, or evaluation of generated text. It deserves a serious referee, because the core mechanism is testable and the paper is transparent enough to engage with. A referee should push for either a comprehension-based human study or a substantially softened claim. As it stands, I'd treat the GPT-4 comparison as \"competitive, not superior\" and focus on the cheap-reader-probe idea, which is the real contribution.","headline":"The reader-as-probe loop is a genuinely neat idea, but the paper's headline claim that small open models beat GPT-4 on accessibility rests entirely on surface readability formulas, and the paper's own human evaluation does not support it.","tokens_in":19411,"tokens_out":1835,"would_cite":true,"duration_ms":19607,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a loop of three small open-source LLMs—journalist, reader, and editor—produces science articles with lower readability-formula scores than GPT-4, across three benchmark corpora.","keywords":["science journalism","large language models","multi-agent collaboration","readability","lay summarization","iterative revision","general audience","automatic summarization"],"falsifier":"Give non-expert readers a comprehension test on paired articles generated by JRE-L and GPT-4 without revealing which is which; if the JRE-L articles are not understood better despite lower readability-formula scores, the paper's central measure of accessibility fails. The paper's own human evaluation already hints at this, since raters found no statistically significant difference in readability between JRE-L and GPT-4.","tokens_in":18519,"feed_emoji":"📰","tokens_out":5581,"duration_ms":42723,"temperature":0.7,"pith_summary":"The paper proposes JRE-L, a writing loop in which three open-source language models take the roles of journalist, general reader, and editor, and argues that this loop makes automatically generated popular science articles more readable than existing methods. The concrete claim is that two 7B-parameter models and one 1.8B-parameter model, cooperating without any training, produce articles whose Coleman-Liau, Flesch-Kincaid, and Dale-Chall scores are lower than those of articles prompted from GPT-4 and other strong baselines, with the gap statistically significant at p<0.01. The mechanism is that a deliberately weak reader model takes notes on the journalist's draft, and its note quality exposes writing problems that the editor then turns into concrete revision advice. If true, the result matters because it suggests that resource-light, parameter-free role-play can match or beat much larger closed models on an accessibility metric for science communication.","feed_headline":"Small open-source LLMs beat GPT-4 on science-writing readability","feed_subtitle":"Two 7B and one 1.8B models playing journalist, reader, and editor outscore GPT-4 on readability formulas.","key_machinery":"The load-bearing object is the JRE-L loop itself: an iterative, training-free cycle in which a 7B journalist model, a 1.8B reader model, and a 7B editor model pass prompts in a write-read-assess-revise sequence. The reader's note-taking task is the distinctive mechanism; the authors deliberately choose a smaller, weaker model so that unclear writing produces incomplete or confused notes, converting article readability into measurable signal through error propagation. The editor reads the paper, the article, and the notes, then issues specific revision advice; the journalist incorporates the advice and the cycle repeats, with readability gains concentrated in the first iterations and plateauing by the third.","core_discovery":"JRE-L's central claim is that the three-role collaboration loop—journalist writes, a smaller reader model reads and takes notes, an editor evaluates the notes and issues revision suggestions, journalist revises—produces popular science articles with lower automatic readability scores (CLI, FKGL, DCRS) than single-LLM prompting, including GPT-4, across three benchmark corpora (SCITech, eLife, PLOS). The paper further claims that the reader LLM is the load-bearing element: because the 1.8B reader has weaker comprehension, its notes become comprehensive only when the article explains technical terms clearly, so note quality propagates readability problems back to the editor. In human evaluation, JRE-L received numerically higher readability ratings than GPT-4 without a statistically significant difference on any dimension; the paper interprets this as achieving GPT-4-level accessibility with far smaller models.","pith_inferences":["The paper's headline claim is about readability formulas, not verified comprehension; a fair follow-up would test whether lower CLI, FKGL, and DCRS actually translate into better understanding by lay readers.","The reader-as-error-propagation idea could extend beyond journalism to any text simplification task where the target audience's likely misunderstandings can be simulated by a weaker model.","If human raters do not perceive a significant readability difference from GPT-4, the practical advantage of JRE-L may lie in cost and openness rather than in raw accessibility.","The framework's editor depends on the reader's notes; replacing note-taking with a direct readability judgment failed in pilot experiments, which suggests the value comes from the proxy reader rather than from LLM self-assessment."],"forward_implications":["Science journalism pipelines could achieve GPT-4-level readability using open-source 7B and 1.8B models, cutting cost and avoiding closed APIs.","The framework requires no fine-tuning or parameter updates, so it can be applied to new domains by changing the input paper alone.","Ablation shows that removing the reader, the editor, or the collaboration degrades readability scores, implying all three roles contribute to the improvement.","The loop generalizes across model families (Qwen and LLaMA) and across scientific domains (computer science, biomedical, and life sciences).","Readability gains saturate around the third iteration, suggesting a practical stopping rule for the revision cycle."],"supporting_citations":[{"why":"Supplies the Coleman-Liau Index, one of the three readability formulas used as the paper's main automatic measure.","marker":"Coleman and Liau (1975)"},{"why":"Supplies the Flesch-Kincaid Grade Level formula used in automatic evaluation.","marker":"Kincaid et al. (1975)"},{"why":"Supplies the Dale-Chall Readability Score and its familiar-words list used in automatic evaluation.","marker":"Dale and Chall (1948)"},{"why":"Provides the eLife and PLOS corpora for lay summarisation and the fine-tuned BART baseline.","marker":"Goldsack et al. (2022)"},{"why":"Provides the SCITech corpus, the discourse-structure ASJ baseline, and the human-evaluation protocol and agreement values compared against.","marker":"Cardenas et al. (2023)"},{"why":"Introduces the automatic science journalism task and the sequence-to-sequence approach the paper extends.","marker":"Dangovski et al. (2021)"},{"why":"Supports the premise that LLMs can act as evaluators, which underlies the editor role.","marker":"Zheng et al. (2024a)"},{"why":"Provides ChatDev, one of the multi-LLM collaboration frameworks the paper adapts and compares against.","marker":"Qian et al. (2024)"},{"why":"Provides CollabStory, the other multi-LLM collaboration baseline adapted for comparison.","marker":"Venkatraman et al. (2024)"},{"why":"Documents GPT-4, the strongest single-LLM baseline whose readability scores JRE-L claims to beat.","marker":"Achiam et al. (2023)"}],"fun_headline_variants":["Small LLM trio beats GPT-4 at popular science writing","JRE-L: Two 7B and one 1.8B LLM beat GPT-4 readability","Three open-source LLMs simulate editor to beat GPT-4 on readability","Tiny LLM journalist-reader-editor loop outperforms GPT-4","LLM trio achieves GPT-4-level readability with far fewer parameters"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's quantitative claim rests on treating the Coleman-Liau Index, Flesch-Kincaid Grade Level, and Dale-Chall Readability Score as valid measures of how accessible an article is to the general public.","fun_headline_variants_meta":{"raw":{"variants":["Small LLM trio beats GPT-4 at popular science writing","JRE-L: Two 7B and one 1.8B LLM beat GPT-4 readability","Three open-source LLMs simulate editor to beat GPT-4 on readability","Tiny LLM journalist-reader-editor loop outperforms GPT-4","LLM trio achieves GPT-4-level readability with far fewer parameters"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000355,"raw_usage":{"total_tokens":1905,"prompt_tokens":898,"completion_tokens":1007,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":514,"completion_tokens_details":{"reasoning_tokens":905}},"tokens_in":514,"tokens_out":1007,"duration_ms":9104,"temperature":1.0,"reasoning_tokens":905,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T10:04:13.875213+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give non-expert readers a comprehension test on paired articles generated by JRE-L and GPT-4 without revealing which is which; if the JRE-L articles are not understood better despite lower readability-formula scores, the paper's central measure of accessibility fails. The paper's own human evaluation already hints at this, since raters found no statistically significant difference in readability between JRE-L and GPT-4.","supporting_citations":[],"review_version":1}