{"id":"0ac2a9b1-62ce-4264-82c4-28757ffa1183","arxiv_id":"2507.14096","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Across two TREC shared-task years, top LLM systems matched human writers on factual accuracy and completeness but not on simplicity or brevity, while common automatic metrics correlated poorly with manual judgments.","lead":"This paper reports the outcomes of a two-year shared task in which research teams used language models to rewrite biomedical abstracts into plain language for patients. The best systems matched human writers on factual accuracy and completeness, but simplicity and brevity lagged, and standard automatic evaluation metrics did not track human judgments well.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'not simplicity or brevity' half of the headline claim rests on manual axes with poor reliability (SEN alpha=0.1748, FLU alpha=-0.0255) and on a brevity axis with no human baseline; the accuracy/completeness half is robust.","rationale":"The reader identified manual evaluation reliability as the key risk, and I agree it is the central issue, but it is not uniform across axes. The paper's own Appendix F and §4.2 show the unreliability is concentrated in SEN and FLU—exactly the axes feeding SIM, which underlies the 'not simplicity' conclusion. The accuracy/completeness comparison (PLABA_1 ACC 97.65 vs Manual_avg 96.21) uses high-agreement axes and is a real, robust finding. Therefore the single most load-bearing concern is that the negative half of the headline claim is not established: the SIM gap is within annotation noise on unreliable axes, and BRV has no human baseline at all. This does not invalidate the paper's contributions—the shared-task infrastructure, the expert accuracy judgments, and the metric-correlation analysis (which is conservative under noisy labels) remain valuable—but it means the abstract overstates one component of the headline finding. Since the reader already classified the paper as CONDITIONAL and this concern reinforces that classification (the authors should qualify or remove the 'not simplicity or brevity' claim and add confidence intervals or sensitivity analysis for SIM), the verdict remains UNCHANGED. My agreement with the reader is partial: I concur that manual evaluation reliability is the load-bearing worry, but I localize it to the negative half of the claim and note that the positive half is well supported by high-agreement axes.","tokens_in":29429,"tokens_out":8431,"duration_ms":531286,"concrete_test":"Recompute the Round 2 comparison in Table 3 replacing SIM with two restricted composites: (a) SIM_noFLU = mean(SEN, TRM, TAC) and (b) SIM_reliable = mean(TRM, TAC), for PLABA_1 and Manual_avg. Then construct 95% bootstrap confidence intervals for each PLABA_1-minus-Manual_avg difference by resampling sentences (or abstracts) from the existing sentence-level judgments. If either interval contains zero, the 'not simplicity' claim is not robust to the unreliable SEN/FLU axes.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim's negative half—'top-performing models rivaled human levels of factual accuracy and completeness, but not simplicity or brevity'—is not adequately supported. In Table 3 (Round 2), PLABA_1 SIM=91.51 vs Manual_avg=93.56, a 2.05-point gap on an interpolated 0-100 scale. SIM is the average of SEN, TRM, TAC, and FLU (Table 1). Appendix F reports sentence-level Krippendorff's alpha for SEN = 0.1748 and FLU = -0.0255. The paper itself attributes FLU's negative alpha to near-perfect scores, and in §4.2 says SEN's low agreement came from annotators conflating sentence length with brevity, 'which caused confusion among annotators and contributed to poor inter-annotator agreement.' With two of four sub-axes unreliable, the 2.05-point SIM gap could be annotation noise; the paper reports no confidence intervals, significance tests, or sensitivity analyses. Separately, 'not brevity' has no human baseline: BRV was introduced in 2024, when no manually written reference adaptations were created (§5.2), so no human BRV score exists for comparison. In contrast, the positive half of the claim is on firmer ground: sentence-level alpha for COM and FTH is 0.8833 and 0.8148, respectively. Thus the fragile part of the headline is precisely the part the paper itself flags in Appendix F and §4.2, yet the abstract states it as a settled finding.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports on the Plain Language Adaptation of Biomedical Abstracts (PLABA) shared task at TREC 2023 and 2024. Task 1 asked teams to rewrite biomedical abstracts sentence-by-sentence into plain language; Task 2 asked teams to identify, classify, and generate replacements for difficult terms. The paper describes the data, the manual and automatic evaluation protocols, the participating systems, and the results, and it draws conclusions about the state of LLM-based plain language adaptation, the reliability of automatic metrics, and the risk of hallucinations.","tokens_in":29652,"tokens_out":6728,"duration_ms":73745,"significance":"If the reported findings hold, the paper is a substantial empirical contribution to evaluation methodology for biomedical text simplification. Its concrete strengths include a two-year shared task with twelve participating teams and thirty-eight runs, a four-fold professionally written reference set for the 2023 Task 1 test set, extensive manual evaluation of system outputs, and transparent appendices documenting inter-annotator agreement, the sentence-alignment pipeline, system metadata, and hallucination examples. The correlation analysis in Section 5.1.4 provides useful evidence that n-gram-based metrics, including SARI and SAMSA, correlate poorly with manual judgments, while BERTScore correlates more strongly. The paper's headline claims are falsifiable and are presented with enough table-level detail for the reader to inspect the supporting evidence.","major_comments":[{"comment":"The headline claim 'but not brevity' has no human baseline. Brevity (BRV) was introduced only in the 2024 offering (§4.2), and §3.1.3 states that no manually written reference adaptations were created for the 2024 test set; Table 4 therefore lists BRV scores only for system submissions, with no Manual_avg row. Without a human reference score or a pre-specified absolute criterion, the data cannot support the comparative claim that systems did not rival humans on brevity.","section":"Abstract; §5.2, Table 4"},{"comment":"The 'not simplicity' half of the headline rests on the SIM composite, which averages SEN, TRM, TAC, and FLU. Appendix F reports sentence-level Krippendorff's alpha of 0.1748 for SEN and -0.0255 for FLU, and §4.2 explains that SEN was conflated with brevity and that FLU had near-perfect scores. The 2.05-point SIM gap between PLABA_1 (91.51) and Manual_avg (93.56) in Table 3 is therefore within the range that annotation noise could explain, yet no confidence intervals, significance tests, or sensitivity analyses are provided. The authors should either report reliability-aware comparisons or explicitly qualify the simplicity claim.","section":"Appendix F; Table 3; §4.2"},{"comment":"The positive half of the headline, that top-performing models rivaled human levels of factual accuracy and completeness, is supported by a single run: PLABA_1 in Round 2 of TREC 2023 exceeded Manual_avg on ACC and COM, while other systems in Table 3 scored below Manual_avg on those axes, and no human comparator is available for 2024. The plural 'top-performing models' overgeneralizes; the claim should be restricted to the single best run in that round unless additional supporting evidence is provided.","section":"§5.1.3; §6.3; Table 3"},{"comment":"The metric-correlation analysis uses Pearson correlations on sentence-level averages of a 3-point Likert scale, which produces the striations visible in Appendix G. Pearson r on discrete, bounded data can be sensitive to the chosen aggregation and scale assumptions; reporting Spearman's rho or an ordinal association measure would make the ranking of metrics (SARI and SAMSA lowest, BERTScore highest) more robust. This is a methodological concern rather than a reason to doubt the qualitative conclusion that most reference-based metrics correlate poorly.","section":"§5.1.4; Appendix G"}],"minor_comments":[{"comment":"The sentence 'we omit SAMSA due to lack of poor correlation with manual judgments' appears to have a missing negation; based on §5.1.4 and Appendix G it should read 'due to its poor correlation with manual judgments.'","section":"§5.1.2"},{"comment":"The text says 'Table 1d displays counts of each simplification type,' but the counts appear in Figure 1(d); Table 1 contains the manual evaluation axes. Please correct the cross-reference.","section":"§3.2.2"},{"comment":"The introduction cites Karamcheti et al. (2021) for Mistral, but that reference describes a Stanford CRFM project named Mistral, not the Mistral AI models used later in the paper (e.g., Mistral-Nemo-Instruct-2407 in Appendix D). Please correct or disambiguate the citation.","section":"Introduction; References"},{"comment":"The caption says 'an PLABA abstract pair'; this should be 'a PLABA abstract pair.'","section":"Figure 1(b) caption"}],"recommendation":"major_revision","confidential_remarks":"The positive half of the central claim is on reasonably firm ground, but the negative half is not adequately supported by the data as presented, and the paper's own appendix documents the relevant reliability problems. The manuscript would be publishable after the authors revise the headline claim to match the evidence, add reliability caveats, and re-run or re-report the correlation analysis with an ordinal measure."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Good paper, but the abstract oversells one half of its main claim. What's new is real: two years of PLABA test data, a four-reference 2023 set, a 10,314-term Task 2 lexical simplification dataset, and a metric-correlation analysis showing SARI and SAMSA correlate poorly with manual judgments while BERTScore does better. The manual evaluation is a genuine contribution, especially the human baseline against which top systems are compared.\n\nThe positive half of the headline—that top models rival human factual accuracy and completeness—is well supported. In Table 3, PLABA_1's ACC 97.65 and COM 98.72 beat the human average of 96.21 and 96.16. The inter-annotator agreement for COM and FTH is strong (alpha 0.8833 and 0.8148). So that claim holds.\n\nThe negative half—'but not simplicity or brevity'—is weaker than the abstract suggests. The 2.05-point SIM gap over the human average rests on two sub-axes with poor sentence-level agreement: SEN at 0.1748 and FLU at -0.0255. The paper itself attributes FLU's low alpha to ceiling effects and SEN's to annotators confusing sentence length with brevity. That makes the simplicity gap vulnerable to annotation noise, especially without confidence intervals or significance tests. And 'not brevity' has no human baseline at all: BRV was introduced in 2024, when no reference adaptations were written. You can say systems scored low on brevity; you can't say they fell short of human brevity. The paper flags these issues in Appendix F and Section 4.2, so this is sloppy headline writing rather than a hidden flaw. The hallucination analysis is also explicitly non-systematic, so the cautionary examples are illustrative, not measured rates.\n\nNone of this sinks the paper. The shared-task infrastructure, the public data, and the metric-correlation findings are reusable and valuable. For anyone building or evaluating biomedical text simplification systems, this is a useful reference. The abstract should be softened to match the evidence.\n\nWorth sending to peer review. The reviewer should ask for a revised abstract and some sensitivity analysis or confidence intervals on the manual scores, but the underlying contribution is solid.","headline":"Useful shared-task infrastructure with real new data, but the abstract overstates the evidence on the simplicity and brevity half of the headline claim.","tokens_in":30292,"tokens_out":3439,"would_cite":true,"duration_ms":34052,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The best AI plain-language rewrites of biomedical abstracts matched human accuracy and completeness, but not simplicity or brevity.","keywords":["plain language adaptation","biomedical text simplification","large language models","shared task evaluation","manual evaluation","automatic metrics","patient-facing medical text"],"falsifier":"Re-score the same system outputs with a fresh panel of biomedical expert annotators using the same rubrics; if the new rankings place different systems above the human reference average, or if sentence-simplicity agreement remains near the reported low level, the central claim about rivaling humans does not replicate. A complementary behavioral check would give lay readers either the top system outputs or the original abstracts and test comprehension: if readers understand the original no worse, the practical value of the adaptation is not established.","tokens_in":29145,"feed_emoji":"🩺","tokens_out":6339,"duration_ms":70182,"temperature":0.7,"pith_summary":"This paper reports what a two-year shared task on plain-language adaptation of biomedical abstracts revealed about large language models. The central claim is that top-performing systems rivaled human reference writers on factual accuracy and completeness when rewriting abstracts sentence by sentence, but did not match human simplicity or brevity. The paper also claims that automatic reference-based metrics such as SARI, SAMSA, and BLEU correlate poorly with expert manual judgments, with BERTScore correlating best. If these claims hold, current language models can already support accurate patient-facing summaries of medical literature, but evaluating them requires expert review rather than standard automatic scoring.","feed_headline":"AI abstracts match human accuracy but not simplicity","feed_subtitle":"Two years of expert judging show top systems rival human factuality while automatic metrics mislead.","key_machinery":"The machinery is the PLABA task design itself: a sentence-aligned rewriting task in which each source sentence of a biomedical abstract has a corresponding plain-language output, evaluated on separate axes rather than one overall score. The organizers built four professionally written references per abstract for automatic scoring, an automatic sentence-alignment pipeline for document-level systems, and a manual evaluation rubric that judged simplicity and accuracy in 2023, streamlined to simplicity, accuracy, completeness, and brevity in 2024. The argument for the central claim runs by contrasting top system scores against the human reference average on these manual axes, then correlating each automatic metric with the manual scores.","core_discovery":"On its own terms, the paper's core discovery is that sentence-level plain-language adaptation of biomedical abstracts is largely effective for factual content but not for readability. In the 2023 manual evaluation, the top system scored 97.65 for accuracy and 98.72 for completeness against a human average of 96.21 and 96.16, while its simplicity score of 91.51 fell below the human average of 93.56. In 2024, the best systems pushed accuracy and completeness even higher, and large-language-model-based term replacements in Task 2 also earned high manual scores for accuracy, completeness, and simplicity, though not brevity. Across both years, systems struggled to identify difficult terms and to classify the right replacement strategy. The paper further reports that reference-based automatic metrics generally did not align with manual judgments: SARI and SAMSA had the worst correlations, and BERTScore had the best.","pith_inferences":["Beyond the paper, if the poor metric correlations generalize outside this track, new automatic evaluation for simplification should be validated sentence-level against expert judgments before being used in medical settings.","Beyond the paper, the human baseline here came from reference adaptations written for the dataset; a stronger practical test would compare system outputs against clinicians' real plain-language explanations through a comprehension study with patients.","Beyond the paper, the low inter-annotator agreement on sentence simplicity suggests the simplicity gap may partly be an artifact of rubric ambiguity, and a sharper definition of brevity versus simplicity could change system rankings.","Beyond the paper, given the documented hallucinations, a targeted automatic check that verifies numerical results and named entities against the source sentence could catch the most dangerous errors even when overall accuracy looks high."],"forward_implications":["Current instruction-tuned language models, prompted with abstract context, can generate plain-language versions of biomedical abstracts that preserve factual content and completeness at the level of human-written adaptations.","Because the top systems lagged on simplicity and brevity, further gains require explicit optimization of readability, not just factual fidelity.","Reference-based automatic metrics are not reliable ranking tools for this task, so shared tasks in patient-facing medical text should budget for manual evaluation.","Term-level adaptation is the harder subproblem: identifying difficult terms and choosing replacement strategies scored much lower than generating the replacements themselves.","Document-level systems can participate through automatic sentence alignment, but the alignment cannot handle sentence transposition or merging, which limits some rewriting strategies."],"supporting_citations":[{"why":"Supplies the PLABA dataset and task format that define the sentence-aligned rewriting task and its training data.","marker":"Attal et al., 2023"},{"why":"Defines SARI, the primary reference-based simplification metric whose correlations with manual judgments the paper finds deficient.","marker":"Xu et al., 2016"},{"why":"Defines SAMSA, the reference-free metric reported as having the worst correlation with manual judgments.","marker":"Sulem et al., 2018"},{"why":"Defines BERTScore, the automatic metric that correlated best with manual simplicity and accuracy.","marker":"Zhang et al., 2019"},{"why":"Earlier finding that automatic metrics are unsuitable for text simplification, which the paper's correlation analysis concurs with.","marker":"Alva-Manchego et al., 2021"},{"why":"Survey of biomedical text simplification that frames the manual evaluation axes and the shift from grammatical to factual errors in neural systems.","marker":"Ondov et al., 2022"}],"fun_headline_variants":["AI abstracts match human accuracy but fall short on simplicity","Automatic metrics mislead on plain-language AI abstracts","LLM abstracts: accurate and complete, but not simple or brief","PLABA track: AI rivals human accuracy, not simplicity"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the contracted biomedical experts' manual ratings are reliable ground truth; the paper itself reports low inter-annotator agreement on sentence simplicity and negative agreement on fluency, so if those ratings are noisy the headline comparisons and metric correlations could shift.","fun_headline_variants_meta":{"raw":{"variants":["AI abstracts match human accuracy but fall short on simplicity","Automatic metrics mislead on plain-language AI abstracts","LLM abstracts: accurate and complete, but not simple or brief","PLABA track: AI rivals human accuracy, not simplicity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000542,"raw_usage":{"total_tokens":2633,"prompt_tokens":1015,"completion_tokens":1618,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":631,"completion_tokens_details":{"reasoning_tokens":1551}},"tokens_in":631,"tokens_out":1618,"duration_ms":15221,"temperature":1.0,"reasoning_tokens":1551,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:01:31.470400+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-score the same system outputs with a fresh panel of biomedical expert annotators using the same rubrics; if the new rankings place different systems above the human reference average, or if sentence-simplicity agreement remains near the reported low level, the central claim about rivaling humans does not replicate. A complementary behavioral check would give lay readers either the top system outputs or the original abstracts and test comprehension: if readers understand the original no worse, the practical value of the adaptation is not established.","supporting_citations":[{"cited_title":", author Ondov, B","cited_arxiv_id":null,"evidence_quote":"Supplies the PLABA dataset and task format that define the sentence-aligned rewriting task and its training data."},{"cited_title":", author Abend, O","cited_arxiv_id":null,"evidence_quote":"Defines SAMSA, the reference-free metric reported as having the worst correlation with manual judgments."},{"cited_title":", author Scarton, C","cited_arxiv_id":null,"evidence_quote":"Earlier finding that automatic metrics are unsuitable for text simplification, which the paper's correlation analysis concurs with."},{"cited_title":", author Attal, K","cited_arxiv_id":null,"evidence_quote":"Survey of biomedical text simplification that frames the manual evaluation axes and the shift from grammatical to factual errors in neural systems."}],"review_version":1}