{"id":"cdcb371c-a4ac-4502-b447-e681040fe42f","arxiv_id":"2505.22280","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A scoping review of 129 studies organizes NLP-for-EBM research along the five-step evidence-based medicine cycle, identifies task trends, and lists benchmark gaps.","lead":"This paper is a scoping review of 129 research studies that use natural language processing to support evidence-based medicine. It maps these studies onto the five steps of the EBM cycle (Ask, Acquire, Appraise, Apply, Assess) and identifies gaps in benchmarks and future research needs.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"PRISMA flow arithmetic is internally inconsistent, so the headline 129-study count is not reproducible from the paper's own numbers.","rationale":"The reader's weakest assumption concerns auditing of the screening process; I agree that auditability is central. My stress-test finds a more concrete, internal version of that problem: the PRISMA flow diagram and the prose in Section 2.4 cannot both be right, so the exact number of screened records, and hence the exact number of included studies, is not determined by the manuscript as written. This is a correctness risk specific to the paper's central claim rather than a disagreement with external consensus. I do not recommend rejection because the discrepancy is plausibly a reporting error, and the descriptive trends in Figure 2 (entity extraction/classification and evaluation dominance) are likely robust to a one-count correction. The condition is exactly the one the reader already imposed: the authors must release the screening metadata, reconcile the PRISMA arithmetic, and confirm the 129 unique-study count. With that, the verdict remains CONDITIONAL rather than ACCEPT; no further adjustment is needed.","tokens_in":24991,"tokens_out":5828,"duration_ms":67419,"concrete_test":"Reconstruct the Covidence export from the exact A.3 queries and the nine additional sources: record per-database hits and duplicates, then check whether screened = database hits + other sources - duplicates equals 603. Independently enumerate the 129 included rows from Supplementary Table 1 and deduplicate by DOI/title; if screened is not 603 or the deduplicated count is not 129, correct the headline and re-plot Figure 2. If the reconstruction yields 603 and 129 unique studies, the census claim stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is a census: '129 research studies.' The paper's own PRISMA reconstruction does not add up. Section 2.4 says the pool was '601 papers retrieved from databases and 9 additional sources' and that 8 duplicates were removed; 601 + 9 - 8 = 602, yet Figure 1 reports 603 studies screened. The figure also lists duplicate removals as 6 manual + 1 Covidence = 7, not 8. Since the included count is obtained as screened - 386 - 88, an off-by-one in the denominator changes 129 to 128 or, if screened is interpreted as 601 - 8 = 593, to 119. In other words, the headline count survives only under one unstated interpretation of the source/duplicate arithmetic, and the manuscript does not give the per-database hit counts, the list of the 9 additional sources, or a versioned GitHub snapshot needed to resolve it. A separate count risk: Kim et al. (2023a) and Kim et al. (2023b) in Supplementary Table 1 and the reference list are the same title, venue, and page range, entered as two rows, which would also reduce the unique-study count.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This scoping review surveys NLP applications in evidence-based medicine (EBM), following PRISMA reporting. The authors report 129 included studies, organize them by the five EBM steps (Ask, Acquire, Appraise, Apply, Assess), review NLP techniques per step, summarize benchmark datasets, and discuss limitations and future directions. The manuscript includes a PRISMA flow diagram, detailed supplementary tables of included studies and benchmarks, explicit search queries, and a public GitHub repository.","tokens_in":25226,"tokens_out":4788,"duration_ms":52070,"significance":"If the numbers are corrected, this would be a useful structured map of a fast-moving interdisciplinary area. The paper's strengths are its systematic selection pipeline, the explicit and reproducible search strategy, the public release of study metadata, and the organization around the 5A EBM framework. The census of 129 studies, however, is the backbone of the review, so the internal arithmetic of the PRISMA flow and the presence of duplicate entries are consequential rather than cosmetic issues. The benchmark overview and task-trend summary are genuinely useful resources for researchers entering this area.","major_comments":[{"comment":"The PRISMA flow arithmetic is internally inconsistent. The text states that the pool was \"601 papers retrieved from databases and 9 additional sources\" and that \"8 duplicates\" were removed; 601 + 9 - 8 = 602, yet Figure 1 reports 603 studies screened. Figure 1 also lists only 7 duplicate removals (6 manual + 1 Covidence) in the \"References removed\" box, not 8. Because the final count is obtained as screened (603) - 386 - 88 = 129, an alternative reading in which screened should be 593 would yield 119 included studies. The authors should report per-database hit counts, list the 9 additional sources, and provide a single consistent duplicate-removal count so that the headline number 129 is reproducible.","section":"Section 2.4 and Figure 1"},{"comment":"The entries Kim et al. (2023a) and Kim et al. (2023b) refer to the same paper: the reference list contains two identical full citations (same title, venue, and page range), and Supplementary Table 1 lists them as two different included studies with different model/disease annotations. This duplication implies that the unique-study count is at most 128. The table should be de-duplicated, and the task-distribution statistics in Figure 2 should be recomputed accordingly.","section":"Supplementary Table 1 and reference list"},{"comment":"The screening process is described as having two annotators cross-verify study selection and metadata extraction with third-party arbitration, but no inter-annotator agreement measure is reported. Because the 129-study corpus and all downstream trends depend on the consistency of inclusion and exclusion decisions, reporting Cohen's kappa or a comparable statistic for the title/abstract and full-text screens is necessary for readers to assess the reliability of the corpus.","section":"Section 2.3"}],"minor_comments":[{"comment":"The search string line \"{ebm_keywords)\" has a mismatched closing parenthesis; the prose says the Boolean strategy combines any NLP term with any EBM term, but the displayed queries combine the full keyword groups with AND, not individual terms.","section":"Appendix A.3"},{"comment":"The task legend uses the code \"S\" for both Abstractive Summarization and Extractive Summarization, which makes it impossible to tell which task is assigned to a given paper. A distinct code for extractive summarization would remove the ambiguity.","section":"Supplementary Table 1 legend"},{"comment":"The row label \"Apply & Asses\" contains a typo and should read \"Apply & Assess.\"","section":"Table 1"},{"comment":"The phrase \"domain-aware pertaining benefits\" should be \"domain-aware pre-training benefits.\"","section":"Section 5.1"},{"comment":"The task names in Figure 2 (e.g., \"Evaluation\", \"Summarization\") are not fully aligned with the task codes in Supplementary Table 1 (e.g., \"Evaluation of Performance\", two separate S codes); a supplementary table of per-year counts per task would make the trends easier to verify.","section":"Figure 2"}],"recommendation":"major_revision","confidential_remarks":"The review is within the scope of the journal and the underlying idea is sound, but the census claim is currently not reproducible from the manuscript's own numbers. The PRISMA inconsistencies and the duplicated Kim entry should be fixed before publication, and the authors should be asked to provide a corrected figure, a de-duplicated study list, and, ideally, a versioned snapshot of the metadata repository so that the 129-study count can be independently verified."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely useful scoping review that maps 129 papers onto the 5A cycle of evidence-based medicine. If you're planning work in NLP for EBM, the benchmark table and the GitHub list save you a lot of searching. The paper does what a good scoping review should: clear inclusion criteria, PRISMA flow, dual-annotator screening with third-annotator arbitration, and an honest Limitations section that concedes the English-only, 2019-2024, four-database scope.\n\nThe soft spot is the census itself, and it's not cosmetic. Section 2.4 says 601 database records plus 9 from other sources, with 8 duplicates removed. That's 602 to screen. Figure 1 reports 603 screened, and its own duplicate breakdown sums to 7 (6 manual + 1 Covidence). So the numbers only reconcile if you assume the '8' is a typo for '7' or that one of the 9 additional sources wasn't actually added. The paper gives no per-database hit counts and no versioned snapshot of the screening metadata to untangle it. A second issue: in Supplementary Table 1 and the reference list, Kim et al. (2023a) and Kim et al. (2023b) are the same title, venue, and page range entered twice. That knocks the unique-study count from 129 to 128, unless the duplication is just a table artifact. For a paper whose output is a count and a taxonomy, these inconsistencies are the load-bearing part. They should have been caught before submission.\n\nThe concerns about missing kappa values and exclusion lists are fair but secondary; the appendix does give the actual query templates, which is more than many reviews do. The lack of inter-annotator agreement statistics weakens the reliability claim a bit.\n\nBottom line: the review is worth engaging. It's a solid organizing contribution for researchers entering the area. I'd send it to peer review, but I'd ask for a mandatory reproducibility check: reconcile the PRISMA cascade numbers, provide the excluded-studies list and kappa, and fix the duplicate entry.\n\nWho's this for: researchers entering the NLP-for-EBM area, or anyone writing a grant or background section on clinical NLP. Not a groundbreaking methods paper, but a serviceable map for the field.\n\nRecommendation: accept with major revision, conditional on auditable screening metadata and corrected counts.","headline":"Useful 5A-organized scoping review of NLP for EBM, but the PRISMA arithmetic for the headline 129-study count does not add up as printed.","tokens_in":25764,"tokens_out":3829,"would_cite":true,"duration_ms":36064,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper maps 129 NLP studies onto the five steps of evidence-based medicine, showing where the field is concentrated and where it is thin.","keywords":["scoping review","evidence-based medicine","natural language processing","large language models","PICO extraction","clinical trial matching","evidence synthesis","benchmark datasets"],"falsifier":"Re-run the same five-step mapping with a broader search that adds databases such as Embase and Web of Science and terms like 'systematic review automation', 'trial matching', or 'risk of bias', and count how many additional evidence-synthesis, appraisal, and QA papers surface. If the distribution of tasks shifts materially—or if many of the 129 papers are reclassified under a different annotator—the paper's claim that extraction and evaluation dominate would be weakened.","tokens_in":24690,"feed_emoji":"📚","tokens_out":4316,"duration_ms":38773,"temperature":0.7,"pith_summary":"This paper is a scoping review asking how natural language processing (NLP) actually supports evidence-based medicine (EBM). The authors screened 603 candidate papers and selected 129 that apply NLP to EBM tasks, then mapped them onto the five steps of the EBM cycle—Ask, Acquire, Appraise, Apply, and Assess. Their central claim is that this corpus reveals a clear division of labor: entity extraction (especially PICO elements) and performance evaluation dominate the field, while evidence synthesis, formal appraisal, and question-answering benchmarks are comparatively under-served. The review also documents a shift from rule-based and recurrent neural network systems toward transformer models and large language models, and it argues that trustworthiness, hallucination, and a scarcity of medical-specific benchmarks are now the main blockers. A sympathetic reader would take the paper to establish a structured map of where NLP is actually being used in EBM and where the gaps are.","feed_headline":"129 NLP studies mapped onto EBM's five steps","feed_subtitle":"A scoping review shows extraction and evaluation dominate, while evidence synthesis and appraisal benchmarks lag.","key_machinery":"The organizing device is the 5A framework of evidence-based medicine (Ask, Acquire, Appraise, Apply, Assess), used as a fixed taxonomy to classify each of the 129 studies. Within that taxonomy, the PICO schema (Patient/Population, Intervention, Comparison, Outcome) is the central object that most extraction, retrieval, and synthesis work is built around; the paper treats PICO extraction and normalization as the canonical NLP task for EBM. The screening machinery is the PRISMA flow, which provides the reproducibility of the corpus, and the paper's supplementary metadata tables attach every study to a model type, a disease, and an EBM task.","core_discovery":"The paper's discovery is a taxonomy, not a new algorithm. By organizing 129 studies according to the 5A's, it shows that NLP-for-EBM research is concentrated in the early and late parts of the pipeline—searching and retrieving literature (Acquire) and applying evidence in the clinic (Apply/Assess)—with entity extraction and classification as the most frequent task across all years, peaking with the 2023 wave of large-language-model papers. It further claims that quality assessment, evidence ranking, evidence synthesis, and summarization remain thinner, and that existing benchmarks for synthesis and appraisal are rarely medical-specific. The implied finding is that the middle of the EBM workflow—appraising and synthesizing evidence—is the least automated and the most open for NLP development, even though LLMs have begun to be tested there.","pith_inferences":["A testable extension the paper does not run: use the same 5A map to classify LLM-era papers published after 2024 and see whether the QA and synthesis categories grow faster than extraction—if they do, the 'under-served middle' claim would need updating within a year or two.","The review's exclusion of non-English studies likely understates NLP-for-EBM activity in health systems where English is not the clinical language; a parallel scoping review in Chinese, Spanish, or German clinical NLP would be a natural complement.","The emphasis on scarcity of medical-specific benchmarks suggests an opportunity: adapting existing general summarization benchmarks like CNN-DailyMail—which the paper flags as non-medical—into clinically validated test sets could shift evaluations toward the steps the field currently neglects."],"forward_implications":["If the map is right, new NLP research for EBM should focus on evidence synthesis, appraisal, and question answering, since those steps are the least covered and LLM evaluation is still early.","The 5A taxonomy gives a common vocabulary for comparing future systems: a new tool can be placed at a specific step and benchmarked against the corpora the review catalogues (EBM-NLP, Chia, PICO-Corpus, MS^2, Trialstreamer).","The shift toward LLMs documented in 2022–2024 papers implies that clinicians can expect more conversational, QA-style EBM tools, but the review's challenges section says these need retrieval augmentation and source attribution before clinical use.","The scarcity of medical-specific benchmarks for synthesis and appraisal means that progress in those steps will require dataset construction as much as model innovation."],"supporting_citations":[{"why":"Supplies the definition of evidence-based medicine that frames the entire review.","marker":"Sackett et al., 1996"},{"why":"Provides the 5A framework (Ask, Acquire, Appraise, Apply, Assess) that structures the paper's taxonomy.","marker":"Ratnani et al., 2023"},{"why":"Introduces the EBM-NLP corpus, the foundational dataset for PICO extraction used across many of the reviewed studies.","marker":"Nye et al., 2018"},{"why":"Describes Trialstreamer, a living database of RCT reports that serves as a benchmark for retrieval and evidence tasks.","marker":"Marshall et al., 2020"},{"why":"Compares PubMed search filters used in the Ask step, grounding the review's discussion of literature retrieval methods.","marker":"Navarro-Ruan and Haynes, 2022"},{"why":"Presents the PICO-Corpus, which supports automatic data extraction and synthesis in the Acquire step.","marker":"Mutinda et al., 2022a"}],"fun_headline_variants":["NLP for evidence-based medicine lags at synthesis and appraisal","NLP for EBM: strong at search and apply, weak at synthesis","EBM automation gap: NLP skips appraisal and synthesis","LLMs touch EBM's edges, not its evidence core"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review's map of the field is only as complete as its keyword-based search; the assumption is that requiring NLP terms and EBM terms to co-occur in titles or abstracts, searching four databases, and restricting to English-language work from 2019–2024 captured the relevant literature without systematic bias.","fun_headline_variants_meta":{"raw":{"variants":["NLP for evidence-based medicine lags at synthesis and appraisal","NLP for EBM: strong at search and apply, weak at synthesis","EBM automation gap: NLP skips appraisal and synthesis","LLMs touch EBM's edges, not its evidence core"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000895,"raw_usage":{"total_tokens":3818,"prompt_tokens":866,"completion_tokens":2952,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":482,"completion_tokens_details":{"reasoning_tokens":2879}},"tokens_in":482,"tokens_out":2952,"duration_ms":20768,"temperature":1.0,"reasoning_tokens":2879,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:09:44.147209+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same five-step mapping with a broader search that adds databases such as Embase and Web of Science and terms like 'systematic review automation', 'trial matching', or 'risk of bias', and count how many additional evidence-synthesis, appraisal, and QA papers surface. If the distribution of tasks shifts materially—or if many of the 129 papers are reclassified under a different annotator—the paper's claim that extraction and evaluation dominate would be weakened.","supporting_citations":[],"review_version":1}