{"id":"9f06ff3c-2580-477c-a329-6143cf4e7f74","arxiv_id":"2504.20113","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of 54 automated meta-analysis studies finds automation concentrated in data processing, with bias assessment and full-process automation rarely achieved.","lead":"This paper reviews 54 studies on automated meta-analysis and reports that most automation focuses on data extraction and statistical modeling, while bias assessment and full-pipeline automation remain rare. It offers a framework for mapping automation gaps across medical and non-medical domains, useful for researchers building AI tools for evidence synthesis.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Appendix coding contradicts \"only one full-process study,\" undermining the 2% headline statistic.","rationale":"After reading the paper, the strongest claim is the quantitative distribution of automation across PPS stages, introduced in §4 and repeated in the abstract. The paper's own Appendix 1 is the only place where per-study stage assignments are operationalized, using a single marker per stage. Row 12 (study [16]) and row 13 (study [82]) show all three phases marked, along with row 24 (study [39]), yet §6.2 says 'Only one study [39] has explored full automation across all AMA stages.' This is not a minor typo: the abstract's 2% figure and the paper's central narrative about the 'critical gap' in full-process automation depend on this count. Either the appendix is right, in which case the abstract is wrong, or the appendix markers are unreliable, in which case the quantitative findings lack a transparent basis. The paper's own limitation statement (§6.5) concedes that the automation-level criteria are 'subjective and qualitative,' but the conclusion nevertheless reports precise percentages. I agree with the reader's verdict that this inconsistency warrants rejection of the central quantitative claim. The qualitative direction of the review (that automation clusters around data processing and advanced synthesis remains rare) is plausible and supported by many individual study descriptions, but the headline statistics are not internally consistent or auditable. The proposed test—independent recoding of the appendix—would settle the discrepancy by determining whether the actual full-process count is one or three, or whether the markers were applied inconsistently. This is the single most load-bearing concern because it attacks the paper's main contribution (a quantitative map of the AMA landscape) from its own data.","tokens_in":33837,"tokens_out":5186,"duration_ms":49689,"concrete_test":"Recode all 54 studies in Appendix 1 with an independent second rater, using the PPS definitions in §3.2: a study counts as full-process automation only if its full text shows automated execution in all three phases. Compute the number of studies satisfying this criterion and compare with the abstract's '2% (one study).' Specifically, check the full texts of [16], [39], and [82]: if [16] or [82] performs automated steps in all three phases, the 'only one study' claim fails; if neither does, then Appendix 1's markers are coding errors for those rows, which still invalidates the stated statistics.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim (abstract: 'Just one study (2%) explored preliminary full-process automation'; Section 6.2: 'Only one study [39] has explored full automation across all AMA stages') is contradicted by the paper's own Appendix 1. Rows 12, 13, and 24 mark studies [16] (Michelson 2014), [82] (Neupane et al. 2014), and [39] as covering pre-processing, processing, and post-processing. If those markers mean the studies address all three PPS phases, then the 'one study' count is wrong by a factor of three. If they mean something weaker, the stage-marker scheme is not reliable enough to support the precise percentages in the abstract (57% processing, 17% advanced synthesis, 2% full automation). The review provides no coding rubric for the markers, and §6.5 admits the criteria for assessing automation level 'remain subjective and qualitative.' The headline numbers therefore rest on an unvalidated, self-inconsistent coding, making the central claim non-reproducible from the provided materials.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports a PRISMA-guided systematic review of automated meta-analysis (AMA), screening 978 records and analyzing 54 studies published between 2006 and 2024. The authors introduce a Progressive Phase Structure (PPS) that splits AMA into pre-processing, processing, and post-processing, and combine it with Task-Technology Fit (TTF) judgments. They report that research is concentrated in the data-processing stage (57%), that only 17% address advanced synthesis (heterogeneity, bias, sensitivity), and that just one study (2%) attempts full-process automation. They also compare CMA versus NMA, medical versus non-medical domains, and propose future research directions, including fine-tuned LLMs, living AMA, and interpretability standards.","tokens_in":1296,"tokens_out":1295,"duration_ms":74590,"significance":"If the quantitative claims were reliable, the review would fill a real gap: it is broader than an earlier clinical-trials-focused AMA review [23] and it provides a 54-study appendix organized by field, methodology, and automation stage. The PRISMA process is described, the snowballing extension is a useful design choice, and the domain-level comparison (medical vs. non-medical, CMA vs. NMA) is a genuinely useful contribution. The paper is also candid about its limitations (§6.5). However, the headline quantitative findings rest on a coding scheme that is not defined in the manuscript and is contradicted by the paper's own appendix; as a result, the central claims are not reproducible from the supplied materials.","major_comments":[{"comment":"The abstract states that \"just one study (2%)\" explored preliminary full-process automation, and §6.2 repeats \"Only one study [39] has explored full automation across all AMA stages.\" Appendix 1, however, marks three studies with all three PPS stage markers: [16] (row 12), [82] (row 13), and [39] (row 24) each have pre-processing, processing, and post-processing markers. If those markers mean the study addresses all three PPS phases, the count is three out of 54 (5.6%), not one (2%). If the markers mean something weaker, then the appendix does not support the \"one study\" claim and the counting rule is not defined. Either way, the central quantitative headline is internally inconsistent and must be reconciled.","section":"Abstract & §6.2 vs. Appendix 1 (rows 12, 13, 24)"},{"comment":"No coding rubric is provided for the Appendix 1 stage markers or for the High/Moderate/Low TTF ratings used throughout Tables 2–6, and §6.5 concedes that \"the criteria for assessing the level of automation remain subjective and qualitative.\" This is load-bearing because the abstract's percentages (57% processing, 17% advanced synthesis, 2% full automation) are computed from exactly this coding. Without an explicit protocol, a worked example of stage assignment, or at least a cross-tabulation linking Figure 4B's percentages to Appendix 1's rows, the quantitative results cannot be independently verified. The \"advanced synthesis\" category (17%) is especially problematic because PPS defines only three broad phases and the appendix does not separately code heterogeneity, bias, or sensitivity analysis.","section":"§3.2, §6.5, and Figure 4B"},{"comment":"The results state that 89% of studies focused on a single AMA step and 11% addressed multiple stages, but the manuscript does not list which studies are counted as multi-stage or provide the mapping from that binary classification to Appendix 1's three-column markers. This compounds the counting problem in the previous comment: the 11% multi-stage figure and the 2% full-automation figure are not derivable from the appendix as presented. The authors should provide a transparent study-by-study coding table, or revise the quantitative claims to be consistent with the markers they actually report.","section":"§4, first paragraph"}],"minor_comments":[{"comment":"The text says inclusion criteria require publication from 2014 to 2024, then says snowballing \"expanded our temporal scope to 2006-2024,\" while Table 1 lists all dates from 2006 to 2024 as accepted; please reconcile these three statements.","section":"§3.1 and Table 1"},{"comment":"Several small writing and consistency issues should be fixed: \"medical filed\" (§4.4.1), \"Enanced accessibility\" (§4.4.2), \"frontier fro future\" (§6), and inconsistent spelling of \"metaGWASmanager\"/\"MetaGW ASmanager\" in §4.4.1 and Appendix 1.","section":"Throughout"},{"comment":"The caption and text say line thickness reflects application frequency, but the figure itself has no legend defining line thickness; please add one for reproducibility.","section":"Figure 6"}],"recommendation":"major_revision","confidential_remarks":"The core problem is the internal contradiction between the appendix coding and the headline \"one study\" claim. I do not see this as a fatal flaw beyond repair: the qualitative conclusion that full-process automation is rare and that advanced synthesis is underdeveloped is likely robust even if the correct count is three rather than one. However, the authors must either supply a coding rubric and reconcile the appendix, or substantially soften the precise percentages. If the authors cannot provide such a rubric, the paper's quantitative contribution should be withdrawn and the findings presented only qualitatively."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bad news first: the abstract's headline numbers don't survive contact with the paper's own appendix. Section 6.2 says only one study [39] achieved full automation across all stages, and the abstract rounds this to 2%. But Appendix 1 marks three studies—[16], [82], and [39]—as having automated steps in all three PPS phases. Whether that means 'full automation' is unclear, and that's the problem: the paper never defines what a checkmark in the appendix means, and §6.5 concedes the automation-level criteria 'remain subjective and qualitative.' Without a coding rubric, the precise percentages (57%, 17%, 2%) are not reproducible, and the central claim that full-process automation is vanishingly rare is not supported by the paper's own evidence.\n\nThe paper is still worth engaging. It's the first cross-domain systematic review of automated meta-analysis that I know of, with a PRISMA process, a curated set of 54 studies, and a decent appendix table. The PPS/TTF framework is a reasonable organizing device, even if it's more descriptive than analytical. The qualitative findings—automation concentrates in data extraction and statistical modeling, while advanced synthesis like bias and heterogeneity assessment lags—are plausible and consistent with prior reviews of SLR automation. The medical/non-medical comparison is a useful addition. The limitations section is honest, which is more than many such papers manage.\n\nThe fixes are clear: align the appendix markers with a defined rubric, correct the count of full-process studies, and either drop the precise percentages or label them as approximate based on subjective classification. If the authors do that, the review would be a solid reference for the field. As it stands, I'd caution anyone against citing the 2% figure.\n\nWho's it for? People mapping the evidence-synthesis automation landscape, or researchers looking for a starting list of tools. It deserves peer review—not desk reject—because the corpus and framework are valuable and the flaws are fixable. But the referee should demand the coding be made auditable before acceptance.","headline":"Useful survey of automated meta-analysis with a curated corpus, but the headline '2% full automation' is contradicted by the paper's own appendix and rests on subjective coding.","tokens_in":34532,"tokens_out":3561,"would_cite":false,"duration_ms":34947,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Automated meta-analysis remains concentrated in data extraction and statistical modeling, with full-process automation nearly absent.","keywords":["automated meta-analysis","systematic review","evidence synthesis","large language models","task-technology fit","progressive phase structure","PRISMA","automation stages"],"falsifier":"Re-code the 54 included studies using the paper's own PPS definitions. If two or more studies beyond [39] are found to automate pre-processing, processing, and post-processing—as Appendix 1's rows for [16] and [82] suggest—the 2% full-automation claim is refuted; if independent coding reproduces a single full-process study, the claim survives.","tokens_in":33666,"feed_emoji":"📊","tokens_out":5773,"duration_ms":54452,"temperature":0.7,"pith_summary":"This paper is a PRISMA systematic review of 54 automated meta-analysis (AMA) studies, screened from 978 records published 2006–2024. It argues that AMA research is lopsided: most automation effort targets data pre-processing and processing steps such as literature retrieval, information extraction, and statistical modeling, while higher-order synthesis tasks like heterogeneity assessment, bias evaluation, and integrated reporting remain largely manual. The headline evidence is a stage distribution in which 57% of studies automate data processing, 17% touch advanced synthesis, and only one study (2%) attempts full-process automation. A sympathetic reader would take this as a map of where the field actually is, and as an argument that LLM-driven reasoning should now be aimed at the synthesis stages rather than at extraction.","feed_headline":"Most meta-analysis automation stops at data processing","feed_subtitle":"A 54-study review finds 57% of effort on extraction and modeling, 2% on full-process automation.","key_machinery":"The machinery is the Progressive Phase Structure (PPS) with Task-Technology Fit (TTF). PPS divides meta-analysis into three phases—data pre-processing (problem definition, query design, literature retrieval), data processing (information extraction and statistical modeling, or network construction in NMA), and data post-processing (database building, diagnostics, reporting, visualization)—and TTF grades how well a technology's characteristics match each phase's tasks. This framework does the paper's quantitative work: it is the instrument by which the 54 studies are tagged by stage, the 57%/17%/2% distribution is produced, and the medical versus non-medical comparison is explained.","core_discovery":"The central claim is that automated meta-analysis has matured as a set of point solutions rather than as an end-to-end pipeline. Using a three-stage Progressive Phase Structure (pre-processing, processing, post-processing) paired with Task-Technology Fit, the authors classify 54 studies and find that 89% automate a single meta-analysis step, only 11% address multiple stages, and just one study ([39]) spans all three. They further report that 57% of the work concentrates on data processing, compared with 17% on advanced synthesis, and that LLMs have entered extraction and screening but remain underused in statistical modeling and higher-order synthesis such as heterogeneity and bias assessment. Medical applications (67% of the dataset) show stronger task-technology fit because clinical trial data is standardized, while non-medical fields (33%) struggle with heterogeneous data and reporting styles. The paper concludes that achieving seamless, end-to-end automation remains an open challenge, and that LLMs with reasoning capability are the natural next step for closing it.","pith_inferences":["The paper's own Appendix 1 marks [16] and [82] as covering pre-processing, processing, and post-processing; if those markers are read literally, the claim that only [39] achieved full-process automation would need to be revised upward, undercutting the 2% headline.","A natural next step, not pursued in the paper, is to turn the TTF fit ratings into a quantitative rubric and have independent coders re-score the 54 studies, which would test whether the stage distribution is stable.","The same framework could be applied prospectively: new LLM-based tools could be benchmarked by which PPS stages they cover and how well they handle heterogeneity and bias, giving the field a shared scorecard."],"forward_implications":["If the distribution is accurate, the highest-value next target for AMA is not more extraction tools but automated heterogeneity assessment, bias evaluation, and sensitivity analysis.","Full-process AMA remains an open problem, so near-term systems should be semi-automated, with expert oversight at synthesis and interpretation steps.","In standardized medical data, automation fit is strong and near-term deployment can proceed, while non-medical domains need more adaptable, less format-dependent tools.","LLM-based extraction is viable but not yet trustworthy for quantitative outcomes, so hallucination control and validation benchmarks are prerequisites for clinical use.","A living AMA that continuously updates evidence will require new infrastructure for monitoring, version control, and reconciling conflicting new studies."],"supporting_citations":[{"why":"Supplies the Task-Technology Fit lens used to assess how well automation aligns with meta-analysis tasks.","marker":"[36]"},{"why":"The prior dedicated AMA review focused narrowly on clinical trials, which this review extends across domains.","marker":"[23]"},{"why":"Provides the snowball sampling method used for backward and forward citation chaining.","marker":"[37]"},{"why":"The one study counted as full-process automation in Section 6.2, making it the load-bearing example for the 2% claim.","marker":"[39]"},{"why":"An early RCT automation study that Appendix 1 marks across all three stages, relevant to the full-automation count.","marker":"[16]"},{"why":"A review of NMA R packages that Appendix 1 also marks across all three stages, relevant to the full-automation count.","marker":"[82]"},{"why":"Prior systematic literature review of automation that frames the field and its gaps.","marker":"[8]"}],"fun_headline_variants":["Meta-analysis automation: 57% on data, 2% on full process","Only 1 of 54 studies automates entire meta-analysis","AI meta-analysis: data prep automated, synthesis not","Automated meta-analysis stuck in data processing stage","From 54 studies, just one fully automates meta-analysis"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline percentages rest on the authors' stage-by-stage coding of which studies automate which phases, and specifically on counting only study [39] as full-process automation; the appendix's own markers for [16] and [82] appear to contradict that count.","fun_headline_variants_meta":{"raw":{"variants":["Meta-analysis automation: 57% on data, 2% on full process","Only 1 of 54 studies automates entire meta-analysis","AI meta-analysis: data prep automated, synthesis not","Automated meta-analysis stuck in data processing stage","From 54 studies, just one fully automates meta-analysis"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000254,"raw_usage":{"total_tokens":1609,"prompt_tokens":1028,"completion_tokens":581,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":644,"completion_tokens_details":{"reasoning_tokens":496}},"tokens_in":644,"tokens_out":581,"duration_ms":5488,"temperature":1.0,"reasoning_tokens":496,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:53:08.222436+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-code the 54 included studies using the paper's own PPS definitions. If two or more studies beyond [39] are found to automate pre-processing, processing, and post-processing—as Appendix 1's rows for [16] and [82] suggest—the 2% full-automation claim is refuted; if independent coding reproduces a single full-process study, the claim survives.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The prior dedicated AMA review focused narrowly on clinical trials, which this review extends across domains."}],"review_version":1}