REVIEW 3 major objections 5 minor 11 references
Systematic Literature Reviews With Two Multi-Agentic Systems And Human-In-The-Loop
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Two multi-agent systems with human-in-the-loop can reproduce a published network meta-analysis and add 17 eligible trials that change treatment rankings.
desk verdict The screening MAS evidence is solid and worth a look, but the network-meta-analysis heading is not reproducible as reported and the 'uniform improvement' claim overstates Table 3. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the separation of screening and extraction into two fit-to-purpose multi-agent systems. Screening uses N agents with heterogeneous personas (strict regulator, permissive clinician, statistician, clinical pharmacologist) plus an inspector that generates targeted follow-up questions and a result agent; consensus finalizes, while disagreement and 'maybe' go to a human. Extraction uses specialized treatment, subgroup, and endpoint agents to build a standardized mock table, an extraction agent that proposes values and re-reads the source document for T rounds, assigns High, Medium, or Low confidence from cross-round consistency, and retrieves only matching evidence clips to control token cost. These mechanisms convert individual-model failure modes, such as over-conservative screening, hallucinated numbers, and missed subgroups, into measurable disagreement signals that a human can audit.
What would settle it
Re-run the NSCLC screening set with an independent panel of reviewers who do not see the original consensus labels, then compare the panel's final decisions with the MAS decisions; if the panel labels a substantial fraction of the MAS's high-confidence inclusions or exclusions as errors, or if the sensitivity gain from 0.58 to 0.93 reverses under the new labels, the uniform-improvement claim fails. A cheaper check is to apply the full pipeline to a completed Cochrane or similar review with published decisions and extracted data and count exact agreement on both screening and numeric extraction.
Extended reading notes
Core claim
The central claim is that a multi-agent architecture, rather than any one model, is what makes LLM-based systematic review trustworthy enough for clinical use. Screening is performed by several agents with deliberately different personas who vote and then answer targeted follow-up questions from an inspector until consensus or a turn limit; the decision rule sends disagreement, 'maybe,' and externally discovered trials to a human. In 30 Monte Carlo runs on 2,236 non-small-cell lung cancer trials, this raises F1 from 0.736 to 0.955 for the most conservative base model and improves F1 for all three tested models over their single-agent baselines. Extraction is similarly decomposed into standardization, iterative correction with confidence labels, and retrieval-based context control; on 804 endpoint cells across three cancers, 92.2% are high-confidence, 90.4% of those exactly match the source, and correcting flagged cells raises overall accuracy to about 97%. Reproducing a published metastatic colorectal cancer network meta-analysis, the system recovered all 29 originally included trials (seven of them outside the ClinicalTrials.gov knowledge base), surfaced 17 additional eligible studies, and the updated network changes the relative standing of bevacizumab-based combinations.
Load-bearing premise
The evaluation treats the final human-consensus labels as ground truth, and in the NSCLC screening study the three reviewers reached only moderate initial agreement (Fleiss' Kappa = 0.70) before reconciling, so if that consensus is biased the measured accuracy gains are measured against a flawed yardstick.
Editorial extensions
If this is right
- Screening time can decouple from the number of studies because agents process the knowledge base in parallel, approaching the latency of a single inference cycle rather than a linear per-paper cost.
- The architecture is model-agnostic, so stronger underlying language models, retrieval engines, or image-recognition tools should improve both modules without redesign.
- The confidence labels plus source-risk flags concentrate human audit on roughly 10% of extracted data points while catching most errors, making the workload of human oversight explicit and learnable.
- If the reproduction result holds, published network meta-analyses can be revisited cheaply as new trials appear, and evidence syntheses used in regulatory and clinical decisions can be kept current.
- The uniform screening improvement over single-agent baselines suggests that agent disagreement itself is a signal worth routing to human review, not just a voting byproduct.
Reading between the lines
- Editorial extension: because the evaluation is oncology-only and tied to registry-based NCT identifiers, the strongest untested check is a cross-specialty gold-standard benchmark, such as a completed Cochrane review in cardiology or infectious disease, where inclusion criteria are structured differently.
- Editorial extension: the paper recovers pre-2007 and non-NCT trials through external web-search agents rather than through the knowledge base itself; a natural next test is whether the screening MAS can ingest those registries natively and sustain the same recall gain.
- Editorial extension: the confidence-label mechanism could be validated prospectively as a triage instrument by recording, in new disease areas, whether human correction rates rise monotonically from High to Medium to Low confidence cells.
- Editorial extension: the 17 newly identified trials in the colorectal cancer reproduction come from a single published review; re-running the same pipeline on several independent network meta-analyses would show whether the missed-trial rate is typical of manual reviews or specific to this benchmark.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes two multi-agentic LLM systems (MAS) for systematic literature review: a screening MAS that combines heterogeneous personas, multi-round cross-review, and human-in-the-loop adjudication, and an extraction MAS that standardizes target fields, iteratively corrects extractions, assigns confidence labels, and uses retrieval-based context control. The screening system is evaluated on a 2,245-trial NSCLC benchmark with 30 replicates, comparing each MAS against a same-base-model single-agent baseline under two adjudication rules. The extraction system is evaluated on 804 entries from TNBC, NSCLC, and gastric cancer trials. As a real-world application, the authors reproduce a published network meta-analysis (NMA) of first-line metastatic colorectal cancer, claim to recover all 29 original trials, identify 17 additional eligible trials, and report updated clinical conclusions.
Significance. If the screening and extraction results hold, the paper makes a useful contribution to semi-automated systematic reviewing. The screening experiments are well designed in an important respect: the MAS is compared with a single-LLM baseline using the same base model, so gains are attributable to the multi-agent workflow rather than to a stronger foundation model. The 30-run replication design, the reporting of majority-vote (Rule A) results without human correction, and the explicit prompt library in the appendix all strengthen the credibility of the screening comparison. The extraction confidence-label scheme is also a practical contribution, and the claim that human-review flags concentrate errors is valuable for workload estimation. However, the headline claim of 'updated clinical conclusions' from the reproduced NMA is not currently reproducible or auditable, and the 'uniform improvement' wording is stronger than Table 3 supports. These issues are load-bearing for the abstract and conclusion but appear fixable within the manuscript's scope.
major comments (3)
- [Section 5 and Table 3] The conclusion that 'MAS uniformly improves accuracy' is not supported by the metrics reported in Table 3. For GemF, Rule A gives sensitivity 0.988 versus 0.991 for the single agent and WSS 0.865 versus 0.866. For Hope, Rule A ties sensitivity at 0.941 and slightly lowers WSS (0.824 versus 0.825). For DS, Rule A lowers specificity from 1.000 to 0.995 and PPV from 0.997 to 0.941. If 'uniform improvement' refers only to the accuracy column, the text should say so; otherwise the claim should be revised or the trade-offs explicitly discussed.
- [Section 4.2 and Appendix C.2] The abstract's claim that the system leads to 'updated clinical conclusions' is not reproducible from the manuscript. Section 4.2 reports P-scores and hazard ratios without specifying the NMA model, estimation method, prior distributions, consistency assumptions, or software. Appendix C.1 further states that many overall-survival hazard ratios were estimated by IPDfromKM or SynthIPD from Kaplan-Meier curves rather than by the extraction MAS, and Appendix C.2 explicitly acknowledges that P-scores for biomarker-selected regimens 'may be inflated' because such trials are pooled with unselected populations. The authors should provide the full NMA specification, the reconstructed hazard-ratio inputs, and sensitivity analyses excluding or stratifying reconstructed and biomarker-restricted data; otherwise the 'updated clinical conclusions' claim should be removed or substantially weakened.
- [Section 2.3 versus Section 3.2] The statement that 'our extraction MAS with the same LLM model recovers all of Table 2 correctly' is not demonstrated. The one-shot baseline in Table 2 uses Gemini 3.1 Pro on the IPSOS trial, while the MAS extraction evaluation in Section 3.2 uses GemF, Hope, and DS on different trials and endpoints. No controlled MAS-versus-SAS extraction comparison on the same inputs is reported. The authors should either add a like-for-like comparison (same model, same trials, same query) or explicitly label the Table 2 contrast as an illustrative example rather than an experimental result.
minor comments (5)
- [Section 3.1 and Table 3] The text states that 2,245 trials were screened, but the Table 3 caption reports n = 2,236; these numbers should be reconciled.
- [Section 3.1] Fleiss' Kappa of 0.70 is described as 'moderate'; standard conventions usually classify values in 0.61-0.80 as substantial, so the wording should be adjusted or justified.
- [Table 5] Table 5 contains 19 rows while the text says 17 newly identified studies; the table should make explicit that some trials contribute multiple treatment comparisons and should report a unique-study count.
- [Appendix C.1] The sentence '13 (resp. 14) included overall survival results and included in our analysis' is grammatically unclear and should state the counts for original and newly identified trials separately.
- [Data availability statement] The data availability statement names only public data sources, while the CRC answer sheet in Appendix A.3.1 is 'available upon request'; the authors should state where the MAS workflow code, prompts, and analysis scripts will be deposited.
Circularity Check
No significant circularity: MAS evaluation is anchored to external human-consensus and source-publication benchmarks.
full rationale
The load-bearing claims in this paper are evaluated against external benchmarks rather than against the system's own outputs. The screening MAS is scored against a human-consensus gold standard constructed by three blinded reviewers (Section 3.1), with the consensus reached before comparison; flagged-case adjudication is reported separately (Rule A vs Rule B), so the accuracy gain is not definitionally forced. Extraction accuracy is judged by exact match to primary publications (Section 3.2), an external source; confidence labels are used to stratify, not to define, correctness. The NMA application uses the original Xu et al. (2021) inclusion criteria and published trial list as external reference points, and the HR reconstruction via IPDfromKM/SynthIPD is a cited tool-citation rather than a derivation from the MAS outputs; the model specification gap and the possible P-score inflation flagged in C.2 are reproducibility/correctness concerns, not circularity. No equation, fitted parameter, or benchmark is renamed as a prediction, and no load-bearing premise reduces to a self-citation. Therefore no circular step is identified.
Assumptions & free parameters
free parameters (5)
- Number of screening agents N =
5 in evaluation, 6 in NMA application
- Maximum revision rounds T =
3 in screening evaluation, 5 in NMA application, 3 in extraction
- Sampling temperature =
0.7
- top_p =
0.9
- Confidence label thresholds for extraction =
ceil(T/2) <= T' < T for Medium; T' < ceil(T/2) for Low
assumptions (5)
- domain assumption The three-reviewer consensus after discussion is the correct benchmark for screening eligibility
- domain assumption Heterogeneous personas reduce LLM mode collapse and improve decision diversity
- domain assumption Repeated extraction agreement (self-consistency) is a valid proxy for extraction correctness
- domain assumption HRs reconstructed from Kaplan-Meier plots via IPDfromKM and SynthIPD are unbiased
- domain assumption The adopted network meta-analysis model and connectivity rules are appropriate
Cite this review
Pith. "Pith review of Systematic Literature Reviews With Two Multi-Agentic Systems And Human-In-The-Loop." pith.science (2026). https://pith.science/paper/OF4UE5FZ
@misc{pith2026260721920,
author = {Pith},
title = {Pith review of: Systematic Literature Reviews With Two Multi-Agentic Systems And Human-In-The-Loop},
year = {2026},
howpublished = {\url{https://pith.science/paper/OF4UE5FZ}},
note = {Machine review of arXiv:2607.21920}
}
read the original abstract
Systematic literature review of clinical trials drives regulatory decision-making, but conventional screening and extraction are time-consuming, labor-intensive, and vulnerable to study selection bias. We propose two fit-to-purpose multi-agentic systems (MAS) for systematic literature review, with human-in-the-loop. The screening MAS uses multiple LLM agents with heterogeneous personas and multiround cross-review, and uniformly improves accuracy over a single-LLM baseline. The extraction MAS combines standardization, an iterative correction loop, and retrieval-based context control to ensure accuracy and scalability. Both MAS are specifically designed to support Human-In-The-Loop which is essential for clinical decisions. The novelty of the proposed approach lies in the system architecture rather than in any single foundation tools: the system can naturally benefit from future improvements in the underlying tools, for instance, stronger LLM agents, retrieval engines, image recognition methods, etc. As a real-world application, a published network meta-analysis is reproduced by the MAS. The result recovers all trials from the original study and identifies additional eligible trials missed by manual review, leading to updated clinical conclusions.
Figures
Reference graph
Works this paper leans on
-
[8]
Clara Montagut Viladot and Eva Martinez Balibrea. Genotype-based selection of treatment for patients with advanced colorectal cancer (seticc): a pharmacogenetic-based randomized phase ii trial.ANNALS OF ONCOLOGY, 2018, vol. 29, núm. 2, p. 439-444,
work page 2018
-
[10]
Accessed: 2026-05-05. Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdh- ery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171,
arXiv 2026
-
[11]
Zhang, Tim Kraska, and Omar Khattab
Alex L. Zhang, Tim Kraska, and Omar Khattab. Recursive language models, 2025a. URL https: //arxiv.org/abs/2512.24601. Jiayi Zhang, Simon Yu, Derek Chong, Anthony Sicilia, Michael R Tomz, Christopher D Manning, and Weiyan Shi. Verbalized sampling: How to mitigate mode collapse and unlock llm diversity. arXiv preprint arXiv:2510.01171, 2025b. Zixuan Zhao, Z...
-
[1999]
T Aparicio, S Lavau-Denes, JM Phelip, E Maillard, JL Jouve, D Gargot, M Gasmi, C Locher, X Adhoute, P Michel, et al. Randomized phase iii trial in elderly patients comparing lv5fu2 with or without irinotecan for first-line treatment of metastatic colorectal cancer (ffcd 2001–02).Annals of Oncology, 27(1):121–127,
work page 2001
-
[2005]
Large language models streamline automated systematic review: A preliminary study
Xi Chen and Xue Zhang. Large language models streamline automated systematic review: A preliminary study.arXiv preprint arXiv:2502.15702,
-
[2013]
Michel Ducreux, David Malka, Jean Mendiboure, Pierre-Luc Etienne, Patrick Texereau, Dominique Auby, Philippe Rougier, Mohamed Gasmi, Marine Castaing, Moncef Abbas, et al. Sequential versus combination chemotherapy for the treatment of advanced colorectal cancer (ffcd 2000–05): an open-label, randomised, phase 3 trial.The lancet oncology, 12(11):1032–1044,
work page 2000
-
[2017]
Automation of systematic reviews with large language models.medRxiv, pages 2025–06, 2025a
Christian Cao, Rohit Arora, Paul Cento, Katherine Manta, Elina Farahani, Matthew Cecere, Anabel Selemon, Jason Sang, Ling Xi Gong, Robert Kloosterman, et al. Automation of systematic reviews with large language models.medRxiv, pages 2025–06, 2025a. Christian Cao, Jason Sang, Rohit Arora, David Chen, Robert Kloosterman, Matthew Cecere, Jaswanth Gorla, Rich...
work page 2025
-
[2019]
The prisma 2020 statement: an updated guideline for reporting systematic reviews.bmj, 372,
24 Matthew J Page, Joanne E McKenzie, Patrick M Bossuyt, Isabelle Boutron, Tammy C Hoffmann, Cynthia D Mulrow, Larissa Shamseer, Jennifer M Tetzlaff, Elie A Akl, Sue E Brennan, et al. The prisma 2020 statement: an updated guideline for reporting systematic reviews.bmj, 372,
work page 2020
Show all 11 references
-
[2023]
Enhancing systematic literature review with advanced ai: A study based on llm screening
Dan Li, Leihong Wu, Svitlana Shpyleva, Ting Li, and Joshua Xu. Enhancing systematic literature review with advanced ai: A study based on llm screening. Technical report, FDA 2024 Digital Transformation Symposium,
2024
-
[2025]
Exploring large language model based intelligent agents: Definitions, methods, and prospects.arXiv preprint arXiv:2401.03428,
Yuheng Cheng, Ceyao Zhang, Zhengwen Zhang, Xiangrui Meng, Sirui Hong, Wenhao Li, Zihao Wang, Zekai Wang, Feng Yin, Junhua Zhao, et al. Exploring large language model based intelligent agents: Definitions, methods, and prospects.arXiv preprint arXiv:2401.03428,
-
[2026]
Ying Li, Surabhi Datta, Majid Rastegar-Mojarad, Kyeryoung Lee, Hunki Paek, Julie Glasgow, Chris Liston, Long He, Xiaoyan Wang, and Yingxin Xu
doi: 10.1017/rsm.2025.10065. Ying Li, Surabhi Datta, Majid Rastegar-Mojarad, Kyeryoung Lee, Hunki Paek, Julie Glasgow, Chris Liston, Long He, Xiaoyan Wang, and Yingxin Xu. Enhancing systematic literature reviews with generative artificial intelligence: development, application...
2025
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.