{"id":"9de04dec-c497-4535-8204-3fb9a2a9329c","arxiv_id":"2505.17065","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of LLMs for rare disease diagnosis that catalogs datasets, ontologies, and challenges, but whose promised original experiment is absent and whose study counts do not add up.","lead":"This paper is a review of how large language models are being used to help diagnose rare diseases, covering 19 studies from 2022 to 2024. It promises an original experiment comparing several LLMs on diagnostic questionnaires, but that experiment is missing from the text.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Advertised original experimentation is entirely absent; the conclusion's 'promising results' claim has no supporting protocol, data, or results in the manuscript.","rationale":"I concur with the reader's REJECT verdict. The most load-bearing issue is not the PRISMA arithmetic (though that is a real defect) but the complete absence of the experimentation promised in the abstract and cited in the conclusion. The abstract explicitly advertises 'a section on experimentation'; Section 7 asserts that this experimentation 'showed promising results.' Neither the protocol nor the results appear anywhere in the paper. This makes the central empirical claim unfalsifiable and unreproducible. The reader's stated weakest_assumption focused on PRISMA selection consistency; I agree that the 30-21=19 accounting error is serious, but it affects the survey's representativeness, whereas the missing experimentation directly voids a stated contribution. Fixing the counts would not repair the experimentation claim. Therefore, as submitted, the paper cannot support its conclusion, and rejection is appropriate. I would encourage the authors to either remove all claims of original experimentation or provide the full experimental section with model names, questionnaire items, evaluation metrics, and results.","tokens_in":14550,"tokens_out":4198,"duration_ms":40288,"concrete_test":"Search the manuscript for any section or subsection that would constitute the promised 'section on experimentation': identify original use of multiple LLMs with structured questionnaires for diagnosis, including named models, prompt templates, data collection, and results. Inspect the table of contents and section headings; if no such section exists between Sections 2 and 7, the Section 7 claim is unsupported. Additionally, attempt to reconstruct the experimental results from any tables or text; if no quantitative outcomes are reported, the claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim, as stated in the abstract and Section 7, includes an original experimentation section that 'utilizes multiple LLMs alongside structured questionnaires' and yielded 'promising results' for diagnosis. The full text contains no such section. The section structure runs Introduction (1), PRISMA methodology (2), literature classification (3), datasets (4), challenges (5), future perspectives (6), and Conclusion (7). None of these describes an original experimental protocol, model configuration, questionnaire design, evaluation metrics, or numerical results. Section 3 describes prompting strategies used in the reviewed studies, not the authors' own experiments. Consequently, the assertion 'Our experimentation ... showed promising results' is a claim without derivation. This is not a matter of external validity or consensus; it is an internal inconsistency between the declared contributions and the manuscript's content. Even if the PRISMA accounting were corrected to 9 or 19, the missing experimentation would remain an unsupported central claim. For a systematic review, the experimental claim is superfluous if removed, but as submitted it is advertised as a contribution and used to support the conclusion. A reader cannot verify, replicate, or evaluate the diagnostic promise claimed.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is presented as a PRISMA systematic review of large language models (LLMs) applied to the diagnosis of rare diseases, with an additional advertised original experimentation section. The paper describes the PRISMA selection procedure, sketches a classification framework for the literature, surveys datasets and ontologies, lists challenges, and offers future perspectives. The abstract promises a section on experimentation that 'utilizes multiple LLMs alongside structured questionnaires' for diagnosis, and the conclusion states that 'Our experimentation with different LLMs... showed promising results regarding their potential to assist in diagnosis.' However, the full text contains no such experimental section, no protocol, no data, and no results. In addition, the PRISMA accounting in Section 2 is arithmetically inconsistent, the list of the 19 supposedly included studies is not provided, and several citation-supported claims rest on references that do not address the cited topics. The manuscript therefore does not currently support its stated central contributions.","tokens_in":14626,"tokens_out":6027,"duration_ms":62557,"significance":"If the advertised systematic review and experimentation were actually present and correct, the paper would be a useful contribution to a growing area: it would map the current text-only focus of LLM work in rare diseases, identify multimodal integration as the frontier, and assemble a helpful inventory of datasets, ontologies, and limitations. The paper has some genuinely informative components, notably Table 1's summaries of four disease-specific studies and Table 5's structured list of limitations. The difficulty is that these components do not compensate for the absence of the promised experiments and the irreproducibility of the review's selection. As submitted, the central claims about 'promising results' and about what 'the selected literature' shows are unsupported, so the paper's significance cannot be assessed beyond its descriptive parts.","major_comments":[{"comment":"The abstract promises 'a section on experimentation that utilizes multiple LLMs alongside structured questionnaires, specifically designed for diagnostic purposes,' but no such section exists anywhere in the manuscript. Section 7 then asserts, 'Our experimentation with different LLMs, however, showed promising results regarding their potential to assist in diagnosis,' without providing any protocol, model names, questionnaire design, data, evaluation metrics, or numerical results. This is an internal inconsistency between the declared contributions and the actual content, and the diagnostic promise claim is therefore unverifiable and unreproducible.","section":"Abstract and Section 7"},{"comment":"The PRISMA accounting cannot be reconstructed. Section 2 states that 30 studies met the initial inclusion criteria and were assessed by full-text review, that 21 articles were excluded during this phase, and that 19 studies were ultimately included; however, 30 - 21 = 9, not 19. The manuscript also does not provide the list of the 19 included studies; Table 1 summarizes only four disease-specific studies, and Section 3's classification framework is not applied to any enumerated set of 19 papers. In addition, the stated exclusion criterion 'lack of peer review' is contradicted by the inclusion of reference [28], a medRxiv preprint, in Table 1. These problems make it impossible to verify which studies the survey's generalizations are based on.","section":"Section 2 (PRISMA flow)"},{"comment":"Several load-bearing citation claims are unsupported by the cited references. The claim that MIMIC-III has been 'extensively adopted for patient phenotyping, disease classification, and predictive modelling' cites [17], which is a book on biological network analysis, not a MIMIC-III adoption study. The corresponding claim about MIMIC-IV cites [21], a survey on disease spreading modeling, which likewise does not discuss MIMIC-IV. In the Introduction, the need for a systematic survey is supported in part by [34], a paper on age and gender differences in SARS-CoV-2 outcomes, which is not relevant to that point. Because these citations do not support the assertions they accompany, the survey's factual grounding is in need of systematic verification.","section":"Section 4.1 and Introduction (References [17], [21], [34])"},{"comment":"The proposed classification framework is described in general terms but never actually applied to the included studies. Section 3 repeats the framework description almost verbatim twice, includes an unresolved 'Table ??' cross-reference, and presents no completed classification of the reviewed papers. Statements such as 'In all the articles analyzed, the models were tested in closed environments and independently' are based on only the four studies in Table 1, not on the 19 studies claimed to be included. The survey's synthesis is therefore not supported by the evidence actually presented.","section":"Section 3 (classification framework)"}],"minor_comments":[{"comment":"The sentence 'affecting an estimated 3.5–5.9' is missing the unit and the citation; it should read '3.5%–5.9% of the global population,' as correctly stated later in Section 1.1 with reference [37].","section":"Section 1 (first paragraph)"},{"comment":"A paragraph beginning 'Research into the use of Large Language Models...' is duplicated almost verbatim, and the cross-reference to the framework table remains as 'Table ??'. Please remove the duplication and fix the cross-reference.","section":"Section 3"},{"comment":"The phrase 'lunar avascular necrosis' should be 'lunate avascular necrosis.'","section":"Section 3.0.1"},{"comment":"The sentence 'It is [43] a social network structured around various topics-focused forums' is ungrammatical; consider revising to 'Reddit [43] is a social network structured around topic-focused forums.'","section":"Section 4.4"},{"comment":"The sentence 'According to Hasani et al., this kind of cooperation is essential...' cites no reference number; if reference [19] is intended, it should be cited explicitly at that point.","section":"Section 7"},{"comment":"Tables 2, 3, and 4 are presented without in-text callouts in the running text; please add explicit references to each table in the relevant subsection.","section":"Section 4"}],"recommendation":"reject","confidential_remarks":"This manuscript has a central internal inconsistency: the abstract and conclusion advertise an original experimentation section that does not exist, and the PRISMA selection accounting is irreparably unclear. I would not invite a revision unless the authors either add a genuine experimental section with full methods and data or explicitly resubmit as a survey-only paper after correcting the PRISMA numbers, providing the list of included studies, and fixing the citation errors involving references [17], [21], and [34]. The self-citation pattern in those references is worth checking carefully before any resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: The paper is a systematic review of LLMs in rare-disease diagnosis that advertises an original experimentation section which simply is not in the manuscript. The abstract and conclusion both say the authors' own experiments with multiple LLMs and structured questionnaires 'showed promising results,' but no experiment, protocol, or data is presented anywhere. That unsupported claim is the load-bearing problem.\n\nWhat's actually new and good: the rare-disease focus is a real gap in the medicine-LLM survey literature, and the descriptive content is largely accurate. The dataset summaries (MIMIC-III/IV, RareDis, RAMEDIS) and ontology descriptions (HPO, Orphanet, ORDO, OMIM) are faithful to the sources. The framing of the field as mostly text-only with multimodal integration as the open frontier is reasonable. If the study list were complete and verifiable, this would be a useful map.\n\nSoft spots, in order of severity. One: the missing experimental claim. It is not a missing optional extra; the conclusion explicitly relies on it. A reader cannot verify or evaluate a diagnostic promise that has no associated methods. Two: the PRISMA accounting is off by ten. The text says 30 studies were assessed by full-text review, 21 were excluded, and 19 were included. 30 minus 21 is 9. Either the numbers are wrong or the included list is incomplete; either way, the synthesis rests on unreported selections. Three: Section 3 repeats the same paragraphs nearly verbatim, with dangling cross-references like 'Table ??' and an unresolved citation. That suggests a lack of copyediting that matters more than it would elsewhere because the paper presents itself as a systematic review.\n\nThe citation pattern is mostly normal. There are a few self-citations used for background claims about MIMIC and disease modeling; that alone is not a problem, but given the other issues it doesn't inspire confidence.\n\nBottom line: this paper deserves a serious referee, not a desk reject, because the topic is important and the survey material is real. But the referee should treat the phantom experiments and the numerical inconsistency as blocking. I would not cite it in its current form. If the authors remove the experimental claim, correct the PRISMA counts, list the 19 included studies, and cut the duplicate text, it could be a solid contribution.","headline":"A useful rare-disease survey undermined by a phantom experimental claim and PRISMA arithmetic that doesn't add up.","tokens_in":15287,"tokens_out":3694,"would_cite":false,"duration_ms":36310,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A systematic review of 19 studies claims LLMs can assist rare-disease diagnosis from text, while true multimodal integration remains the open frontier.","keywords":["large language models","rare diseases","diagnosis","systematic review","multimodal data","questionnaires","clinical natural language processing"],"falsifier":"Ask the authors to release the list of the 19 included studies and the prompts, questionnaires, and outputs behind their claimed experiments; if the counts cannot be reconstructed (30 assessed by full-text review minus 21 excluded is 9, not 19) or the experimental results cannot be reproduced, the review's empirical generalizations lose their basis. A simpler check is to rerun the stated literature search with the stated keywords and date range and see whether the same 19 studies emerge.","tokens_in":14232,"feed_emoji":"🩺","tokens_out":11723,"duration_ms":101200,"temperature":0.7,"pith_summary":"The paper is a systematic review that sets out to establish what large language models (LLMs) currently can and cannot do for the diagnosis of rare diseases. Following a systematic-review protocol, it selects 19 studies and sorts them into two methodological camps: studies that test LLMs with structured questionnaires and studies that use LLMs to extract or synthesize knowledge from unstructured text. Across the selected literature, it finds that the models are always used standalone, closed-source, and on a single input modality, never combining genetic, imaging, or electronic health record data. The conclusion asserts that the authors' own experiments with multiple LLMs and structured questionnaires showed promising diagnostic results, but the body of the paper does not report those experiments. A sympathetic reader would take the paper's thesis to be that LLMs are already plausible assistive tools for rare-disease diagnosis from text, and that multimodal integration is the necessary next step.","feed_headline":"LLM diagnosis of rare diseases: text works, multimodal is next","feed_subtitle":"Review of 19 studies: LLMs help diagnose rare diseases from text; genomic and imaging data remain untapped.","key_machinery":"The machinery that carries the argument is a systematic literature selection plus a five-dimension classification framework. The selection process is what lets the authors speak about 'the selected literature' as a single corpus; the classification framework (disease focus, study objective, input data modality, LLM type and access, and pipeline role) is what turns that corpus into a diagnosis of the field. The last two dimensions do the heaviest lifting: because every selected study uses a closed-source model in a standalone, monomodal setting, the framework yields the paper's central generalization that LLM-based rare-disease diagnosis is not yet integrated with genomic, imaging, or electronic health record data.","core_discovery":"On the paper's own terms, the central claim is that current LLM applications in rare-disease diagnosis are promising but structurally narrow. The detailed cases the review examines—glaucoma, sarcoidosis, Kienböck's disease, and amyloidosis—all use closed-source models in standalone mode with a single input type, either a structured questionnaire or raw patient-generated text. The review's classification framework organizes the field along five dimensions (disease focus, objective, input modality, model type and access, and pipeline role), and every selected study lands on the same side of the last two dimensions, which is what supports the conclusion that no study yet realizes a multimodal pipeline. The paper also asserts, in its conclusion, that the authors' own experiments with multiple LLMs and structured questionnaires produced promising results for diagnostic assistance, with genomic, imaging, laboratory, and longitudinal patient data named as the field's open frontier.","pith_inferences":["The authors leave implicit that their 'diagnostic odyssey' framing suggests a concrete outcome measure, time-to-diagnosis; a natural extension would compare time-to-diagnosis in cohorts where an LLM triaged the initial patient text against standard care.","A cheap falsifiable benchmark would rerun the four questionnaire studies (glaucoma, sarcoidosis, Kienböck's disease, amyloidosis) with one genetic or imaging feature added per case and measure whether diagnostic accuracy changes.","The review's focus on closed-source models implies a reproducibility hazard: later API versions may not reproduce published accuracy numbers, so an open-weight replication of the same questionnaires would give more durable evidence.","Because the claimed experiments appear only in the conclusion and lack a methods or results section, the fair reading is that the authors intend them as a pointer for future work rather than as established evidence."],"forward_implications":["If the review is right, clinicians considering an LLM for rare-disease diagnosis today should expect a text-only, closed-source assistant validated mainly through questionnaires, not a tool embedded in clinical workflow.","The dominance of monomodal studies implies that the next measurable gains will come from datasets and pipelines that pair clinical text with genetic variants, imaging, and structured laboratory results.","The proposed classification framework gives future evaluations a shared vocabulary, allowing new LLM studies to be compared on disease focus, input modality, model access, and pipeline role rather than as isolated accuracy figures.","Because every reviewed study is standalone and closed-source, reproducibility and auditability become the first governance issues for clinical deployment.","Taken at face value, the authors' own 'promising results' point to near-term use in triage and patient-facing explanation rather than autonomous diagnosis."],"supporting_citations":[{"why":"Supplies the glaucoma and retina GPT-4 evaluation, the review's primary questionnaire-based study.","marker":"[22]"},{"why":"Supplies the sarcoidosis social-media mining study, the review's primary unstructured-text example.","marker":"[61]"},{"why":"Supplies the Kienböck's disease FAQ study, used to show high accuracy but limited readability.","marker":"[3]"},{"why":"Supplies the amyloidosis ChatGPT knowledge evaluation, a load-bearing questionnaire study.","marker":"[28]"},{"why":"RareBench, cited as evidence of LLM difficulties on rare diseases and of dynamic few-shot prompting as a remedy.","marker":"[9]"},{"why":"Qualitative study of generative AI assistance for rare and complex diagnoses, supporting the diagnosis-assistance claim.","marker":"[1]"},{"why":"Zebra-LLaMA, classified as a monomodal standalone LLM for rare-disease knowledge.","marker":"[50]"},{"why":"Cited for the argument that multimodal AI frameworks can transform rare-disease diagnostics by combining diverse data sources.","marker":"[42]"}],"fun_headline_variants":["LLMs crack rare diseases from text, multimodal next","Rare-disease LLMs: text-based now, multimodal future","Survey: LLM rare-disease diagnosis limited to text data","Text-only LLMs help rare diseases, imaging untapped","Rare disease diagnosis via LLM: text works, multimodal not yet"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the 19 studies the selection process nominally included are the relevant, representative literature on LLMs and rare-disease diagnosis; the text cannot substantiate that because its own counts (30 assessed by full-text review, 21 excluded, 19 included) do not reconcile, and the claimed promising experiments are asserted in the conclusion without a reported method or results.","fun_headline_variants_meta":{"raw":{"variants":["LLMs crack rare diseases from text, multimodal next","Rare-disease LLMs: text-based now, multimodal future","Survey: LLM rare-disease diagnosis limited to text data","Text-only LLMs help rare diseases, imaging untapped","Rare disease diagnosis via LLM: text works, multimodal not yet"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000193,"raw_usage":{"total_tokens":1347,"prompt_tokens":936,"completion_tokens":411,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":552,"completion_tokens_details":{"reasoning_tokens":325}},"tokens_in":552,"tokens_out":411,"duration_ms":4137,"temperature":1.0,"reasoning_tokens":325,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:33:39.564034+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Ask the authors to release the list of the 19 included studies and the prompts, questionnaires, and outputs behind their claimed experiments; if the counts cannot be reconstructed (30 assessed by full-text review minus 21 excluded is 9, not 19) or the experimental results cannot be reproduced, the review's empirical generalizations lose their basis. A simpler check is to rerun the stated literature search with the stated keywords and date range and see whether the same 19 studies emerge.","supporting_citations":[{"cited_title":"Assessment of a large language model’s responses to questions and cases about glaucoma and retina management","cited_arxiv_id":null,"evidence_quote":"Supplies the glaucoma and retina GPT-4 evaluation, the review's primary questionnaire-based study."},{"cited_title":"Understanding sarcoidosis using large language models and social media data","cited_arxiv_id":null,"evidence_quote":"Supplies the sarcoidosis social-media mining study, the review's primary unstructured-text example."},{"cited_title":"High accuracy but limited readability of large language model-generated responses to fre- quently asked questions about kienb¨ ock’s disease","cited_arxiv_id":null,"evidence_quote":"Supplies the Kienböck's disease FAQ study, used to show high accuracy but limited readability."},{"cited_title":"A multidisciplinary assessment of chat- gpt’s knowledge of amyloidosis","cited_arxiv_id":null,"evidence_quote":"Supplies the amyloidosis ChatGPT knowledge evaluation, a load-bearing questionnaire study."},{"cited_title":"Rarebench: Can llms serve as rare diseases specialists? In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , pages 4850–4861, 2024","cited_arxiv_id":null,"evidence_quote":"RareBench, cited as evidence of LLM difficulties on rare diseases and of dynamic few-shot prompting as a remedy."},{"cited_title":"Learning to make rare and complex diagnoses with generative ai assistance: quali- tative study of popular large language models","cited_arxiv_id":null,"evidence_quote":"Qualitative study of generative AI assistance for rare and complex diagnoses, supporting the diagnosis-assistance claim."},{"cited_title":"Artificial intelligence empowering rare diseases: a bibliometric perspective over the last two decades","cited_arxiv_id":null,"evidence_quote":"Cited for the argument that multimodal AI frameworks can transform rare-disease diagnostics by combining diverse data sources."}],"review_version":1}