Pith. sign in

REVIEW 4 major objections 6 minor 66 references

Decoding Rarity: Large Language Models in the Diagnosis of Rare Diseases

T0 review · 4 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read A systematic review of 19 studies claims LLMs can assist rare-disease diagnosis from text, while true multimodal integration remains the open frontier.

desk verdict A useful rare-disease survey undermined by a phantom experimental claim and PRISMA arithmetic that doesn't add up. read the letter →

arxiv 2505.17065 v1 pith:K5LETQPT submitted 2025-05-18 cs.CL cs.LG

classification cs.CLcs.LG
keywords largelanguagemodelsrarediseasesdiagnosissystematicreviewmultimodaldataquestionnairesclinicalnaturalprocessing
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper is a systematic review that sets out to establish what large language models (LLMs) currently can and cannot do for the diagnosis of rare diseases. Following a systematic-review protocol, it selects 19 studies and sorts them into two methodological camps: studies that test LLMs with structured questionnaires and studies that use LLMs to extract or synthesize knowledge from unstructured text. Across the selected literature, it finds that the models are always used standalone, closed-source, and on a single input modality, never combining genetic, imaging, or electronic health record data. The conclusion asserts that the authors' own experiments with multiple LLMs and structured questionnaires showed promising diagnostic results, but the body of the paper does not report those experiments. A sympathetic reader would take the paper's thesis to be that LLMs are already plausible assistive tools for rare-disease diagnosis from text, and that multimodal integration is the necessary next step.

What carries the argument

The machinery that carries the argument is a systematic literature selection plus a five-dimension classification framework. The selection process is what lets the authors speak about 'the selected literature' as a single corpus; the classification framework (disease focus, study objective, input data modality, LLM type and access, and pipeline role) is what turns that corpus into a diagnosis of the field. The last two dimensions do the heaviest lifting: because every selected study uses a closed-source model in a standalone, monomodal setting, the framework yields the paper's central generalization that LLM-based rare-disease diagnosis is not yet integrated with genomic, imaging, or electronic health record data.

What would settle it

Ask the authors to release the list of the 19 included studies and the prompts, questionnaires, and outputs behind their claimed experiments; if the counts cannot be reconstructed (30 assessed by full-text review minus 21 excluded is 9, not 19) or the experimental results cannot be reproduced, the review's empirical generalizations lose their basis. A simpler check is to rerun the stated literature search with the stated keywords and date range and see whether the same 19 studies emerge.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central claim is that current LLM applications in rare-disease diagnosis are promising but structurally narrow. The detailed cases the review examines—glaucoma, sarcoidosis, Kienböck's disease, and amyloidosis—all use closed-source models in standalone mode with a single input type, either a structured questionnaire or raw patient-generated text. The review's classification framework organizes the field along five dimensions (disease focus, objective, input modality, model type and access, and pipeline role), and every selected study lands on the same side of the last two dimensions, which is what supports the conclusion that no study yet realizes a multimodal pipeline. The paper also asserts, in its conclusion, that the authors' own experiments with multiple LLMs and structured questionnaires produced promising results for diagnostic assistance, with genomic, imaging, laboratory, and longitudinal patient data named as the field's open frontier.

Load-bearing premise

The load-bearing assumption is that the 19 studies the selection process nominally included are the relevant, representative literature on LLMs and rare-disease diagnosis; the text cannot substantiate that because its own counts (30 assessed by full-text review, 21 excluded, 19 included) do not reconcile, and the claimed promising experiments are asserted in the conclusion without a reported method or results.

Editorial extensions

If this is right

  • If the review is right, clinicians considering an LLM for rare-disease diagnosis today should expect a text-only, closed-source assistant validated mainly through questionnaires, not a tool embedded in clinical workflow.
  • The dominance of monomodal studies implies that the next measurable gains will come from datasets and pipelines that pair clinical text with genetic variants, imaging, and structured laboratory results.
  • The proposed classification framework gives future evaluations a shared vocabulary, allowing new LLM studies to be compared on disease focus, input modality, model access, and pipeline role rather than as isolated accuracy figures.
  • Because every reviewed study is standalone and closed-source, reproducibility and auditability become the first governance issues for clinical deployment.
  • Taken at face value, the authors' own 'promising results' point to near-term use in triage and patient-facing explanation rather than autonomous diagnosis.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The authors leave implicit that their 'diagnostic odyssey' framing suggests a concrete outcome measure, time-to-diagnosis; a natural extension would compare time-to-diagnosis in cohorts where an LLM triaged the initial patient text against standard care.
  • A cheap falsifiable benchmark would rerun the four questionnaire studies (glaucoma, sarcoidosis, Kienböck's disease, amyloidosis) with one genetic or imaging feature added per case and measure whether diagnostic accuracy changes.
  • The review's focus on closed-source models implies a reproducibility hazard: later API versions may not reproduce published accuracy numbers, so an open-weight replication of the same questionnaires would give more durable evidence.
  • Because the claimed experiments appear only in the conclusion and lack a methods or results section, the fair reading is that the authors intend them as a pointer for future work rather than as established evidence.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This manuscript is presented as a PRISMA systematic review of large language models (LLMs) applied to the diagnosis of rare diseases, with an additional advertised original experimentation section. The paper describes the PRISMA selection procedure, sketches a classification framework for the literature, surveys datasets and ontologies, lists challenges, and offers future perspectives. The abstract promises a section on experimentation that 'utilizes multiple LLMs alongside structured questionnaires' for diagnosis, and the conclusion states that 'Our experimentation with different LLMs... showed promising results regarding their potential to assist in diagnosis.' However, the full text contains no such experimental section, no protocol, no data, and no results. In addition, the PRISMA accounting in Section 2 is arithmetically inconsistent, the list of the 19 supposedly included studies is not provided, and several citation-supported claims rest on references that do not address the cited topics. The manuscript therefore does not currently support its stated central contributions.

Significance. If the advertised systematic review and experimentation were actually present and correct, the paper would be a useful contribution to a growing area: it would map the current text-only focus of LLM work in rare diseases, identify multimodal integration as the frontier, and assemble a helpful inventory of datasets, ontologies, and limitations. The paper has some genuinely informative components, notably Table 1's summaries of four disease-specific studies and Table 5's structured list of limitations. The difficulty is that these components do not compensate for the absence of the promised experiments and the irreproducibility of the review's selection. As submitted, the central claims about 'promising results' and about what 'the selected literature' shows are unsupported, so the paper's significance cannot be assessed beyond its descriptive parts.

major comments (4)
  1. [Abstract and Section 7] The abstract promises 'a section on experimentation that utilizes multiple LLMs alongside structured questionnaires, specifically designed for diagnostic purposes,' but no such section exists anywhere in the manuscript. Section 7 then asserts, 'Our experimentation with different LLMs, however, showed promising results regarding their potential to assist in diagnosis,' without providing any protocol, model names, questionnaire design, data, evaluation metrics, or numerical results. This is an internal inconsistency between the declared contributions and the actual content, and the diagnostic promise claim is therefore unverifiable and unreproducible.
  2. [Section 2 (PRISMA flow)] The PRISMA accounting cannot be reconstructed. Section 2 states that 30 studies met the initial inclusion criteria and were assessed by full-text review, that 21 articles were excluded during this phase, and that 19 studies were ultimately included; however, 30 - 21 = 9, not 19. The manuscript also does not provide the list of the 19 included studies; Table 1 summarizes only four disease-specific studies, and Section 3's classification framework is not applied to any enumerated set of 19 papers. In addition, the stated exclusion criterion 'lack of peer review' is contradicted by the inclusion of reference [28], a medRxiv preprint, in Table 1. These problems make it impossible to verify which studies the survey's generalizations are based on.
  3. [Section 4.1 and Introduction (References [17], [21], [34])] Several load-bearing citation claims are unsupported by the cited references. The claim that MIMIC-III has been 'extensively adopted for patient phenotyping, disease classification, and predictive modelling' cites [17], which is a book on biological network analysis, not a MIMIC-III adoption study. The corresponding claim about MIMIC-IV cites [21], a survey on disease spreading modeling, which likewise does not discuss MIMIC-IV. In the Introduction, the need for a systematic survey is supported in part by [34], a paper on age and gender differences in SARS-CoV-2 outcomes, which is not relevant to that point. Because these citations do not support the assertions they accompany, the survey's factual grounding is in need of systematic verification.
  4. [Section 3 (classification framework)] The proposed classification framework is described in general terms but never actually applied to the included studies. Section 3 repeats the framework description almost verbatim twice, includes an unresolved 'Table ??' cross-reference, and presents no completed classification of the reviewed papers. Statements such as 'In all the articles analyzed, the models were tested in closed environments and independently' are based on only the four studies in Table 1, not on the 19 studies claimed to be included. The survey's synthesis is therefore not supported by the evidence actually presented.
minor comments (6)
  1. [Section 1 (first paragraph)] The sentence 'affecting an estimated 3.5–5.9' is missing the unit and the citation; it should read '3.5%–5.9% of the global population,' as correctly stated later in Section 1.1 with reference [37].
  2. [Section 3] A paragraph beginning 'Research into the use of Large Language Models...' is duplicated almost verbatim, and the cross-reference to the framework table remains as 'Table ??'. Please remove the duplication and fix the cross-reference.
  3. [Section 3.0.1] The phrase 'lunar avascular necrosis' should be 'lunate avascular necrosis.'
  4. [Section 4.4] The sentence 'It is [43] a social network structured around various topics-focused forums' is ungrammatical; consider revising to 'Reddit [43] is a social network structured around topic-focused forums.'
  5. [Section 7] The sentence 'According to Hasani et al., this kind of cooperation is essential...' cites no reference number; if reference [19] is intended, it should be cited explicitly at that point.
  6. [Section 4] Tables 2, 3, and 4 are presented without in-text callouts in the running text; please add explicit references to each table in the relevant subsection.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the paper is a literature survey with no derivational chain; its missing experimentation section is an unsupported assertion, not a circular reduction.

full rationale

This manuscript is a PRISMA-style systematic review and does not present a quantitative derivation, fitted model, or prediction whose output could be equivalent to its inputs by construction. The central defect is that the abstract and conclusion advertise an original experimentation section — 'we present a section on experimentation that utilizes multiple LLMs alongside structured questionnaires' and 'Our experimentation with different LLMs, however, showed promising results regarding their potential to assist in diagnosis' — but no such protocol, data, or results appear in Sections 1-7. That is an internal consistency and completeness problem, not circularity: there is no equation, fitted parameter, or derivation chain to compare with a claimed output. The self-citations in the bibliography, including refs [17], [21], and [34] involving co-authors, are used only as general background support for statements about MIMIC adoption and the value of surveys; they are not invoked as uniqueness theorems, nor do they carry the paper's conclusions. The PRISMA accounting is also internally inconsistent (30 studies assessed by full-text review, 21 excluded, yet 19 reported as included), but arithmetic inconsistency is a reporting defect rather than input-output circularity. No load-bearing step in the survey reduces to its own inputs, so the circularity score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper is a review: it introduces no fitted parameters, no new model, and no new entities. Its load-bearing assumptions are methodological: that the reported PRISMA accounting is correct, though Section 2 reports 30 assessed, 21 excluded, and 19 included, which does not sum; that the 19 cited studies are representative and their results reliable, since no independent verification is performed; and that the promised original experimentation exists, since the conclusion cites its results but no experiment section appears.

assumptions (3)
  • domain assumption The 19 studies reported as included in the PRISMA review constitute a valid and representative sample of LLM research in rare diseases.
    Every general statement in the survey about 'the selected literature' inherits this assumption. Section 2's counts are internally inconsistent, with 30 assessed, 21 excluded, and 19 included, which cannot be reconciled, so the assumption that the review process was executed as reported is weakened.
  • domain assumption The accuracy and capability claims of the 19 cited studies are taken at face value.
    The survey performs no independent evaluation and re-runs no experiments; its conclusions about LLM diagnostic accuracy inherit the validity and error bars of the cited papers, such as the four disease-specific evaluations in Table 1.
  • domain assumption The premise that LLMs can meaningfully assist rare disease diagnosis is testable and true.
    This is the field-level premise the survey organizes evidence around. The paper's own intended test, the promised experimentation section, is absent from the full text, so the premise rests entirely on cited work rather than on anything demonstrated by this manuscript.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Decoding Rarity: Large Language Models in the Diagnosis of Rare Diseases." pith.science (2026). https://pith.science/paper/K5LETQPT

@misc{pith2026250517065,
  author       = {Pith},
  title        = {Pith review of: Decoding Rarity: Large Language Models in the Diagnosis of Rare Diseases},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/K5LETQPT}},
  note         = {Machine review of arXiv:2505.17065}
}
read the original abstract

Recent advances in artificial intelligence, particularly large language models LLMs, have shown promising capabilities in transforming rare disease research. This survey paper explores the integration of LLMs in the analysis of rare diseases, highlighting significant strides and pivotal studies that leverage textual data to uncover insights and patterns critical for diagnosis, treatment, and patient care. While current research predominantly employs textual data, the potential for multimodal data integration combining genetic, imaging, and electronic health records stands as a promising frontier. We review foundational papers that demonstrate the application of LLMs in identifying and extracting relevant medical information, simulating intelligent conversational agents for patient interaction, and enabling the formulation of accurate and timely diagnoses. Furthermore, this paper discusses the challenges and ethical considerations inherent in deploying LLMs, including data privacy, model transparency, and the need for robust, inclusive data sets. As part of this exploration, we present a section on experimentation that utilizes multiple LLMs alongside structured questionnaires, specifically designed for diagnostic purposes in the context of different diseases. We conclude with future perspectives on the evolution of LLMs towards truly multimodal platforms, which would integrate diverse data types to provide a more comprehensive understanding of rare diseases, ultimately fostering better outcomes in clinical settings.

Figures

Figures reproduced from arXiv: 2505.17065 by the authors.

Figure 1
Figure 1. Flowchart of the diagnostic journey for rare disease patients. The path [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. PRISMA flow diagram illustrating the systematic review process. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

66 extracted references · 61 canonical work pages

  1. [28]

    A multidisciplinary assessment of chat- gpt’s knowledge of amyloidosis

    Ryan C King, Jamil S Samaan, Yee Hui Yeo, David C Kunkel, Ali A Habib, and Roxana Ghashghaei. A multidisciplinary assessment of chat- gpt’s knowledge of amyloidosis. medRxiv, pages 2023–07, 2023

  2. [17]

    Biological network analysis: Trends, approaches, graph theory, and algorithms, 2020

    Pietro Hiram Guzzi and Swarup Roy. Biological network analysis: Trends, approaches, graph theory, and algorithms, 2020

  3. [21]

    Disease spreading modeling and analysis: A survey

    Pietro Hiram Guzzi, Francesco Petrizzelli, and Tommaso Mazza. Disease spreading modeling and analysis: A survey. Briefings in Bioinformatics , 23(4):bbac230, 2022. 19

  4. [34]

    Exploiting the molecular basis of age and gen- der differences in outcomes of sars-cov-2 infections

    Daniele Mercatelli, Elisabetta Pedace, Pierangelo Veltri, Federico M Giorgi, and Pietro Hiram Guzzi. Exploiting the molecular basis of age and gen- der differences in outcomes of sars-cov-2 infections. Computational and Structural Biotechnology Journal, 19:4092–4100, 2021

  5. [1]

    Learning to make rare and complex diagnoses with generative ai assistance: quali- tative study of popular large language models

    Tassallah Abdullahi, Ritambhara Singh, Carsten Eickhoff, et al. Learning to make rare and complex diagnoses with generative ai assistance: quali- tative study of popular large language models. JMIR Medical Education, 10(1):e51391, 2024

  6. [2]

    Evaluating llms for temporal entity extraction from pediatric clinical text in rare diseases context

    Judith Jeyafreeda Andrew, Marc Vincent, Anita Burgun, and Nicolas Garcelon. Evaluating llms for temporal entity extraction from pediatric clinical text in rare diseases context. In Proceedings of the First Workshop on Patient-Oriented Language Processing (CL4Health)@ LREC-COLING 2024, pages 145–152, 2024

  7. [3]

    High accuracy but limited readability of large language model-generated responses to fre- quently asked questions about kienb¨ ock’s disease

    Zeynel Mert Asfuro˘ glu, Hilal Ya˘ gar, and Ender G¨ um¨ u¸ so˘ glu. High accuracy but limited readability of large language model-generated responses to fre- quently asked questions about kienb¨ ock’s disease. BMC Musculoskeletal Disorders, 25(1):879, 2024

  8. [4]

    Development of a comprehensive heart disease knowl- edge questionnaire

    Hannah E Bergman, Bryce B Reeve, Richard P Moser, Sarah Scholl, and William MP Klein. Development of a comprehensive heart disease knowl- edge questionnaire. American journal of health education , 42(2):74–87, 2011

Show all 66 references
  1. [5]

    Duchenne muscular dystrophy: disease mechanism and therapeutic strategies

    Addeli Bez Batti Angulski, Nora Hosny, Houda Cohen, Ashley A Martin, Dongwoo Hahn, Jack Bauer, and Joseph M Metzger. Duchenne muscular dystrophy: disease mechanism and therapeutic strategies. Frontiers in physiology, 14:1183101, 2023

  2. [6]

    The unified medical language system (umls): inte- grating biomedical terminology

    Olivier Bodenreider. The unified medical language system (umls): inte- grating biomedical terminology. Nucleic acids research, 32(suppl 1):D267– D270, 2004

  3. [7]

    An automatic and end-to-end system for rare disease knowledge graph construction based on ontology- enhanced large language models: Development study

    Lang Cao, Jimeng Sun, Adam Cross, et al. An automatic and end-to-end system for rare disease knowledge graph construction based on ontology- enhanced large language models: Development study. JMIR Medical In- formatics, 12(1):e60665, 2024

  4. [8]

    Trends in rare disease drug development

    Rui Chen, Sen Liu, Jiashu Han, Shuhua Zhou, Yang Liu, Xiaoyuan Chen, and Shuyang Zhang. Trends in rare disease drug development. Nat Rev Drug Discov, 23(3):168–169, 2024

  5. [9]

    Rarebench: Can llms serve as rare diseases specialists? In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , pages 4850–4861, 2024

    Xuanzhong Chen, Xiaohao Mao, Qihan Guo, Lun Wang, Shuyang Zhang, and Ting Chen. Rarebench: Can llms serve as rare diseases specialists? In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , pages 4850–4861, 2024

  6. [10]

    Large language model capa- bilities in perioperative risk prediction and prognostication

    Philip Chung, Christine T Fong, Andrew M Walters, Nima Aghaeepour, Meliha Yetisgen, and Vikas N O’Reilly-Shah. Large language model capa- bilities in perioperative risk prediction and prognostication. JAMA surgery, 159(8):928–937, 2024. 18

  7. [11]

    The drug repurposing hub: a next-generation drug library and information resource

    Samantha M Corsello, Joshua A Bittker, Zihan Liu, John Gould, Patrick McCarren, Jeremy E Hirschman, Steven E Johnston, Andrej Vrcic, Bonnie Wong, Mojtaba Khan, et al. The drug repurposing hub: a next-generation drug library and information resource. Nature medicine, 23(4):405–...

  8. [12]

    Potential of large language models in health care: Del- phi study

    Kerstin Denecke, Richard May, LLMHealthGroup, and Octavio Rivera Romero. Potential of large language models in health care: Del- phi study. Journal of Medical Internet Research , 26:e52399, 2024

  9. [13]

    Assessing dxgpt: Diagnosing rare diseases with various large language models

    Juanjo do Olmo, Javier Logrono, Carlos Mascias, Marcelo Martinez, and Juli´ an Isla. Assessing dxgpt: Diagnosing rare diseases with various large language models. MedRxiv, pages 2024–05, 2024

  10. [14]

    Rare diseases

    European Commission. Rare diseases. https://health.ec.europa. eu/system/files/2023-01/rare_diseases_en_0.pdf, 2023. Accessed: 2025-05-01

  11. [15]

    Rare diseases - european commission, 2024

    European Commission. Rare diseases - european commission, 2024. Ac- cessed: 2025-05-04

  12. [16]

    Time to diagnosis and determinants of diagnostic delays of people living with a rare disease: re- sults of a rare barometer retrospective patient survey

    Fatoumata Faye, Claudia Crocione, Roberta Anido de Pe˜ na, Simona Bel- lagambi, Luciana Escati Pe˜ naloza, Amy Hunter, Lene Jensen, Cor Oost- erwijk, Eva Schoeters, Daniel de Vicente, et al. Time to diagnosis and determinants of diagnostic delays of people living with a rare d...

  13. [18]

    A systematic overview of rare disease patient reg- istries: challenges in design, quality management, and maintenance

    Isabel C Hageman, Iris ALM van Rooij, Ivo de Blaauw, Misel Trajanovska, and Sebastian K King. A systematic overview of rare disease patient reg- istries: challenges in design, quality management, and maintenance. Or- phanet Journal of Rare Diseases , 18(1):106, 2023

  14. [19]

    Artificial intelligence in medical imaging and its impact on the rare disease community: threats, challenges and oppor- tunities

    Navid Hasani, Faraz Farhadi, Michael A Morris, Moozhan Nikpanah, Ar- man Rhamim, Yanji Xu, Anne Pariser, Michael T Collins, Ronald M Sum- mers, Elizabeth Jones, et al. Artificial intelligence in medical imaging and its impact on the rare disease community: threats, challenges ...

  15. [20]

    A survey of large lan- guage models in medicine and healthcare.arXiv preprint arXiv:2306.13542, 2023

    Tianran He, Zihan Wang, Ying Liu, and Pengtao Xie. A survey of large lan- guage models in medicine and healthcare.arXiv preprint arXiv:2306.13542, 2023

  16. [22]

    Assessment of a large language model’s responses to questions and cases about glaucoma and retina management

    Andy S Huang, Kyle Hirabayashi, Laura Barna, Deep Parikh, and Louis R Pasquale. Assessment of a large language model’s responses to questions and cases about glaucoma and retina management. JAMA ophthalmology, 142(4):371–375, 2024

  17. [23]

    What disease does this patient have? a large-scale open domain question answering dataset from medical exams

    Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences, 11(14):6421, 2021

  18. [24]

    Pubmedqa: A dataset for biomedical research question an- swering

    Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W Cohen, and Xinghua Lu. Pubmedqa: A dataset for biomedical research question an- swering. arXiv preprint arXiv:1909.06146 , 2019

  19. [25]

    Mimic-iv, a freely accessible electronic health record dataset

    Alistair EW Johnson, Lucas Bulgarelli, Lu Shen, Alvin Gayles, Ayad Sham- mout, Steven Horng, Tom J Pollard, Sicheng Hao, Benjamin Moody, Brian Gow, et al. Mimic-iv, a freely accessible electronic health record dataset. Scientific data, 10(1):1, 2023

  20. [26]

    Mimic-iii, a freely accessible critical care database

    Alistair EW Johnson, Tom J Pollard, Lu Shen, Li-wei H Lehman, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo An- thony Celi, and Roger G Mark. Mimic-iii, a freely accessible critical care database. Scientific data, 3(1):1–9, 2016

  21. [27]

    Assessing the utility of large language models for phenotype-driven gene prioritization in the diagnosis of rare genetic disease

    Junyoung Kim, Kai Wang, Chunhua Weng, and Cong Liu. Assessing the utility of large language models for phenotype-driven gene prioritization in the diagnosis of rare genetic disease. The American Journal of Human Genetics, 111(10):2190–2202, 2024

  22. [29]

    The human phenotype ontology in 2021

    Sebastian K¨ ohler, Michael Gargano, Nicolas Matentzoglu, Leigh C Car- mody, David Lewis-Smith, Nicole A Vasilevsky, Daniel Danis, Ganna Bal- agura, Gareth Baynam, Amy M Brower, et al. The human phenotype ontology in 2021. Nucleic acids research, 49(D1):D1207–D1217, 2021

  23. [30]

    Biogpt: generative pre-trained transformer for biomedical text generation and mining

    Wonjin Lee, Hyunjae Sung, Jaewoo Kang, Woomyoung Yoon, Yelong Zhang, and Jianfeng Gao. Biogpt: generative pre-trained transformer for biomedical text generation and mining. Briefings in Bioinformatics , 24(1):bbac409, 2023

  24. [31]

    Treatment of amyloidosis: present and future

    Maria Teresa Mallus and Vittoria Rizzello. Treatment of amyloidosis: present and future. European Heart Journal Supplements , 25(Supple- ment B):B99–B103, 2023

  25. [32]

    The raredis corpus: a corpus anno- tated with rare diseases, their signs and symptoms

    Claudia Mart´ ınez-deMiguel, Isabel Segura-Bedmar, Esteban Chac´ on- Solano, and Sara Guerrero-Aspizua. The raredis corpus: a corpus anno- tated with rare diseases, their signs and symptoms. Journal of biomedical informatics, 125:103961, 2022. 20

  26. [33]

    Mendelian inheritance in man and its online version, omim

    Victor A McKusick. Mendelian inheritance in man and its online version, omim. The American Journal of Human Genetics , 80(4):588–604, 2007

  27. [35]

    Small data challenges of studying rare diseases

    Aya A Mitani and Sebastien Haneuse. Small data challenges of studying rare diseases. JAMA network open , 3(3):e201965–e201965, 2020

  28. [36]

    Estimating cumulative point prevalence of rare diseases: analysis of the orphanet database

    St´ ephanie Nguengang Wakap, Deborah M Lambert, Annie Olry, Char- lotte Rodwell, Charlotte Gueydan, Val´ erie Lanneau, Daniel Murphy, Yann Le Cam, and Ana Rath. Estimating cumulative point prevalence of rare diseases: analysis of the orphanet database. European journal of huma...

  29. [37]

    Lambert, Annie Olry, Char- lotte Rodwell, Charlotte Gueydan, Val´ erie Lanneau, Daniel Murphy, Yann Le Cam, and Ana Rath

    St´ ephanie Nguengang Wakap, Deborah M. Lambert, Annie Olry, Char- lotte Rodwell, Charlotte Gueydan, Val´ erie Lanneau, Daniel Murphy, Yann Le Cam, and Ana Rath. Estimating cumulative point prevalence of rare diseases: analysis of the orphanet database. European Journal of Hum...

  30. [38]

    Deep learning for genomics: A concise overview

    Tuan Nguyen, Thanh Tran, and Thao Nguyen. Deep learning for genomics: A concise overview. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 12(2):e1449, 2022

  31. [39]

    A scoping review on generative ai and large language models in mitigating medication related harm

    Jasmine Chiat Ling Ong, Michael Hao Chen, Ning Ng, Kabilan Elangovan, Nichole Yue Ting Tan, Liyuan Jin, Qihuang Xie, Daniel Shu Wei Ting, Rosa Rodriguez-Monguio, David W Bates, et al. A scoping review on generative ai and large language models in mitigating medication related ...

  32. [40]

    Cystic fibrosis: a review

    Thida Ong and Bonnie W Ramsey. Cystic fibrosis: a review. Jama, 329(21):1859–1871, 2023

  33. [41]

    Large language models vote: Prompting for rare dis- ease identification

    David Oniani, Jordan Hilsman, Hang Dong, Fengyi Gao, Shiven Verma, and Yanshan Wang. Large language models vote: Prompting for rare dis- ease identification. arXiv preprint arXiv:2308.12890 , 2023

  34. [42]

    Artificial intelligence empowering rare diseases: a bibliometric perspective over the last two decades

    Peiling Ou, Ru Wen, Linfeng Shi, Jian Wang, and Chen Liu. Artificial intelligence empowering rare diseases: a bibliometric perspective over the last two decades. Orphanet Journal of Rare Diseases , 19(1):345, 2024

  35. [43]

    Studying reddit: A systematic overview of disciplines, approaches, methods, and ethics

    Nicholas Proferes, Naiyan Jones, Sarah Gilbert, Casey Fiesler, and Michael Zimmer. Studying reddit: A systematic overview of disciplines, approaches, methods, and ethics. Social Media+ Society , 7(2):20563051211019004, 2021. 21

  36. [44]

    Rare disease diagnosis: a review of web search, social media and large-scale data-mining approaches

    Srivatsan Rao, Gino Fung, and Sangeeta Rao. Rare disease diagnosis: a review of web search, social media and large-scale data-mining approaches. Rare Diseases, 5(1):e1355661, 2017

  37. [45]

    Rare disease policies to im- prove care for patients in europe

    Charlotte Rodwell and S´ egol` ene Aym´ e. Rare disease policies to im- prove care for patients in europe. Biochimica et Biophysica Acta (BBA)- Molecular Basis of Disease , 1852(10):2329–2335, 2015

  38. [46]

    Why rare diseases are an important medical and social issue

    Arrigo Schieppati, Jan-Inge Henter, Erica Daina, and Anita Aperia. Why rare diseases are an important medical and social issue. The Lancet , 371(9629):2039–2041, 2008

  39. [47]

    A systematic review of large language model (llm) evaluations in clinical medicine

    Sina Shool, Sara Adimi, Reza Saboori Amleshi, Ehsan Bitaraf, Reza Golpira, and Mahmood Tara. A systematic review of large language model (llm) evaluations in clinical medicine. BMC Medical Informatics and De- cision Making, 25(1):117, 2025

  40. [48]

    Identifying and extracting rare diseases and their phenotypes with large language models

    Cathy Shyr, Yan Hu, Lisa Bastarache, Alex Cheng, Rizwan Hamid, Paul Harris, and Hua Xu. Identifying and extracting rare diseases and their phenotypes with large language models. Journal of Healthcare Informatics Research, 8(2):438–461, 2024

  41. [49]

    Large language models encode clinical knowledge

    Karan Singhal, Shekoofeh Azizi, Timo Tu, Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Samuel Coley, Joon Lee, et al. Large language models encode clinical knowledge. Nature, 620(7972):172–180, 2023

  42. [50]

    Zebra-llama: A context-aware large language model for democratizing rare disease knowledge

    Karthik Soman, Andrew Langdon, Catalina Villouta, Chinmay Agrawal, Lashaw Salta, Braian Peetoom, Gianmarco Bellucci, and Orion J Buske. Zebra-llama: A context-aware large language model for democratizing rare disease knowledge. arXiv preprint arXiv:2411.02657 , 2024

  43. [51]

    Huntington’s disease: Diagnosis and management

    Thomas B Stoker, Sarah L Mason, Julia C Greenland, Simon T Holden, Helen Santini, and Roger A Barker. Huntington’s disease: Diagnosis and management. Practical neurology, 22(1):32–41, 2022

  44. [52]

    Release and impact of china’s” second list of rare diseases”

    Mi Tang, Yan Yang, Ziping Ye, Peipei Song, Chunlin Jin, Qi Kang, and Jiangjiang He. Release and impact of china’s” second list of rare diseases”. Intractable & Rare Diseases Research, 12(4):251–256, 2023

  45. [53]

    Ramedis: the rare metabolic diseases database

    Thoralf T¨ opel, Ralf Hofest¨ adt, Dagmar Scheible, and Friedrich Trefz. Ramedis: the rare metabolic diseases database. Applied bioinformatics , 5:115–118, 2006

  46. [54]

    Food and Drug Administration

    U.S. Food and Drug Administration. Rare diseases: About rare diseases. https://www.fda.gov/patients/rare-diseases-fda, 2023. Accessed: 2025-05-01

  47. [55]

    Food and Drug Administration

    U.S. Food and Drug Administration. Rare diseases at fda, 2024. Accessed: 2025-05-04. 22

  48. [56]

    Ordo: an ontology connecting rare disease, epidemiology and genetic data

    Drashtti Vasant, Laetitia Chanas, James Malone, Marc Hanauer, Annie Olry, Simon Jupp, Peter N Robinson, Helen Parkinson, and Ana Rath. Ordo: an ontology connecting rare disease, epidemiology and genetic data. In Proceedings of ISMB, volume 30, pages 1–4. researchgate. net, 2014

  49. [57]

    Mondo Dis- ease Ontology: harmonizing disease concepts across the world, volume 2807

    Nicole Vasilevsky, Shahim Essaid, Nico Matentzoglu, Nomi L Harris, Melissa Haendel, Peter Robinson, and Christopher J Mungall. Mondo Dis- ease Ontology: harmonizing disease concepts across the world, volume 2807. eScholarship, University of California, 2020

  50. [58]

    The impact of artificial intelligence in the odyssey of rare diseases

    Anna Visibelli, Bianca Roncaglia, Ottavia Spiga, and Annalisa San- tucci. The impact of artificial intelligence in the odyssey of rare diseases. Biomedicines, 11(3):887, 2023

  51. [59]

    Pubmed 2.0

    Jacob White. Pubmed 2.0. Medical reference services quarterly, 39(4):382– 387, 2020

  52. [60]

    A hybrid framework with large language models for rare disease phenotyping

    Jinge Wu, Hang Dong, Zexi Li, Haowei Wang, Runci Li, Arijit Patra, Chengliang Dai, Waqar Ali, Phil Scordis, and Honghan Wu. A hybrid framework with large language models for rare disease phenotyping. BMC Medical Informatics and Decision Making , 24(1):289, 2024

  53. [61]

    Understanding sarcoidosis using large language models and social media data

    Nan Miles Xi, Hong-Long Ji, and Lin Wang. Understanding sarcoidosis using large language models and social media data. Journal of Healthcare Informatics Research, pages 1–26, 2024

  54. [62]

    Rdguru: a conversa- tional intelligent agent for rare diseases

    Jian Yang, Liqi Shu, Huilong Duan, and Haomin Li. Rdguru: a conversa- tional intelligent agent for rare diseases. IEEE Journal of Biomedical and Health Informatics, 2024

  55. [63]

    Machine learning for drug discovery

    Kevin Yang, Kristina Swanson, Wengong Jin, Connor Coley, Paul Eiden, Hua Gao, Alberto Guzman-Perez, Tom Hopper, Bryan Kelley, Matthias Mathea, et al. Machine learning for drug discovery. Nature Reviews Drug Discovery, 18(6):463–477, 2019

  56. [64]

    Diagnostic accuracy of a custom large language model on rare pediatric disease case reports

    Cameron C Young, Ellie Enichen, Christian Rivera, Corinne A Auger, Nathan Grant, Arya Rao, and Marc D Succi. Diagnostic accuracy of a custom large language model on rare pediatric disease case reports. Amer- ican Journal of Medical Genetics Part A , 197(2):e63878, 2025

  57. [65]

    Innovations in medicine: Exploring chatgpt’s im- pact on rare disorder management

    Stefania Zampatti, Cristina Peconi, Domenica Megalizzi, Giulia Calvino, Giulia Trastulli, Raffaella Cascella, Claudia Strafella, Carlo Caltagirone, and Emiliano Giardina. Innovations in medicine: Exploring chatgpt’s im- pact on rare disorder management. Genes, 15(4):421, 2024

  58. [66]

    A survey of large language models for healthcare

    Jie Zhao, Di Jin, Yifan Yang, and Fei Liu. A survey of large language models for healthcare. Journal of Biomedical Informatics , 144:104408, 2023. 23 Table 5: Limitations of Large Language Models (LLMs) in Supporting Rare Disease Diagnosis Limitation Description Hallucination ...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.