Pith. sign in

REVIEW 12 cited by

Large Language Models for Disease Diagnosis: A Scoping Review

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.00097 v3 pith:CEJYNLXF submitted 2024-08-27 cs.CL cs.AI

classification cs.CLcs.AI
keywords diseaseclinicaldiagnosisllmsreviewmethodscomprehensivedata
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Automatic disease diagnosis has become increasingly valuable in clinical practice. The advent of large language models (LLMs) has catalyzed a paradigm shift in artificial intelligence, with growing evidence supporting the efficacy of LLMs in diagnostic tasks. Despite the increasing attention in this field, a holistic view is still lacking. Many critical aspects remain unclear, such as the diseases and clinical data to which LLMs have been applied, the LLM techniques employed, and the evaluation methods used. In this article, we perform a comprehensive review of LLM-based methods for disease diagnosis. Our review examines the existing literature across various dimensions, including disease types and associated clinical specialties, clinical data, LLM techniques, and evaluation methods. Additionally, we offer recommendations for applying and evaluating LLMs for diagnostic tasks. Furthermore, we assess the limitations of current research and discuss future directions. To our knowledge, this is the first comprehensive review for LLM-based disease diagnosis.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FinLoRA: Benchmarking LoRA Methods for Fine-Tuning LLMs on Financial Datasets

    cs.CE 2025-05 conditional novelty 6.0 of 10

    A benchmark showing LoRA fine-tuning improves large language models on financial tasks, with four new XBRL-format analysis datasets from 150 SEC filings.

  2. Early Diagnosis of Atrial Fibrillation Recurrence: A Large Tabular Model Approach with Structured and Unstructured Clinical Data

    cs.LG 2025-05 conditional novelty 6.0 of 10

    TabPFN outperforms SVM and the CHADS2-VASc, HATCH, and APPLE scores at predicting AF recurrence within two years of onset, using NLP-enriched EHR features, but all models remain weak in absolute terms.

  3. Fragments to Facts: Partial-Information Fragment Inference from LLMs

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Fine-tuned LLMs leak private fragment-level information to adversaries holding only a few unordered public fragments, as shown by two probe attacks (LR-Attack and PRISM) on medical and legal summarization tasks.

  4. RAMIE: Retrieval-Augmented Multi-task Information Extraction with Large Language Models on Dietary Supplements

    cs.CL 2024-11 conditional novelty 6.0 of 10

    RAMIE, a retrieval-augmented multi-task instruction-tuned framework, improves LLM information extraction for dietary supplements from clinical records, with RAG recovering accuracy lost in multi-task training.

  5. Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes

    cs.CL 2025-07 reject novelty 5.0 of 10

    The paper reports that a fact-checker trained on replaced entities, used both as a reward and as the evaluation metric, raises measured step-factuality of small open LLMs by up to 49.9 percentage points.

  6. AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity

    cs.LG 2025-05 conditional novelty 5.0 of 10

    AnchorAttention uses the maximum attention score from initial and local tokens as an anchor to threshold-select important key-value positions at stripe granularity, achieving faster prefill with comparable accuracy.

  7. A Multi-granularity Concept Sparse Activation and Hierarchical Knowledge Graph Fusion Framework for Rare Disease Diagnosis

    cs.AI 2025-07 conditional novelty 4.0 of 10

    A retrieval and knowledge-graph framework lifts rare-disease QA accuracy by 0.12 on average (0.22 for the weaker LLM, 0.02 for the stronger) on 100 BioASQ questions, reaching 0.89.

  8. Retrieval-augmented in-context learning for multimodal large language models in disease classification

    cs.AI 2025-05 conditional novelty 4.0 of 10

    RAICL retrieves similar image-text examples as demonstrations and improves multimodal LLM disease classification accuracy, but the reported gains are measured against zero-shot prompting, not against random in-context...

  9. A Case Study Exploring the Current Landscape of Synthetic Medical Record Generation with Commercial LLMs

    cs.CL 2025-04 conditional novelty 4.0 of 10

    Commercial LLMs generate usable synthetic ICU records only for small feature sets, with fidelity and downstream prediction quality degrading sharply as dimensionality grows.

  10. Continually Evolved Multimodal Foundation Models for Cancer Prognosis

    cs.LG 2025-01 reject novelty 4.0 of 10

    A cancer prognosis model that grows with new data modalities via LoRA adapters and gated query fusion is claimed to beat fusion baselines, but the supporting table is incomplete and internally inconsistent.

  11. Large Language models for Time Series Analysis: Techniques, Applications, and Challenges

    cs.LG 2025-05 reject novelty 3.0 of 10

    A review of LLM-based time series analysis that proposes several taxonomies, but is undermined by citation errors and a lack of systematic methodology.

  12. Structured Semantics from Unstructured Notes: Language Model Approaches to EHR-Based Decision Support

    cs.IR 2025-06 unverdicted novelty 1.0 of 10

    A position paper surveying language model approaches to EHR decision support, with illustrative, not real, experimental results.

Pith tools