REVIEW 12 cited by
Large Language Models for Disease Diagnosis: A Scoping Review
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Automatic disease diagnosis has become increasingly valuable in clinical practice. The advent of large language models (LLMs) has catalyzed a paradigm shift in artificial intelligence, with growing evidence supporting the efficacy of LLMs in diagnostic tasks. Despite the increasing attention in this field, a holistic view is still lacking. Many critical aspects remain unclear, such as the diseases and clinical data to which LLMs have been applied, the LLM techniques employed, and the evaluation methods used. In this article, we perform a comprehensive review of LLM-based methods for disease diagnosis. Our review examines the existing literature across various dimensions, including disease types and associated clinical specialties, clinical data, LLM techniques, and evaluation methods. Additionally, we offer recommendations for applying and evaluating LLMs for diagnostic tasks. Furthermore, we assess the limitations of current research and discuss future directions. To our knowledge, this is the first comprehensive review for LLM-based disease diagnosis.
Forward citations
Cited by 12 Pith papers
-
FinLoRA: Benchmarking LoRA Methods for Fine-Tuning LLMs on Financial Datasets
A benchmark showing LoRA fine-tuning improves large language models on financial tasks, with four new XBRL-format analysis datasets from 150 SEC filings.
-
Early Diagnosis of Atrial Fibrillation Recurrence: A Large Tabular Model Approach with Structured and Unstructured Clinical Data
TabPFN outperforms SVM and the CHADS2-VASc, HATCH, and APPLE scores at predicting AF recurrence within two years of onset, using NLP-enriched EHR features, but all models remain weak in absolute terms.
-
Fragments to Facts: Partial-Information Fragment Inference from LLMs
Fine-tuned LLMs leak private fragment-level information to adversaries holding only a few unordered public fragments, as shown by two probe attacks (LR-Attack and PRISM) on medical and legal summarization tasks.
-
RAMIE: Retrieval-Augmented Multi-task Information Extraction with Large Language Models on Dietary Supplements
RAMIE, a retrieval-augmented multi-task instruction-tuned framework, improves LLM information extraction for dietary supplements from clinical records, with RAG recovering accuracy lost in multi-task training.
-
Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes
The paper reports that a fact-checker trained on replaced entities, used both as a reward and as the evaluation metric, raises measured step-factuality of small open LLMs by up to 49.9 percentage points.
-
AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity
AnchorAttention uses the maximum attention score from initial and local tokens as an anchor to threshold-select important key-value positions at stripe granularity, achieving faster prefill with comparable accuracy.
-
A Multi-granularity Concept Sparse Activation and Hierarchical Knowledge Graph Fusion Framework for Rare Disease Diagnosis
A retrieval and knowledge-graph framework lifts rare-disease QA accuracy by 0.12 on average (0.22 for the weaker LLM, 0.02 for the stronger) on 100 BioASQ questions, reaching 0.89.
-
Retrieval-augmented in-context learning for multimodal large language models in disease classification
RAICL retrieves similar image-text examples as demonstrations and improves multimodal LLM disease classification accuracy, but the reported gains are measured against zero-shot prompting, not against random in-context...
-
A Case Study Exploring the Current Landscape of Synthetic Medical Record Generation with Commercial LLMs
Commercial LLMs generate usable synthetic ICU records only for small feature sets, with fidelity and downstream prediction quality degrading sharply as dimensionality grows.
-
Continually Evolved Multimodal Foundation Models for Cancer Prognosis
A cancer prognosis model that grows with new data modalities via LoRA adapters and gated query fusion is claimed to beat fusion baselines, but the supporting table is incomplete and internally inconsistent.
-
Large Language models for Time Series Analysis: Techniques, Applications, and Challenges
A review of LLM-based time series analysis that proposes several taxonomies, but is undermined by citation errors and a lack of systematic methodology.
-
Structured Semantics from Unstructured Notes: Language Model Approaches to EHR-Based Decision Support
A position paper surveying language model approaches to EHR decision support, with illustrative, not real, experimental results.
Discussion (0). Continue with ORCID to comment.