REVIEW 9 cited by
Large Language Models are Few-Shot Clinical Information Extractors
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
A long-running goal of the clinical NLP community is the extraction of important variables trapped in clinical notes. However, roadblocks have included dataset shift from the general domain and a lack of public clinical corpora and annotations. In this work, we show that large language models, such as InstructGPT, perform well at zero- and few-shot information extraction from clinical text despite not being trained specifically for the clinical domain. Whereas text classification and generation performance have already been studied extensively in such models, here we additionally demonstrate how to leverage them to tackle a diverse set of NLP tasks which require more structured outputs, including span identification, token-level sequence classification, and relation extraction. Further, due to the dearth of available data to evaluate these systems, we introduce new datasets for benchmarking few-shot clinical information extraction based on a manual re-annotation of the CASI dataset for new tasks. On the clinical extraction tasks we studied, the GPT-3 systems significantly outperform existing zero- and few-shot baselines.
Forward citations
Cited by 9 Pith papers
-
Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning
A dual clinical-computational taxonomy for medical LLM reasoning plus a five-level 5k-sample benchmark showing specialists excel at diagnosis and general models at decision support/dialogue.
-
From scratch to silver: Creating trustworthy training data for patent-SDG classification using Large Language Models
A weak-supervision pipeline using LLM-extracted concepts and rank fusion creates silver-standard patent-to-SDG labels that recover known citation-derived associations and show high network modularity.
-
CaseReportBench: An LLM Benchmark Dataset for Dense Information Extraction in Clinical Case Reports
CaseReportBench tests LLMs on dense information extraction from 138 rare-disease case reports and reports that Qwen2.5-7B outperforms GPT-4o under string-based metrics.
-
Structured Extraction of Real World Medical Knowledge using LLMs for Summarization and Search
An LLM plus knowledge graph pipeline extracts HPO phenotypes from EHR notes and identifies 12 suspected undiagnosed BPAN patients among 33.6 million, with benchmark validation on three public datasets.
-
Information Extraction from Clinical Notes: Are We Ready to Switch to Large Language Models?
Instruction-tuned LLaMA models beat BERT on clinical NER and relation extraction, gaining up to 7% F1 on an unseen institution, but with much higher compute cost and lower throughput.
-
Latent Factor Point Processes for Patient Representation in Electronic Health Records
A latent factor point process model plus Fourier spectral embeddings gives new patient-level representations for EHR classification and clustering, but the stated theoretical guarantees contain a diverging error term.
-
CardioCoT: Hierarchical Reasoning for Multimodal Survival Analysis
CardioCoT uses GPT-4o-generated hierarchical reasoning text as an extra input to a multimodal survival model, reporting higher C-index for MACE recurrence but with potential outcome-label leakage and no external validation.
-
Understanding the Dark Side of LLMs' Intrinsic Self-Correction
Intrinsic self-correction makes state-of-the-art LLMs overturn correct answers across four task types, and simple question repetition or tiny fine-tuning reduces this damage.
-
LightLLM: A Versatile Large Language Model for Predictive Light Sensing
A frozen-LLM framework with task-specific encoders, knowledge prompts, and LoRA tuning reports 4.4x and 3.4x improvements over prior models for unseen-environment light-based localization and indoor solar estimation.
Discussion (0). Continue with ORCID to comment.