Pith. sign in

REVIEW 9 cited by

Large Language Models are Few-Shot Clinical Information Extractors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.12689 v2 pith:6KKY5CM3 submitted 2022-05-25 cs.CL cs.AI

classification cs.CLcs.AI
keywords clinicalextractionfew-shotinformationmodelstasksclassificationdataset
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

A long-running goal of the clinical NLP community is the extraction of important variables trapped in clinical notes. However, roadblocks have included dataset shift from the general domain and a lack of public clinical corpora and annotations. In this work, we show that large language models, such as InstructGPT, perform well at zero- and few-shot information extraction from clinical text despite not being trained specifically for the clinical domain. Whereas text classification and generation performance have already been studied extensively in such models, here we additionally demonstrate how to leverage them to tackle a diverse set of NLP tasks which require more structured outputs, including span identification, token-level sequence classification, and relation extraction. Further, due to the dearth of available data to evaluate these systems, we introduce new datasets for benchmarking few-shot clinical information extraction based on a manual re-annotation of the CASI dataset for new tasks. On the clinical extraction tasks we studied, the GPT-3 systems significantly outperform existing zero- and few-shot baselines.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 48 citations worldwide. Full citation record

  1. Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning

    cs.AI 2026-07 accept novelty 6.0 of 10

    A dual clinical-computational taxonomy for medical LLM reasoning plus a five-level 5k-sample benchmark showing specialists excel at diagnosis and general models at decision support/dialogue.

  2. From scratch to silver: Creating trustworthy training data for patent-SDG classification using Large Language Models

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A weak-supervision pipeline using LLM-extracted concepts and rank fusion creates silver-standard patent-to-SDG labels that recover known citation-derived associations and show high network modularity.

  3. CaseReportBench: An LLM Benchmark Dataset for Dense Information Extraction in Clinical Case Reports

    cs.CL 2025-05 conditional novelty 6.0 of 10

    CaseReportBench tests LLMs on dense information extraction from 138 rare-disease case reports and reports that Qwen2.5-7B outperforms GPT-4o under string-based metrics.

  4. Structured Extraction of Real World Medical Knowledge using LLMs for Summarization and Search

    cs.CL 2024-12 conditional novelty 6.0 of 10

    An LLM plus knowledge graph pipeline extracts HPO phenotypes from EHR notes and identifies 12 suspected undiagnosed BPAN patients among 33.6 million, with benchmark validation on three public datasets.

  5. Information Extraction from Clinical Notes: Are We Ready to Switch to Large Language Models?

    cs.CL 2024-11 conditional novelty 6.0 of 10

    Instruction-tuned LLaMA models beat BERT on clinical NER and relation extraction, gaining up to 7% F1 on an unseen institution, but with much higher compute cost and lower throughput.

  6. Latent Factor Point Processes for Patient Representation in Electronic Health Records

    stat.ME 2025-08 reject novelty 5.0 of 10

    A latent factor point process model plus Fourier spectral embeddings gives new patient-level representations for EHR classification and clustering, but the stated theoretical guarantees contain a diverging error term.

  7. CardioCoT: Hierarchical Reasoning for Multimodal Survival Analysis

    cs.AI 2025-05 reject novelty 5.0 of 10

    CardioCoT uses GPT-4o-generated hierarchical reasoning text as an extra input to a multimodal survival model, reporting higher C-index for MACE recurrence but with potential outcome-label leakage and no external validation.

  8. Understanding the Dark Side of LLMs' Intrinsic Self-Correction

    cs.CL 2024-12 conditional novelty 5.0 of 10

    Intrinsic self-correction makes state-of-the-art LLMs overturn correct answers across four task types, and simple question repetition or tiny fine-tuning reduces this damage.

  9. LightLLM: A Versatile Large Language Model for Predictive Light Sensing

    cs.LG 2024-11 conditional novelty 5.0 of 10

    A frozen-LLM framework with task-specific encoders, knowledge prompts, and LoRA tuning reports 4.4x and 3.4x improvements over prior models for unseen-environment light-based localization and indoor solar estimation.

Pith tools