Pith. sign in

REVIEW 2 cited by

How far is Language Model from 100% Few-shot Named Entity Recognition in Medical Domain

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.00186 v2 pith:5GV5XCEE submitted 2023-07-01 cs.CL

classification cs.CL
keywords medicalmodelsfew-shotdomainentityperformancetaskseffective
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements in language models (LMs) have led to the emergence of powerful models such as Small LMs (e.g., T5) and Large LMs (e.g., GPT-4). These models have demonstrated exceptional capabilities across a wide range of tasks, such as name entity recognition (NER) in the general domain. (We define SLMs as pre-trained models with fewer parameters compared to models like GPT-3/3.5/4, such as T5, BERT, and others.) Nevertheless, their efficacy in the medical section remains uncertain and the performance of medical NER always needs high accuracy because of the particularity of the field. This paper aims to provide a thorough investigation to compare the performance of LMs in medical few-shot NER and answer How far is LMs from 100\% Few-shot NER in Medical Domain, and moreover to explore an effective entity recognizer to help improve the NER performance. Based on our extensive experiments conducted on 16 NER models spanning from 2018 to 2023, our findings clearly indicate that LLMs outperform SLMs in few-shot medical NER tasks, given the presence of suitable examples and appropriate logical frameworks. Despite the overall superiority of LLMs in few-shot medical NER tasks, it is important to note that they still encounter some challenges, such as misidentification, wrong template prediction, etc. Building on previous findings, we introduce a simple and effective method called \textsc{RT} (Retrieving and Thinking), which serves as retrievers, finding relevant examples, and as thinkers, employing a step-by-step reasoning process. Experimental results show that our proposed \textsc{RT} framework significantly outperforms the strong open baselines on the two open medical benchmark datasets

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Extracting OPQRST in Electronic Health Records using Large Language Models with Reasoning

    cs.CL 2025-09 conditional novelty 4.0 of 10

    Reasoning-style prompts improve few-shot LLM extraction of OPQRST items from EHR notes, but the result rests on an 85-note single-annotator evaluation with an LLM judge.

  2. Token and Span Classification for Entity Recognition in French Historical Encyclopedias

    cs.CL 2025-06 conditional novelty 4.0 of 10

    On the GeoEDdA corpus of 18th century French encyclopedias, fine-tuned CamemBERT achieves the highest macro-averaged F1, closely followed by Flair, and few-shot GPT models lag behind but may help when labeled data are scarce.

Pith tools