Pith. sign in

REVIEW 1 cited by

Evaluating Named Entity Recognition Using Few-Shot Prompting with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.15796 v2 pith:DGMYQDCZ submitted 2024-08-28 cs.IR cs.AI

classification cs.IRcs.AI
keywords few-shotmodelslargeentityperformancepromptingdatasetslabeled
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper evaluates Few-Shot Prompting with Large Language Models for Named Entity Recognition (NER). Traditional NER systems rely on extensive labeled datasets, which are costly and time-consuming to obtain. Few-Shot Prompting or in-context learning enables models to recognize entities with minimal examples. We assess state-of-the-art models like GPT-4 in NER tasks, comparing their few-shot performance to fully supervised benchmarks. Results show that while there is a performance gap, large models excel in adapting to new entity types and domains with very limited data. We also explore the effects of prompt engineering, guided output format and context length on performance. This study underscores Few-Shot Learning's potential to reduce the need for large labeled datasets, enhancing NER scalability and accessibility.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Token and Span Classification for Entity Recognition in French Historical Encyclopedias

    cs.CL 2025-06 conditional novelty 4.0 of 10

    On the GeoEDdA corpus of 18th century French encyclopedias, fine-tuned CamemBERT achieves the highest macro-averaged F1, closely followed by Flair, and few-shot GPT models lag behind but may help when labeled data are scarce.

Pith tools