Pith. sign in

REVIEW 2 cited by

Human Still Wins over LLM: An Empirical Study of Active Learning on Domain-Specific Annotation Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.09825 v1 pith:NMP6NLPJ submitted 2023-11-16 cs.CL

classification cs.CL
keywords modelstasksannotationdomain-specificexperthumanlearningllms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) have demonstrated considerable advances, and several claims have been made about their exceeding human performance. However, in real-world tasks, domain knowledge is often required. Low-resource learning methods like Active Learning (AL) have been proposed to tackle the cost of domain expert annotation, raising this question: Can LLMs surpass compact models trained with expert annotations in domain-specific tasks? In this work, we conduct an empirical experiment on four datasets from three different domains comparing SOTA LLMs with small models trained on expert annotations with AL. We found that small models can outperform GPT-3.5 with a few hundreds of labeled data, and they achieve higher or similar performance with GPT-4 despite that they are hundreds time smaller. Based on these findings, we posit that LLM predictions can be used as a warmup method in real-world applications and human experts remain indispensable in tasks involving data annotation driven by domain-specific knowledge.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. A Multi-Stage Large Language Model Framework for Extracting Suicide-Related Social Determinants of Health

    cs.CL 2025-08 conditional novelty 5.0 of 10

    A multi-stage LLM pipeline improves extraction of infrequent suicide-related social determinants from death narratives, but some evaluation results are compromised by using the test set to tune the system.

  2. Conflicting narratives and polarization on social media

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Conflicting narrative roles and selected actors in German Twitter debates reveal discursive fault lines that may explain issue alignment between left and right camps.

Pith tools