Pith. sign in

Scaling Public Health Text Annotation: Zero-Shot Learning vs. Crowdsourcing for Improved Efficiency and Labeling Accuracy

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Public health researchers are increasingly interested in using social media data to study health-related behaviors, but manually labeling this data can be labor-intensive and costly. This study explores whether zero-shot labeling using large language models (LLMs) can match or surpass conventional crowd-sourced annotation for Twitter posts related to sleep disorders, physical activity, and sedentary behavior. Multiple annotation pipelines were designed to compare labels produced by domain experts, crowd workers, and LLM-driven approaches under varied prompt-engineering strategies. Our findings indicate that LLMs can rival human performance in straightforward classification tasks and significantly reduce labeling time, yet their accuracy diminishes for tasks requiring more nuanced domain knowledge. These results clarify the trade-offs between automated scalability and human expertise, demonstrating conditions under which LLM-based labeling can be efficiently integrated into public health research without undermining label quality.

citation-role summary

background 1

citation-polarity summary

fields

cs.AI 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

support 1

representative citing papers

Enhancing Health Fact-Checking with LLM-Generated Synthetic Data

cs.AI · 2025-08-28 · conditional · novelty 5.0

An LLM-generated synthetic data pipeline, built on sentence-fact entailment tables, improved BERT fact-checking F1 by up to 0.019 on PubHealth and 0.049 on SciFact compared with training on original data alone.

citing papers explorer

Showing 1 of 1 citing paper.

  • Enhancing Health Fact-Checking with LLM-Generated Synthetic Data cs.AI · 2025-08-28 · conditional · none · ref 4 · internal anchor

    An LLM-generated synthetic data pipeline, built on sentence-fact entailment tables, improved BERT fact-checking F1 by up to 0.019 on PubHealth and 0.049 on SciFact compared with training on original data alone.