Pith. sign in

REVIEW 1 cited by

GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.15760 v1 pith:RSZXT6ME submitted 2024-05-24 cs.CL cs.CY

GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction

classification cs.CL cs.CY
keywords benchmarkannotationbiasbiasescommunitygpt-3humanturbo
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Social biases in LLMs are usually measured via bias benchmark datasets. Current benchmarks have limitations in scope, grounding, quality, and human effort required. Previous work has shown success with a community-sourced, rather than crowd-sourced, approach to benchmark development. However, this work still required considerable effort from annotators with relevant lived experience. This paper explores whether an LLM (specifically, GPT-3.5-Turbo) can assist with the task of developing a bias benchmark dataset from responses to an open-ended community survey. We also extend the previous work to a new community and set of biases: the Jewish community and antisemitism. Our analysis shows that GPT-3.5-Turbo has poor performance on this annotation task and produces unacceptable quality issues in its output. Thus, we conclude that GPT-3.5-Turbo is not an appropriate substitute for human annotation in sensitive tasks related to social biases, and that its use actually negates many of the benefits of community-sourcing bias benchmarks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AnnotateThis: Analyzing a human-LLM system for annotating social media data with the concept of climate change mitigation pessimism

    cs.CY 2026-06 unverdicted novelty 5.0

    AnnotateThis lets users improve LLM annotations for climate change mitigation pessimism on social media, yielding 0.15 higher F-Measure and 0.23 higher accuracy than automated prompt refinement when ground truth label...