Pith. sign in

REVIEW 2 cited by

Artificial Artificial Artificial Intelligence: Crowd Workers Widely Use Large Language Models for Text Production Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.07899 v1 pith:N3P5WIIT submitted 2023-06-13 cs.CL cs.CY

classification cs.CLcs.CY
keywords llmscrowddataworkershumanartificialannotationslanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) are remarkable data annotators. They can be used to generate high-fidelity supervised training data, as well as survey and experimental data. With the widespread adoption of LLMs, human gold--standard annotations are key to understanding the capabilities of LLMs and the validity of their results. However, crowdsourcing, an important, inexpensive way to obtain human annotations, may itself be impacted by LLMs, as crowd workers have financial incentives to use LLMs to increase their productivity and income. To investigate this concern, we conducted a case study on the prevalence of LLM usage by crowd workers. We reran an abstract summarization task from the literature on Amazon Mechanical Turk and, through a combination of keystroke detection and synthetic text classification, estimate that 33-46% of crowd workers used LLMs when completing the task. Although generalization to other, less LLM-friendly tasks is unclear, our results call for platforms, researchers, and crowd workers to find new ways to ensure that human data remain human, perhaps using the methodology proposed here as a stepping stone. Code/data: https://github.com/epfl-dlab/GPTurk

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Datasheets Aren't Enough: DataRubrics for Automated Quality Metrics and Accountability

    cs.LG 2025-06 conditional novelty 6.0 of 10

    DataRubrics introduces a structured ten-dimension rubric with an LLM-as-a-judge pipeline to automatically assess dataset quality, positioned as a more measurable alternative to descriptive datasheets for conference review.

  2. Redefining Research Crowdsourcing: Incorporating Human Feedback with LLM-Powered Digital Twins

    cs.HC 2025-05 conditional novelty 5.0 of 10

    A study of an LLM-powered 'digital twin' system for crowd workers shows modest accuracy on Likert-scale surveys, with caveats around threshold tuning and evaluation contamination.

Pith tools