Pith. sign in

REVIEW 1 cited by

A Survey of NLP-Related Crowdsourcing HITs: what works and what does not

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.05241 v1 pith:EB6SDJGP submitted 2021-11-09 cs.CL

classification cs.CL
keywords hitscrowdsourcingrequestersissuesworkerworkerseffectgiving
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Crowdsourcing requesters on Amazon Mechanical Turk (AMT) have raised questions about the reliability of the workers. The AMT workforce is very diverse and it is not possible to make blanket assumptions about them as a group. Some requesters now reject work en mass when they do not get the results they expect. This has the effect of giving each worker (good or bad) a lower Human Intelligence Task (HIT) approval score, which is unfair to the good workers. It also has the effect of giving the requester a bad reputation on the workers' forums. Some of the issues causing the mass rejections stem from the requesters not taking the time to create a well-formed task with complete instructions and/or not paying a fair wage. To explore this assumption, this paper describes a study that looks at the crowdsourcing HITs on AMT that were available over a given span of time and records information about those HITs. This study also records information from a crowdsourcing forum on the worker perspective on both those HITs and on their corresponding requesters. Results reveal issues in worker payment and presentation issues such as missing instructions or HITs that are not doable.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Reading between the Lines: Can LLMs Identify Cross-Cultural Communication Gaps?

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A user study of 57 preselected Goodreads reviews finds culture-specific comprehension gaps in most texts, while GPT-4o identifies the relevant spans with 0.49 precision and 0.65 recall across India, Mexico, and the USA.

Pith tools