REVIEW 4 major objections 3 minor 132 references
When Incentives Backfire, Data Stops Being Human
T0 review · 4 major / 3 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read The paper argues that the contamination of AI training data by AI-generated content is not primarily a filtering problem but a design problem: data collection systems built around task-based pay and fine-grained control quietly destroy…
desk verdict Clearly argued and honest position piece, but the load-bearing causal claim is an untested extrapolation the authors themselves concede. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is motivational crowding-out, the psychological process by which external rewards shift a contributor's self-perception from 'I do this because I enjoy it or it matters' to 'I do this for the pay,' reducing effort, creativity, and long-term engagement. The paper operationalizes this through the overjustification effect (Lepper et al., 1973) and self-determination theory (Deci et al., 2017), and pairs it with a design variable called the resolution of control: how finely a data collection system constrains contributor behavior. High-resolution, task-level control maximizes throughput but triggers crowding-out; low-resolution, environment-level control preserves motivation but risks incoherent data; the paper argues that sustainable designs sit in the middle, using structure and rules to elicit behavior without transactional surveillance.
What would settle it
A controlled field experiment on a mainstream annotation platform that pays one group per task and another a flat participation fee (or none) for the same annotation workload, measuring output quality, diversity, and LLM-use rate over several weeks; if the per-task group matches or beats the intrinsic-motivation group on quality and engagement, the paper's central premise fails.
Extended reading notes
Core claim
The central claim is that over-reliance on external incentives and task fragmentation, the two standard engineering levers of crowdsourcing, erode the intrinsic motivation that sustains high-quality human data, and that this erosion, not the mere presence of LLMs, is what makes modern data collection fragile. The paper assembles evidence from psychology (overjustification, self-perception theory, self-determination theory) and economics (Goodhart's law, perverse incentives) and maps it onto the structure of MTurk-style platforms, where it predicts a vicious cycle: as intrinsic motivation fades, platforms tighten control and raise pay, which accelerates shortcutting behavior such as using LLMs to complete tasks, further degrading data quality. The authors propose a 'resolution of control' spectrum, with tightly controlled task-level systems at one end and laissez-faire community platforms at the other, and locate the desirable design space in the middle, exemplified by product-integrated systems and by games. The paper's positive thesis is that structured environments based on voluntary participation, such as games with a purpose, citizen science, and community platforms, can deliver high-quality, high-quantity data while preserving contributor trust.
Load-bearing premise
The argument stands on the transfer of motivational crowding-out from laboratory and workplace settings to large-scale machine-learning data collection, specifically that lowering task-level financial incentives will improve rather than reduce data quality and quantity.
Editorial extensions
If this is right
- AI data pipelines should be redesigned around autonomy, competence, and relatedness rather than piece-rate payment.
- Detecting and filtering AI-generated content is a stopgap; the durable fix is making human contribution itself harder to replace.
- Fragmented micro-tasking into ever smaller units should be treated as a known risk to data quality, not a neutral scaling tool.
- Games with a purpose and citizen-science platforms become a first-class source of training data rather than a curiosity.
- Trust, the sense that contributors' data is used legitimately, becomes a third axis of data system design alongside quality and quantity.
Reading between the lines
- Beyond the paper, the argument implies that the marginal value of higher annotation pay may be negative for exactly the tasks where authenticity matters most, such as preference data and human-behavior corpora.
- Beyond the paper, the framing suggests a testable prediction: platforms that add transparent recognition, feedback, and autonomy will see lower LLM-copying rates at equal pay.
- Beyond the paper, the same crowding-out logic applies to data contributors inside companies and to voluntary community data, which would extend the paper's scope beyond crowdsourcing marketplaces.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that the rise of LLM-generated content in human data collection—both on crowdwork platforms and on the open Internet—is not primarily a filtering problem but a symptom of flawed system design. Drawing on psychological theories (overjustification, self-perception theory, self-determination theory) and economic concepts (Goodhart's Law, perverse incentives), the authors claim that excessive reliance on external incentives and task fragmentation erodes intrinsic motivation, which in turn degrades data quality and pushes contributors toward LLM shortcutting. They propose shifting from task-level control to a 'resolution-of-control spectrum' with a desirable middle ground, and they advocate games as a promising form of structured, intrinsically motivating data collection. They also discuss trust, compensation, and ethical considerations. The paper contains no new empirical data; it is an argumentative synthesis that frames a research agenda for rethinking data collection systems.
Significance. If the thesis holds, it would reframe an active area of ML research: instead of developing better filters for AI-generated content, the community would invest in designing data collection systems that preserve intrinsic motivation. This is a timely and important direction, and the paper brings together a broad, relevant literature from social psychology, economics, and human-computer interaction. Its honest acknowledgment that no ML data-collection game has yet demonstrated sustained success (Section 8.2) is a strength, as is its clear elucidation of the overjustification mechanism. However, the central causal chain—from incentive design to motivational crowding-out, and from crowding-out to LLM contamination—is asserted rather than tested, and a key empirical citation is used inconsistently. The paper is best read as a hypothesis-rich agenda-setting piece, but it currently overstates the confidence in its own remedy.
major comments (4)
- [Section 5 and Section 6] The paper cites Mason & Watts (2009) in Section 5 as evidence that financial compensation leads to greater effort and better quantity and/or quality, but in Section 6 it cites the same paper (along with Ikeda & Bernstein, 2016) to support the claim that piece-rate or pay-per-task systems degrade output quality. Mason & Watts actually reported that higher pay increased quantity without reducing quality, and Ikeda & Bernstein found that per-task payments reduced productivity. These findings do not support the crowding-out transfer in the direction the paper needs, and using the same citation for opposing claims weakens the empirical foundation of the argument.
- [Section 5, crowdwork versus community platforms] The central causal claim that MTurk-style platforms suffer a 'crowding-out effect' over time, leading to LLM-assisted or automated annotation, is not supported by the cited evidence. The overjustification experiments (Lepper et al., 1973) and self-determination theory (Deci et al., 2017) come from controlled lab or workplace settings; the paper does not provide evidence that crowdwork contributors experience motivational crowding-out, nor that the rise in LLM use documented by Veselovsky et al. (2023, 2025) is caused by that crowding-out rather than by macroeconomic pressures, task design, or the simple availability of cheap AI tools. This missing link is load-bearing for the paper's thesis.
- [Section 8.2] The paper's own admission that 'no data collection games in machine learning have yet demonstrated sustained success' directly undermines the proposed remedy. The successful examples appealed to next—Zooniverse, Foldit, Lab in the Wild—are citizen-science platforms with self-selected volunteers who have strong domain interest; the paper does not show that such models can scale to the volumes or task types that ML annotation pipelines require, especially for tedious, low-intrinsic-appeal tasks. Without at least one concrete case study or small-scale controlled comparison, the claim that games can expand the quality–quantity frontier for ML data remains an untested hypothesis.
- [Section 7.1 and Figure 4] The 'resolution-of-control spectrum' is introduced as a central design variable, but it is not operationally defined: there are no criteria for placing a system on the spectrum, no measurement instrument, and no way to determine when a design has medium resolution. As a result, the recommendation to design for the 'middle' is not falsifiable in its present form. The paper should either specify observable features that determine spectrum position or explicitly frame the spectrum as a heuristic that generates testable comparative predictions.
minor comments (3)
- [Throughout] There are frequent typographical spacing errors in names such as 'V on Ahn', 'GW AP', 'F oldIt', 'V odrahalli', and 'Kahneman & Tversky'; these should be corrected.
- [Section 6] The discussion of System 1 and System 2 is attributed to Kahneman & Tversky (2013), but the cited reference is 'Prospect Theory'; the appropriate citation is Kahneman's 'Thinking, Fast and Slow' or his joint work on dual-process models.
- [Figure 1] The perpetual donkey machine analogy is evocative but does not clearly map onto the paper's argument; consider replacing the figure with a schematic of the overjustification mechanism or adding a caption that explicitly connects the two.
Circularity Check
No circularity: the paper makes no fitted predictions and its argument is grounded in external psychology and economics literature.
full rationale
This is an argumentative position paper, not a derivation: there are no fitted parameters, equations, or quantitative predictions whose outputs could reduce to their inputs. The central claim—that overreliance on external incentives can crowd out intrinsic motivation and degrade human data quality—is supported by external experimental literature (Lepper et al. 1973; Deci et al. 2017; Bem 1972) and by independent empirical studies of LLM use in crowdwork (Veselovsky et al. 2023; 2025), even though one of the latter's authors overlaps with the present paper. That self-citation is not load-bearing: it documents the prevalence of LLM use rather than supplying the motivational argument, and the paper does not invoke any author-specific uniqueness theorem or ansatz. The quality-quantity frontier and resolution-of-control spectrum are organizing frameworks, not results derived from the paper's own definitions; Section 8.2 even concedes that no ML data-collection game has yet demonstrated sustained success, which is a limitation rather than a circular step. Whether the psychological crowding-out effect transfers to large-scale ML annotation is an untested empirical extrapolation, but that is a correctness risk, not circularity. Accordingly no circular step is present.
Assumptions & free parameters
assumptions (5)
- domain assumption The overjustification effect, demonstrated in children drawing, transfers to adult crowdworkers and long-term platform participation.
- domain assumption Intrinsic motivation, rather than compensation, is the primary driver of sustained high-quality data contributions.
- domain assumption The quantity-quality tradeoff is a Pareto frontier that system design can expand.
- ad hoc to paper Games can simultaneously provide structured data and preserve intrinsic motivation at scale.
- domain assumption High-quality data is best proxied by naturalness, defined as unprompted, incentive-free human behavior.
invented entities (1)
-
Resolution-of-control spectrum
Cite this review
Pith. "Pith review of When Incentives Backfire, Data Stops Being Human." pith.science (2026). https://pith.science/paper/H3P32UXC
@misc{pith2026250207732,
author = {Pith},
title = {Pith review of: When Incentives Backfire, Data Stops Being Human},
year = {2026},
howpublished = {\url{https://pith.science/paper/H3P32UXC}},
note = {Machine review of arXiv:2502.07732}
}
read the original abstract
Progress in AI has relied on human-generated data, from annotator marketplaces to the wider Internet. However, the widespread use of large language models now threatens the quality and integrity of human-generated data on these very platforms. We argue that this issue goes beyond the immediate challenge of filtering AI-generated content -- it reveals deeper flaws in how data collection systems are designed. Existing systems often prioritize speed, scale, and efficiency at the cost of intrinsic human motivation, leading to declining engagement and data quality. We propose that rethinking data collection systems to align with contributors' intrinsic motivations -- rather than relying solely on external incentives -- can help sustain high-quality data sourcing at scale while maintaining contributor trust and long-term participation.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...
-
[2]
M., Longpre, S., Lambert, N., Wang, X., Muennighoff, N., Hou, B., Pan, L., Jeong, H., et al
Albalak, A., Elazar, Y., Xie, S. M., Longpre, S., Lambert, N., Wang, X., Muennighoff, N., Hou, B., Pan, L., Jeong, H., et al. A survey on data selection for language models. arXiv preprint arXiv:2402.16827, 2024
arXiv 2024
-
[3]
E., Gershman, S
Allen, K., Br \"a ndle, F., Botvinick, M., Fan, J. E., Gershman, S. J., Gopnik, A., Griffiths, T. L., Hartshorne, J. K., Hauser, T. U., Ho, M. K., et al. Using games to understand the mind. Nature Human Behaviour, pp.\ 1--9, 2024
2024
-
[4]
recaptcha: The brilliant business model that only one man could create, 2018
Anton. recaptcha: The brilliant business model that only one man could create, 2018. URL https://d3.harvard.edu/platform-digit/submission/recaptcha-the-brilliant-business-model-that-only-one-man-could-create/. Digital Innovation and Transformation, Posted on March 26, 2018
2018
-
[5]
P., Busby, E
Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., and Wingate, D. Out of one, many: Using language models to simulate human samples. Political Analysis, 31 0 (3): 0 337--351, 2023
2023
-
[6]
Ashok, D. and May, J. A little human data goes a long way. arXiv preprint arXiv:2410.13098, 2024
arXiv 2024
-
[7]
Your roomba may be mapping your home, collecting data that could be shared
Astor, M. Your roomba may be mapping your home, collecting data that could be shared. The New York Times, 25: 0 186, 2017
2017
-
[8]
Lionbridge vs appen: Which platform should you work for?, 2022
at Home Smart, W. Lionbridge vs appen: Which platform should you work for?, 2022. URL https://workathomesmart.com/lionbridge-vs-appen/. Accessed: 2025-01-19
2022
Show all 132 references
-
[9]
Bem, D. J. Self-perception theory. In Advances in experimental social psychology, volume 6, pp.\ 1--62. Elsevier, 1972
1972
-
[10]
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., D e biak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al. Dota 2 with large scale deep reinforcement learning. arXiv preprint arXiv:1912.06680, 2019
1912 arXiv
-
[11]
S., Brandt, J., Miller, R
Bernstein, M. S., Brandt, J., Miller, R. C., and Karger, D. R. Crowds in two seconds: Enabling realtime crowd-powered interfaces. In Proceedings of the 24th annual ACM symposium on User interface software and technology, pp.\ 33--42, 2011
2011
-
[12]
Social science for pennies
Bohannon, J. Social science for pennies. Science, 2011
2011
-
[13]
A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021
2021 arXiv
-
[14]
Protoqa: A question answering dataset for prototypical common-sense reasoning
Boratko, M., Li, X., O’Gorman, T., Das, R., Le, D., and Mccallum, A. Protoqa: A question answering dataset for prototypical common-sense reasoning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp.\ 1122--1136, 2020
2020
-
[15]
and Prentice, G
Brady, A. and Prentice, G. Are loot boxes addictive? analyzing participant’s physiological arousal while opening a loot box. Games and Culture, 16 0 (4): 0 419--433, 2021
2021
-
[16]
Labor and Monopoly Capital: The Degradation of Work in the Twentieth Century
Braverman, H. Labor and Monopoly Capital: The Degradation of Work in the Twentieth Century. Monthly Review Press, New York, 1974
1974
-
[17]
The rise of ai-generated content in wikipedia
Brooks, C., Eggert, S., and Peskoff, D. The rise of ai-generated content in wikipedia. In Proceedings of the First Workshop on Advancing Natural Language Processing for Wikipedia, pp.\ 67--79, 2024
2024
-
[18]
B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in Neural Information Processing Systems, 33: 0 1877--1901, 2020
1901
-
[19]
and McAfee, A
Brynjolfsson, E. and McAfee, A. The second machine age: Work, progress, and prosperity in a time of brilliant technologies. WW Norton & Company, 2014
2014
-
[20]
P., Bennert, N., Urry, C
Cardamone, C., Schawinski, K., Sarzi, M., Bamford, S. P., Bennert, N., Urry, C. M., Lintott, C., Keel, W. C., Parejko, J., Nichol, R. C., et al. Galaxy zoo green peas: discovery of a class of compact extremely star-forming galaxies. Monthly Notices of the Royal Astronomical So...
2009
-
[21]
data labelers
Cheng, M. Microsoft, google, and openai are getting questioned about their ai “data labelers”, Sep 2023. URL https://qz.com/tech-companies-ai-data-labelers-congress-1850834407
2023
-
[22]
Iconary: A pictionary-based game for testing multimodal communication with drawings and text
Clark, C., Salvador, J., Schwenk, D., Bonafilia, D., Yatskar, M., Kolve, E., Herrasti, A., Choi, J., Mehta, S., Skjonsberg, S., et al. Iconary: A pictionary-based game for testing multimodal communication with drawings and text. In Proceedings of the 2021 Conference on Empiric...
2021
-
[23]
Common Crawl Dataset , 2021
Common Crawl . Common Crawl Dataset , 2021. URL https://commoncrawl.org/
2021
-
[24]
Predicting protein structures with a multiplayer online game
Cooper, S., Khatib, F., Treuille, A., Barbero, J., Lee, J., Beenen, M., Leaver-Fay, A., Baker, D., Popovi \'c , Z., et al. Predicting protein structures with a multiplayer online game. Nature, 466 0 (7307): 0 756--760, 2010
2010
-
[25]
and Joler, V
Crawford, K. and Joler, V. Anatomy of an ai system. Anatomy of an AI System, 2018
2018
-
[26]
Reconsidering the trade-off between expertise and flexibility: A cognitive entrenchment perspective
Dane, E. Reconsidering the trade-off between expertise and flexibility: A cognitive entrenchment perspective. Academy of Management Review, 35 0 (4): 0 579--603, 2010. doi:10.5465/amr.35.4.zok579
2010 doi
-
[27]
Deci, E. L. Effects of externally mediated rewards on intrinsic motivation. Journal of personality and Social Psychology, 18 0 (1): 0 105, 1971
1971
-
[28]
L., Olafsen, A
Deci, E. L., Olafsen, A. H., and Ryan, R. M. Self-determination theory in work organizations: The state of a science. Annual review of organizational psychology and organizational behavior, 4 0 (1): 0 19--43, 2017
2017
-
[29]
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pp.\ 248--255. Ieee, 2009
2009
-
[30]
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technolog...
2019
-
[31]
Global shift: Mapping the changing contours of the world economy
Dicken, P. Global shift: Mapping the changing contours of the world economy. SAGE Publications Ltd, 2007
2007
-
[32]
D., Ewell, P
Douglas, B. D., Ewell, P. J., and Brauer, M. Data quality in online human-subjects research: Comparisons between mturk, prolific, cloudresearch, qualtrics, and sona. Plos one, 18 0 (3): 0 e0279720, 2023
2023
-
[33]
X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P
Dubois, Y., Li, C. X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P. S., and Hashimoto, T. B. Alpacafarm: A simulation framework for methods that learn from human feedback. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[34]
Real or fake text?: Investigating human ability to detect boundaries between human-written and machine-generated text
Dugan, L., Ippolito, D., Kirubarajan, A., Shi, S., and Callison-Burch, C. Real or fake text?: Investigating human ability to detect boundaries between human-written and machine-generated text. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pp.\ 12...
2023
-
[35]
(FAIR)†, M. F. A. R. D. T., Bakhtin, A., Brown, N., Dinan, E., Farina, G., Flaherty, C., Fried, D., Goff, A., Gray, J., Hu, H., et al. Human-level play in the game of diplomacy by combining language models with strategic reasoning. Science, 378 0 (6624): 0 1067--1074, 2022
2022
-
[36]
and Nehmad, E
Fogel, J. and Nehmad, E. Internet social network communities: Risk taking, trust, and privacy concerns. Computers in human behavior, 25 0 (1): 0 153--160, 2009
2009
-
[37]
and Bruckman, A
Forte, A. and Bruckman, A. Why do people write for wikipedia? incentives to contribute to open-content publishing. In Proceedings of the 2005 GROUP Conference. ACM, 2005
2005
-
[38]
Y., Ilharco, G., Fang, A., Hayase, J., Smyrnis, G., Nguyen, T., Marten, R., Wortsman, M., Ghosh, D., Zhang, J., et al
Gadre, S. Y., Ilharco, G., Fang, A., Hayase, J., Smyrnis, G., Nguyen, T., Marten, R., Wortsman, M., Ghosh, D., Zhang, J., et al. Datacomp: In search of the next generation of multimodal datasets. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[39]
Geng, S., Hsieh, C.-Y., Ramanujan, V., Wallingford, M., Li, C.-L., Koh, P. W. W., and Krishna, R. The unmet promise of synthetic training images: Using retrieved real images performs better. Advances in Neural Information Processing Systems, 37: 0 7902--7929, 2024
2024
-
[40]
\"U ber-alienated: Powerless and alone in the gig economy
Glavin, P., Bierman, A., and Schieman, S. \"U ber-alienated: Powerless and alone in the gig economy. Work and Occupations, 48 0 (4): 0 399--431, 2021
2021
-
[41]
and Cohen, V
Gokaslan, A. and Cohen, V. Openwebtext corpus. http://Skylion007.github.io/OpenWebTextCorpus, 2019
2019
-
[42]
Goodhart, C. A. and Goodhart, C. Problems of monetary management: the UK experience. Springer, 1984
1984
-
[43]
Gray, M. L. and Suri, S. Ghost work: How to stop Silicon Valley from building a new global underclass. Eamon Dolan Books, 2019
2019
-
[44]
Google uses esp game to tag images
Guardian, T. Google uses esp game to tag images. The Guardian, September 2006. URL https://www.theguardian.com/technology/blog/2006/sep/03/googleusesesp
2006
-
[45]
‘mass theft’: Thousands of artists call for ai art auction to be cancelled
Guardian, T. ‘mass theft’: Thousands of artists call for ai art auction to be cancelled. The Guardian, February 2025. URL https://www.theguardian.com/technology/2025/feb/10/mass-theft-thousands-of-artists-call-for-ai-art-auction-to-be-cancelled
2025
-
[46]
Gururangan, S., Card, D., Dreier, S., Gade, E., Wang, L., Wang, Z., Zettlemoyer, L., and Smith, N. A. Whose language counts as high quality? measuring language ideologies in text data selection. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Pro...
2022
-
[47]
and Eck, D
Ha, D. and Eck, D. A neural representation of sketch drawings. arXiv preprint arXiv:1704.03477, 2017
2017 arXiv
-
[48]
The unreasonable effectiveness of data
Halevy, A., Norvig, P., and Pereira, F. The unreasonable effectiveness of data. IEEE intelligent systems, 24 0 (2): 0 8--12, 2009
2009
-
[49]
How the ai industry profits from catastrophe
Hao, K. How the ai industry profits from catastrophe. MIT Technology Review, 2022. URL https://www.technologyreview.com/2022/04/20/1050392/ai-industry-appen-scale-data-labels/. Accessed: 2025-02-10
2022
-
[50]
Stack overflow bans users en masse for rebelling against openai partnership
Hardware, T. Stack overflow bans users en masse for rebelling against openai partnership. https://www.tomshardware.com/tech-industry/artificial-intelligence/stack-overflow-bans-users-en-masse-for-rebelling-against-openai-partnership-users-banned-for-deleting-answers-to-prevent...
2024
-
[51]
Ho, C.-J., Slivkins, A., Suri, S., and Vaughan, J. W. Incentivizing high quality crowdwork. In Proceedings of the 24th International Conference on World Wide Web, pp.\ 419--429, 2015
2015
-
[52]
A., Welbl, J., Clark, A., et al
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., et al. Training compute-optimal large language models. In Proceedings of the 36th International Conference on Neural Information Processi...
2022
-
[53]
Homans, G. C. Social behavior as exchange. American Journal of Sociology, 63 0 (6): 0 597--606, 1958. doi:10.1086/222355
1958 doi
-
[54]
From the American system to mass production, 1800-1932: The development of manufacturing technology in the United States
Hounshell, D. From the American system to mass production, 1800-1932: The development of manufacturing technology in the United States. Baltimore, Md.: Johns Hopkins University Press, 1984
1932
-
[55]
and Bernstein, M
Ikeda, K. and Bernstein, M. S. Pay it backward: Per-task payments on crowdsourcing platforms reduce productivity. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems, pp.\ 4111--4121, 2016
2016
-
[56]
Human or not? a gamified approach to the turing test
Jannai, D., Meron, A., Lenz, B., Levine, Y., and Shoham, Y. Human or not? a gamified approach to the turing test. arXiv preprint arXiv:2305.20010, 2023
2023 arXiv
-
[57]
H., Brown, L., Cheng, J., Khan, M., Gupta, A., Workman, D., Hanna, A., Flowers, J., and Gebru, T
Jiang, H. H., Brown, L., Cheng, J., Khan, M., Gupta, A., Workman, D., Hanna, A., Flowers, J., and Gebru, T. Ai art and its impact on artists. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, pp.\ 363--374, 2023
2023
-
[58]
Types of motivation affect study selection, attention, and dropouts in online experiments
Jun, E., Hsieh, G., and Reinecke, K. Types of motivation affect study selection, attention, and dropouts in online experiments. Proceedings of the ACM on Human-Computer Interaction, 1 0 (CSCW): 0 1--15, 2017
2017
-
[59]
and Tversky, A
Kahneman, D. and Tversky, A. Prospect theory: An analysis of decision under risk. In Handbook of the fundamentals of financial decision making: Part I, pp.\ 99--127. World Scientific, 2013
2013
-
[60]
B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020
2001 arXiv
-
[61]
Tesla ai day: Full self-driving and neural network training
Karpathy, A. Tesla ai day: Full self-driving and neural network training. Tesla AI Day Presentation, 2021. URL https://www.youtube.com/watch?v=j0z4FweCy4M
2021
-
[62]
On the folly of rewarding a, while hoping for b
Kerr, S. On the folly of rewarding a, while hoping for b. Academy of Management Journal, 18 0 (4): 0 769--783, 1975. doi:10.2307/255378
1975 doi
-
[63]
C., Group, F
Khatib, F., DiMaio, F., Group, F. C., Group, F. V. C., Cooper, S., Kazmierczyk, M., Gilski, M., Krzywda, S., Zabranska, H., Pichova, I., et al. Crystal structure of a monomeric retroviral protease solved by protein folding game players. Nature structural & molecular biology, 1...
2011
-
[64]
Dynabench: Rethinking benchmarking in nlp
Kiela, D., Bartolo, M., Nie, Y., Kaushik, D., Geiger, A., Wu, Z., Vidgen, B., Prasad, G., Singh, A., Ringshia, P., et al. Dynabench: Rethinking benchmarking in nlp. arXiv preprint arXiv:2104.14337, 2021
2021 arXiv
-
[65]
Soda: Million-scale dialogue distillation with social commonsense contextualization
Kim, H., Hessel, J., Jiang, L., West, P., Lu, X., Yu, Y., Zhou, P., Bras, R., Alikhani, M., Kim, G., et al. Soda: Million-scale dialogue distillation with social commonsense contextualization. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Proce...
2023
-
[66]
V., Bernstein, M., Gerber, E., Shaw, A., Zimmerman, J., Lease, M., and Horton, J
Kittur, A., Nickerson, J. V., Bernstein, M., Gerber, E., Shaw, A., Zimmerman, J., Lease, M., and Horton, J. The future of crowd work. In Proceedings of the 2013 conference on Computer supported cooperative work, pp.\ 1301--1318, 2013
2013
-
[67]
E., and Gurevych, I
Klie, J.-C., de Castilho, R. E., and Gurevych, I. Analyzing dataset annotation quality management in the wild. Computational Linguistics, pp.\ 1--48, 2024 a
2024
-
[68]
On efficient and statistical quality estimation for data annotation
Klie, J.-C., Haladjian, J., Kirchner, M., and Nair, R. On efficient and statistical quality estimation for data annotation. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp.\ 15680--15696, 2024 b
2024
-
[69]
Ai2-thor: An interactive 3d environment for visual ai
Kolve, E., Mottaghi, R., Han, W., VanderBilt, E., Weihs, L., Herrasti, A., Deitke, M., Ehsani, K., Gordon, D., Zhu, Y., et al. Ai2-thor: An interactive 3d environment for visual ai. arXiv preprint arXiv:1712.05474, 2017
2017 arXiv
-
[70]
A Theory of Fun for Game Design
Koster, R. A Theory of Fun for Game Design. Paraglyph Press, Scottsdale, AZ, 2005
2005
-
[71]
Kreitmeir, D. H. and Raschky, P. A. The unintended consequences of censoring digital technology--evidence from italy's chatgpt ban. arXiv preprint arXiv:2304.09339, 2023
2023 arXiv
-
[72]
Krizhevsky, A., Sutskever, I., and Hinton, G. E. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25, 2012
2012
-
[73]
C., and Bol, N
Kruikemeier, S., Boerman, S. C., and Bol, N. Breaching the contract? using social contract theory to explain individuals’ online behavior to safeguard privacy. Media Psychology, 23 0 (2): 0 269--292, 2020
2020
-
[74]
Motivations to participate in online communities
Lampe, C., Wash, R., Velasquez, A., and Ozkaya, E. Motivations to participate in online communities. In CHI '10 Extended Abstracts on Human Factors in Computing Systems. ACM, 2010. doi:10.1145/1753846.1753863
2010
-
[75]
Improving task instructions for data annotators: How clear rules and higher pay increase performance in data annotation in the ai economy
Laux, J., Stephany, F., and Liefgreen, A. Improving task instructions for data annotators: How clear rules and higher pay increase performance in data annotation in the ai economy. arXiv preprint arXiv:2312.14565v2, 2024. URL https://arxiv.org/abs/2312.14565v2
2024 arXiv
-
[76]
How eve online players saved real-world scientists 330 years of research on covid-19, May 2021
LeBlanc, W. How eve online players saved real-world scientists 330 years of research on covid-19, May 2021. URL https://www.ign.com/articles/how-eve-online-players-saved-real-world-scientists-330-years-of-research-on-covid-19
2021
-
[77]
Deduplicating training data makes language models better
Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C., and Carlini, N. Deduplicating training data makes language models better. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp.\ 8424--...
2022
-
[78]
Leonard, T. C. Richard h. thaler, cass r. sunstein, nudge: Improving decisions about health, wealth, and happiness: Yale university press, new haven, ct, 2008, 293 pp, 2008
2008
-
[79]
overjustification
Lepper, M. R., Greene, D., and Nisbett, R. E. Undermining children's intrinsic interest with extrinsic reward: A test of the" overjustification" hypothesis. Journal of Personality and social Psychology, 28 0 (1): 0 129, 1973
1973
-
[80]
Y., Bansal, H., Guha, E., Keh, S
Li, J., Fang, A., Smyrnis, G., Ivgi, M., Jordan, M., Gadre, S. Y., Bansal, H., Guha, E., Keh, S. S., Arora, K., et al. Datacomp-lm: In search of the next generation of training sets for language models. Advances in Neural Information Processing Systems, 37: 0 14200--14282, 2024
2024
-
[81]
Y.-Y., Lee, A.-H., Lo, K., Chang, J
Liao, Z., Antoniak, M., Cheong, I., Cheng, E. Y.-Y., Lee, A.-H., Lo, K., Chang, J. C., and Zhang, A. X. Llms as research tools: A large scale survey of researchers' usage and perceptions. arXiv preprint arXiv:2411.05025, 2024
2024 arXiv
-
[82]
J., Schawinski, K., Slosar, A., Land, K., Bamford, S., Thomas, D., Raddick, M
Lintott, C. J., Schawinski, K., Slosar, A., Land, K., Bamford, S., Thomas, D., Raddick, M. J., Nichol, R. C., Szalay, A., Andreescu, D., et al. Galaxy zoo: morphologies derived from visual inspection of galaxies from the sloan digital sky survey. Monthly Notices of the Royal A...
2008
-
[83]
How to calculate worker compensation for amazon mechanical turk, 2024
Malsburg, T. How to calculate worker compensation for amazon mechanical turk, 2024. URL https://tmalsburg.github.io/mturk-compensation.html. Accessed: 2025-01-19
2024
-
[84]
Martin, D., O’Neill, J., Gupta, N., and Hanrahan, B. V. Turking in a global labour market. Computer Supported Cooperative Work (CSCW), 25: 0 39--77, 2016
2016
-
[85]
Maslow, A. H. A theory of human motivation. Psychological review, 50 0 (4): 0 370, 1943
1943
-
[86]
performance of crowds
Mason, W. and Watts, D. J. Financial incentives and the" performance of crowds". In Proceedings of the ACM SIGKDD workshop on human computation, pp.\ 77--85, 2009
2009
-
[87]
M., and Song, D
Nair, V., Garrido, G. M., and Song, D. Exploring the unprecedented privacy risks of the metaverse. arXiv preprint arXiv:2207.13176, 2022
2022 arXiv
-
[88]
Collaborative dialogue in minecraft
Narayan-Chen, A., Jayannavar, P., and Hockenmaier, J. Collaborative dialogue in minecraft. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp.\ 5405--5415, 2019
2019
-
[89]
Quality not quantity: On the interaction between dataset design and robustness of clip
Nguyen, T., Ilharco, G., Wortsman, M., Oh, S., and Schmidt, L. Quality not quantity: On the interaction between dataset design and robustness of clip. Advances in Neural Information Processing Systems, 35: 0 21455--21469, 2022
2022
-
[90]
Training ai to be loyal
Oh, S., Tyagi, H., and Viswanath, P. Training ai to be loyal. arXiv preprint arXiv:2502.15720, 2025
2025 arXiv
-
[91]
Chatgpt: Openai language model
OpenAI. Chatgpt: Openai language model. https://chat.openai.com, 2023. Accessed: January 26, 2025
2023
-
[92]
Captcha if you can: how you’ve been training ai for years without realising it, 2018
O’Malley, J. Captcha if you can: how you’ve been training ai for years without realising it, 2018. URL https://www.techradar.com/news/captcha-if-you-can-how-youve-been-training-ai-for-years-without-realising-it
2018
-
[93]
S., Popowski, L., Cai, C., Morris, M
Park, J. S., Popowski, L., Cai, C., Morris, M. R., Liang, P., and Bernstein, M. S. Social simulacra: Creating populated prototypes for social computing systems. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology, pp.\ 1--18, 2022
2022
-
[94]
S., O'Brien, J., Cai, C
Park, J. S., O'Brien, J., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, pp.\ 1--22, 2023
2023
-
[95]
Exclusive: Openai used kenyan workers on less than \ 2 per hour to make chatgpt less toxic
Perrigo, B. Exclusive: Openai used kenyan workers on less than \ 2 per hour to make chatgpt less toxic. Time Magazine, 18: 0 2023, 2023
2023
-
[96]
Data scarcity: When will ai hit a wall?, 2025
Pieces. Data scarcity: When will ai hit a wall?, 2025. URL https://pieces.app/blog/data-scarcity-when-will-ai-hit-a-wall. Accessed: 2025-01-19
2025
-
[97]
Prolific vs
Prolific. Prolific vs. mturk, 2024. URL https://www.prolific.com/prolific-vs-mturk. Accessed: 2025-01-19
2024
-
[98]
D., Yang, T.-Y., Partsey, R., Desai, R., Clegg, A
Puig, X., Undersander, E., Szot, A., Cote, M. D., Yang, T.-Y., Partsey, R., Desai, R., Clegg, A. W., Hlavac, M., Min, S. Y., et al. Habitat 3.0: A co-habitat for humans, avatars and robots. arXiv preprint arXiv:2310.13724, 2023
-
[99]
X., Reif, E., Simon, G., Hussein, N., Clement, N., Wexler, J., Cai, C
Qian, C., Liu, M. X., Reif, E., Simon, G., Hussein, N., Clement, N., Wexler, J., Cai, C. J., Terry, M., and Kahng, M. The evolution of llm adoption in industry data curation practices. arXiv preprint arXiv:2412.16089, 2024
2024 arXiv
-
[100]
and Bonds-Raacke, J
Raacke, J. and Bonds-Raacke, J. Myspace and facebook: Applying the uses and gratifications theory to exploring friend-networking sites. Cyberpsychology & behavior, 11 0 (2): 0 169--174, 2008
2008
-
[101]
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21 0 (140): 0 1--67, 2020
2020
-
[102]
and Gajos, K
Reinecke, K. and Gajos, K. Z. Labinthewild: Conducting large-scale online experiments with uncompensated samples. In Proceedings of the 18th ACM conference on computer supported cooperative work & social computing, pp.\ 1364--1378, 2015
2015
-
[103]
Ruggiero, T. E. Uses and gratifications theory in the 21st century. Mass communication & society, 3 0 (1): 0 3--37, 2000
2000
-
[104]
An empirical study & evaluation of modern \ CAPTCHAs \
Searles, A., Nakatsuka, Y., Ozturk, E., Paverd, A., Tsudik, G., and Enkoji, A. An empirical study & evaluation of modern \ CAPTCHAs \ . In 32nd usenix security symposium (usenix security 23), pp.\ 3081--3097, 2023 a
2023
-
[105]
T., and Tsudik, G
Searles, A., Prapty, R. T., and Tsudik, G. Dazed & confused: A large-scale real-world user study of recaptchav2. arXiv preprint arXiv:2311.10911, 2023 b
2023 arXiv
-
[106]
Shah, N. B. and Zhou, D. Double or nothing: Multiplicative incentive mechanisms for crowdsourcing. Journal of Machine Learning Research, 17 0 (165): 0 1--52, 2016
2016
-
[107]
Ai models collapse when trained on recursively generated data
Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., and Gal, Y. Ai models collapse when trained on recursively generated data. Nature, 631 0 (8022): 0 755--759, 2024
2024
-
[108]
Der Kobra-Effekt: Wie man Irrwege der Wirtschaftspolitik vermeidet
Siebert, H. Der Kobra-Effekt: Wie man Irrwege der Wirtschaftspolitik vermeidet. Deutsche Verlags-Anstalt, Stuttgart, Germany, 2001
2001
-
[109]
and Sutton, R
Silver, D. and Sutton, R. S. The era of experience. https://storage.googleapis.com/deepmind-media/Era-of-Experience
-
[110]
J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. Mastering the game of go with deep neural networks and tree search. Nature, 529 0 (7587): 0 484--489, 2016
2016
-
[111]
Dolma: an open corpus of three trillion tokens for language model pretraining research
Soldaini, L., Kinney, R., Bhagia, A., Schwenk, D., Atkinson, D., Authur, R., Bogin, B., Chandu, K., Dumas, J., Elazar, Y., et al. Dolma: an open corpus of three trillion tokens for language model pretraining research. In Proceedings of the 62nd Annual Meeting of the Associatio...
2024
-
[112]
and Kuiper, K
Stafford, L. and Kuiper, K. Social exchange theories: Calculating the rewards and costs of personal relationships. In Engaging theories in interpersonal communication, pp.\ 379--390. Routledge, 2021
2021
-
[113]
Scalability in perception for autonomous driving: Waymo open dataset
Sun, P., Kretzschmar, H., Dotiwalla, X., Chouard, A., Patnaik, V., Tsui, P., Guo, J., Zhou, Y., Chai, Y., Caine, B., et al. Scalability in perception for autonomous driving: Waymo open dataset. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognitio...
2020
-
[114]
The bitter lesson
Sutton, R. The bitter lesson. Incomplete Ideas (blog), 13 0 (1): 0 38, 2019
2019
-
[115]
L., Bhagavatula, C., Goldberg, Y., Choi, Y., and Berant, J
Talmor, A., Yoran, O., Bras, R. L., Bhagavatula, C., Goldberg, Y., Choi, Y., and Berant, J. Commonsenseqa 2.0: Exposing the limits of ai through gamification. arXiv preprint arXiv:2201.05320, 2022
2022 arXiv
-
[116]
and Hashimoto, T
Taori, R. and Hashimoto, T. Data feedback loops: Model-driven amplification of dataset biases. In International Conference on Machine Learning, pp.\ 33883--33920. PMLR, 2023
2023
-
[117]
Stack overflow users sabotage their posts after openai deal
Technica, A. Stack overflow users sabotage their posts after openai deal. https://arstechnica.com/information-technology/2024/05/stack-overflow-users-sabotage-their-posts-after-openai-deal/, 2024. Accessed: 2025-01-03
2024
-
[118]
Tesla's approach to autonomous driving: Autopilot ai and full self-driving
Tesla, I. Tesla's approach to autonomous driving: Autopilot ai and full self-driving. Online, 2021. URL https://www.tesla.com/autopilotAI
2021
-
[119]
Chatgpt is fun, but not an author
Thorp, H. Chatgpt is fun, but not an author. Science, 379 0 (6630): 0 313, 2023. doi:10.1126/science.adg7879. URL https://www.science.org/doi/10.1126/science.adg7879
2023 doi
-
[120]
Upwork vs
Upwork. Upwork vs. fiverr: An in-depth comparison, 2024. URL https://www.upwork.com/resources/upwork-vs-fiverr. Accessed: 2025-01-19
2024
-
[121]
H., and West, R
Veselovsky, V., Ribeiro, M. H., and West, R. Artificial artificial artificial intelligence: Crowd workers widely use large language models for text production tasks. arXiv preprint arXiv:2306.07899, 2023
2023 arXiv
-
[122]
J., Gordon, A., Rothschild, D., and West, R
Veselovsky, V., Horta Ribeiro, M., Cozzolino, P. J., Gordon, A., Rothschild, D., and West, R. Prevalence and prevention of large language model use in crowd work. Communications of the ACM, 68 0 (3): 0 42--47, 2025
2025
-
[123]
M., Mathieu, M., Dudzik, A., Chung, J., Choi, D
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al. Grandmaster level in starcraft ii using multi-agent reinforcement learning. Nature, 575 0 (7782): 0 350--354, 2019
2019
-
[124]
and Zou, J
Vodrahalli, K. and Zou, J. Artwhisperer: A dataset for characterizing human-ai interactions in artistic creations. arXiv preprint arXiv:2306.08141, 2023
2023 arXiv
-
[125]
Games with a purpose
Von Ahn, L. Games with a purpose. Computer, 39 0 (6): 0 92--94, 2006
2006
-
[126]
and Dabbish, L
Von Ahn, L. and Dabbish, L. Esp: Labeling images with a computer game. In AAAI spring symposium: Knowledge collection from volunteer contributors, volume 2, pp.\ 1, 2005
2005
-
[127]
Verbosity: a game for collecting common-sense facts
Von Ahn, L., Kedia, M., and Blum, M. Verbosity: a game for collecting common-sense facts. In Proceedings of the SIGCHI conference on Human Factors in computing systems, pp.\ 75--78, 2006 a
2006
-
[128]
Peekaboom: a game for locating objects in images
Von Ahn, L., Liu, R., and Blum, M. Peekaboom: a game for locating objects in images. In Proceedings of the SIGCHI conference on Human Factors in computing systems, pp.\ 55--64, 2006 b
2006
-
[129]
The exploited labor behind artificial intelligence
Williams, A., Miceli, M., and Gebru, T. The exploited labor behind artificial intelligence. Noema Magazine, 22, 2022
2022
-
[130]
Characteristics of gamers who purchase loot box: A systematic literature review
Yokomitsu, K., Irie, T., Shinkawa, H., and Tanaka, M. Characteristics of gamers who purchase loot box: A systematic literature review. Current Addiction Reports, 8: 0 481--493, 2021
2021
-
[131]
Lima: Less is more for alignment
Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., et al. Lima: Less is more for alignment. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[132]
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In Proceedings of the IEEE international conference on computer vision, pp.\ 19--27, 2015
2015
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.