Pith. sign in

REVIEW 4 major objections 3 minor 132 references

When Incentives Backfire, Data Stops Being Human

T0 review · 4 major / 3 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read The paper argues that the contamination of AI training data by AI-generated content is not primarily a filtering problem but a design problem: data collection systems built around task-based pay and fine-grained control quietly destroy…

desk verdict Clearly argued and honest position piece, but the load-bearing causal claim is an untested extrapolation the authors themselves concede. read the letter →

arxiv 2502.07732 v2 pith:H3P32UXC submitted 2025-02-11 cs.CY cs.AIcs.CLcs.CVcs.HCcs.LG

classification cs.CYcs.AIcs.CLcs.CVcs.HCcs.LG
keywords humandatacollectionintrinsicmotivationmotivationalcrowding-outcrowdsourcingLLMcontaminationqualitygameswithapurposeincentivedesign
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that the contamination of AI training data by AI-generated content is not primarily a filtering problem but a design problem: data collection systems built around task-based pay and fine-grained control quietly destroy the intrinsic motivation that produces rich, authentic human data. It proposes that the next generation of data sourcing should be built to align with contributors' intrinsic motivations, with games and other structured but voluntary environments as the model. The authors read the quantity-quality trade-off in data collection as a Pareto frontier that poor incentive design pulls inward rather than a fixed constraint. If they are right, the practical fix for the AI data crisis is better design of human data collection, not better detectors of synthetic text.

What carries the argument

The carrying mechanism is motivational crowding-out, the psychological process by which external rewards shift a contributor's self-perception from 'I do this because I enjoy it or it matters' to 'I do this for the pay,' reducing effort, creativity, and long-term engagement. The paper operationalizes this through the overjustification effect (Lepper et al., 1973) and self-determination theory (Deci et al., 2017), and pairs it with a design variable called the resolution of control: how finely a data collection system constrains contributor behavior. High-resolution, task-level control maximizes throughput but triggers crowding-out; low-resolution, environment-level control preserves motivation but risks incoherent data; the paper argues that sustainable designs sit in the middle, using structure and rules to elicit behavior without transactional surveillance.

What would settle it

A controlled field experiment on a mainstream annotation platform that pays one group per task and another a flat participation fee (or none) for the same annotation workload, measuring output quality, diversity, and LLM-use rate over several weeks; if the per-task group matches or beats the intrinsic-motivation group on quality and engagement, the paper's central premise fails.

Watch

Extended reading notes

Core claim

The central claim is that over-reliance on external incentives and task fragmentation, the two standard engineering levers of crowdsourcing, erode the intrinsic motivation that sustains high-quality human data, and that this erosion, not the mere presence of LLMs, is what makes modern data collection fragile. The paper assembles evidence from psychology (overjustification, self-perception theory, self-determination theory) and economics (Goodhart's law, perverse incentives) and maps it onto the structure of MTurk-style platforms, where it predicts a vicious cycle: as intrinsic motivation fades, platforms tighten control and raise pay, which accelerates shortcutting behavior such as using LLMs to complete tasks, further degrading data quality. The authors propose a 'resolution of control' spectrum, with tightly controlled task-level systems at one end and laissez-faire community platforms at the other, and locate the desirable design space in the middle, exemplified by product-integrated systems and by games. The paper's positive thesis is that structured environments based on voluntary participation, such as games with a purpose, citizen science, and community platforms, can deliver high-quality, high-quantity data while preserving contributor trust.

Load-bearing premise

The argument stands on the transfer of motivational crowding-out from laboratory and workplace settings to large-scale machine-learning data collection, specifically that lowering task-level financial incentives will improve rather than reduce data quality and quantity.

Editorial extensions

If this is right

  • AI data pipelines should be redesigned around autonomy, competence, and relatedness rather than piece-rate payment.
  • Detecting and filtering AI-generated content is a stopgap; the durable fix is making human contribution itself harder to replace.
  • Fragmented micro-tasking into ever smaller units should be treated as a known risk to data quality, not a neutral scaling tool.
  • Games with a purpose and citizen-science platforms become a first-class source of training data rather than a curiosity.
  • Trust, the sense that contributors' data is used legitimately, becomes a third axis of data system design alongside quality and quantity.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper, the argument implies that the marginal value of higher annotation pay may be negative for exactly the tasks where authenticity matters most, such as preference data and human-behavior corpora.
  • Beyond the paper, the framing suggests a testable prediction: platforms that add transparent recognition, feedback, and autonomy will see lower LLM-copying rates at equal pay.
  • Beyond the paper, the same crowding-out logic applies to data contributors inside companies and to voluntary community data, which would extend the paper's scope beyond crowdsourcing marketplaces.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 3 minor

Summary. This position paper argues that the rise of LLM-generated content in human data collection—both on crowdwork platforms and on the open Internet—is not primarily a filtering problem but a symptom of flawed system design. Drawing on psychological theories (overjustification, self-perception theory, self-determination theory) and economic concepts (Goodhart's Law, perverse incentives), the authors claim that excessive reliance on external incentives and task fragmentation erodes intrinsic motivation, which in turn degrades data quality and pushes contributors toward LLM shortcutting. They propose shifting from task-level control to a 'resolution-of-control spectrum' with a desirable middle ground, and they advocate games as a promising form of structured, intrinsically motivating data collection. They also discuss trust, compensation, and ethical considerations. The paper contains no new empirical data; it is an argumentative synthesis that frames a research agenda for rethinking data collection systems.

Significance. If the thesis holds, it would reframe an active area of ML research: instead of developing better filters for AI-generated content, the community would invest in designing data collection systems that preserve intrinsic motivation. This is a timely and important direction, and the paper brings together a broad, relevant literature from social psychology, economics, and human-computer interaction. Its honest acknowledgment that no ML data-collection game has yet demonstrated sustained success (Section 8.2) is a strength, as is its clear elucidation of the overjustification mechanism. However, the central causal chain—from incentive design to motivational crowding-out, and from crowding-out to LLM contamination—is asserted rather than tested, and a key empirical citation is used inconsistently. The paper is best read as a hypothesis-rich agenda-setting piece, but it currently overstates the confidence in its own remedy.

major comments (4)
  1. [Section 5 and Section 6] The paper cites Mason & Watts (2009) in Section 5 as evidence that financial compensation leads to greater effort and better quantity and/or quality, but in Section 6 it cites the same paper (along with Ikeda & Bernstein, 2016) to support the claim that piece-rate or pay-per-task systems degrade output quality. Mason & Watts actually reported that higher pay increased quantity without reducing quality, and Ikeda & Bernstein found that per-task payments reduced productivity. These findings do not support the crowding-out transfer in the direction the paper needs, and using the same citation for opposing claims weakens the empirical foundation of the argument.
  2. [Section 5, crowdwork versus community platforms] The central causal claim that MTurk-style platforms suffer a 'crowding-out effect' over time, leading to LLM-assisted or automated annotation, is not supported by the cited evidence. The overjustification experiments (Lepper et al., 1973) and self-determination theory (Deci et al., 2017) come from controlled lab or workplace settings; the paper does not provide evidence that crowdwork contributors experience motivational crowding-out, nor that the rise in LLM use documented by Veselovsky et al. (2023, 2025) is caused by that crowding-out rather than by macroeconomic pressures, task design, or the simple availability of cheap AI tools. This missing link is load-bearing for the paper's thesis.
  3. [Section 8.2] The paper's own admission that 'no data collection games in machine learning have yet demonstrated sustained success' directly undermines the proposed remedy. The successful examples appealed to next—Zooniverse, Foldit, Lab in the Wild—are citizen-science platforms with self-selected volunteers who have strong domain interest; the paper does not show that such models can scale to the volumes or task types that ML annotation pipelines require, especially for tedious, low-intrinsic-appeal tasks. Without at least one concrete case study or small-scale controlled comparison, the claim that games can expand the quality–quantity frontier for ML data remains an untested hypothesis.
  4. [Section 7.1 and Figure 4] The 'resolution-of-control spectrum' is introduced as a central design variable, but it is not operationally defined: there are no criteria for placing a system on the spectrum, no measurement instrument, and no way to determine when a design has medium resolution. As a result, the recommendation to design for the 'middle' is not falsifiable in its present form. The paper should either specify observable features that determine spectrum position or explicitly frame the spectrum as a heuristic that generates testable comparative predictions.
minor comments (3)
  1. [Throughout] There are frequent typographical spacing errors in names such as 'V on Ahn', 'GW AP', 'F oldIt', 'V odrahalli', and 'Kahneman & Tversky'; these should be corrected.
  2. [Section 6] The discussion of System 1 and System 2 is attributed to Kahneman & Tversky (2013), but the cited reference is 'Prospect Theory'; the appropriate citation is Kahneman's 'Thinking, Fast and Slow' or his joint work on dual-process models.
  3. [Figure 1] The perpetual donkey machine analogy is evocative but does not clearly map onto the paper's argument; consider replacing the figure with a schematic of the overjustification mechanism or adding a caption that explicitly connects the two.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper makes no fitted predictions and its argument is grounded in external psychology and economics literature.

full rationale

This is an argumentative position paper, not a derivation: there are no fitted parameters, equations, or quantitative predictions whose outputs could reduce to their inputs. The central claim—that overreliance on external incentives can crowd out intrinsic motivation and degrade human data quality—is supported by external experimental literature (Lepper et al. 1973; Deci et al. 2017; Bem 1972) and by independent empirical studies of LLM use in crowdwork (Veselovsky et al. 2023; 2025), even though one of the latter's authors overlaps with the present paper. That self-citation is not load-bearing: it documents the prevalence of LLM use rather than supplying the motivational argument, and the paper does not invoke any author-specific uniqueness theorem or ansatz. The quality-quantity frontier and resolution-of-control spectrum are organizing frameworks, not results derived from the paper's own definitions; Section 8.2 even concedes that no ML data-collection game has yet demonstrated sustained success, which is a limitation rather than a circular step. Whether the psychological crowding-out effect transfers to large-scale ML annotation is an untested empirical extrapolation, but that is a correctness risk, not circularity. Accordingly no circular step is present.

Assumptions & free parameters 0 free parameters · 5 assumptions · 1 invented entities

The paper introduces no fitted parameters. It rests on several domain assumptions imported from psychology and economics, most importantly that overjustification generalizes to crowdwork and that intrinsic motivation can be preserved at scale in structured systems. The resolution-of-control spectrum is a new conceptual lens without independent empirical grounding.

assumptions (5)
  • domain assumption The overjustification effect, demonstrated in children drawing, transfers to adult crowdworkers and long-term platform participation.
    Section 5 applies Lepper et al. (1973) to data collection platforms without direct evidence from those platforms.
  • domain assumption Intrinsic motivation, rather than compensation, is the primary driver of sustained high-quality data contributions.
    Sections 5 and 7 treat this as established from Wikipedia and Reddit examples; no controlled comparison is provided.
  • domain assumption The quantity-quality tradeoff is a Pareto frontier that system design can expand.
    Section 3 and Figure 2 introduce the frontier conceptually; no formal model or data support the claim that it can be expanded.
  • ad hoc to paper Games can simultaneously provide structured data and preserve intrinsic motivation at scale.
    Section 8 cites GWAP, Zooniverse, and FoldIt, but the paper acknowledges that no ML data collection game has yet demonstrated sustained success.
  • domain assumption High-quality data is best proxied by naturalness, defined as unprompted, incentive-free human behavior.
    Section 4 proposes naturalness as a grounding for quality; this is a value-laden definition, not a measured one.
invented entities (1)
  • Resolution-of-control spectrum
    purpose: Classify data collection systems by how tightly contributor behavior is constrained, from environment-level to task-level control.
    Introduced in Section 7.1 and Figure 4 as a design variable. It is a framing device rather than an empirical entity, and the paper provides no falsifiable handle independent of its own examples.

how reviews work

0 comments
Cite this review

Pith. "Pith review of When Incentives Backfire, Data Stops Being Human." pith.science (2026). https://pith.science/paper/H3P32UXC

@misc{pith2026250207732,
  author       = {Pith},
  title        = {Pith review of: When Incentives Backfire, Data Stops Being Human},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/H3P32UXC}},
  note         = {Machine review of arXiv:2502.07732}
}
read the original abstract

Progress in AI has relied on human-generated data, from annotator marketplaces to the wider Internet. However, the widespread use of large language models now threatens the quality and integrity of human-generated data on these very platforms. We argue that this issue goes beyond the immediate challenge of filtering AI-generated content -- it reveals deeper flaws in how data collection systems are designed. Existing systems often prioritize speed, scale, and efficiency at the cost of intrinsic human motivation, leading to declining engagement and data quality. We propose that rethinking data collection systems to align with contributors' intrinsic motivations -- rather than relying solely on external incentives -- can help sustain high-quality data sourcing at scale while maintaining contributor trust and long-term participation.

Figures

Figures reproduced from arXiv: 2502.07732 by the authors.

Figure 1
Figure 1. Perpetual Donkey Machine. It looks like the donkey could walk forever with the carrot just out of reach. But it won’t, not forever. Reward a task the donkey would never do otherwise, and you get shortcuts – actions optimized only to reach the carrot. Reward a task it already does, and you risk erasing the inner drive that moved it in the first place – making it less donkey. Good incentives shape action. Flawed ones … view at source ↗
Figure 2
Figure 2. Illustration of a quantity-quality trade-off in data col￾lection systems. Popular crowdwork platforms (e.g., MTurk / Microtask, Prolific / Survey, and UpWork / Gig) tend optimize for either scale or quality but struggle to achieve both at the same time. In contrast, data from sources not explicitly designed for collection, such as online collectives and communities (e.g., Wikipedia and Reddit), operate outside this … view at source ↗
Figure 3
Figure 3. Intrinsic vs. Extrinsic Motivation: Internally motivated contributors are likely to produce human-like and diverse outputs, grounded in creativity and engagement. Externally motivated systems tend to favor controllability, structure and efficiency, often resulting in more uniform outputs that follow clearly defined goals. The figure illustrates how different motivational contexts can shape the nature and trajectory … view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: The Resolution-of-Control Spectrum. The horizontal axis orders human data pipelines by how tightly contributors’ be￾havior is constrained, from environment-level, low control (left) to task-level, high control (right). Low-control systems include online knowledge bases…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

132 extracted references · 52 canonical work pages

  1. [1]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...

  2. [2]

    M., Longpre, S., Lambert, N., Wang, X., Muennighoff, N., Hou, B., Pan, L., Jeong, H., et al

    Albalak, A., Elazar, Y., Xie, S. M., Longpre, S., Lambert, N., Wang, X., Muennighoff, N., Hou, B., Pan, L., Jeong, H., et al. A survey on data selection for language models. arXiv preprint arXiv:2402.16827, 2024

  3. [3]

    E., Gershman, S

    Allen, K., Br \"a ndle, F., Botvinick, M., Fan, J. E., Gershman, S. J., Gopnik, A., Griffiths, T. L., Hartshorne, J. K., Hauser, T. U., Ho, M. K., et al. Using games to understand the mind. Nature Human Behaviour, pp.\ 1--9, 2024

  4. [4]

    recaptcha: The brilliant business model that only one man could create, 2018

    Anton. recaptcha: The brilliant business model that only one man could create, 2018. URL https://d3.harvard.edu/platform-digit/submission/recaptcha-the-brilliant-business-model-that-only-one-man-could-create/. Digital Innovation and Transformation, Posted on March 26, 2018

  5. [5]

    P., Busby, E

    Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., and Wingate, D. Out of one, many: Using language models to simulate human samples. Political Analysis, 31 0 (3): 0 337--351, 2023

  6. [6]

    and May, J

    Ashok, D. and May, J. A little human data goes a long way. arXiv preprint arXiv:2410.13098, 2024

  7. [7]

    Your roomba may be mapping your home, collecting data that could be shared

    Astor, M. Your roomba may be mapping your home, collecting data that could be shared. The New York Times, 25: 0 186, 2017

  8. [8]

    Lionbridge vs appen: Which platform should you work for?, 2022

    at Home Smart, W. Lionbridge vs appen: Which platform should you work for?, 2022. URL https://workathomesmart.com/lionbridge-vs-appen/. Accessed: 2025-01-19

Show all 132 references
  1. [9]

    Bem, D. J. Self-perception theory. In Advances in experimental social psychology, volume 6, pp.\ 1--62. Elsevier, 1972

  2. [10]

    Dota 2 with large scale deep reinforcement learning

    Berner, C., Brockman, G., Chan, B., Cheung, V., D e biak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al. Dota 2 with large scale deep reinforcement learning. arXiv preprint arXiv:1912.06680, 2019

  3. [11]

    S., Brandt, J., Miller, R

    Bernstein, M. S., Brandt, J., Miller, R. C., and Karger, D. R. Crowds in two seconds: Enabling realtime crowd-powered interfaces. In Proceedings of the 24th annual ACM symposium on User interface software and technology, pp.\ 33--42, 2011

  4. [12]

    Social science for pennies

    Bohannon, J. Social science for pennies. Science, 2011

  5. [13]

    A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M

    Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021

  6. [14]

    Protoqa: A question answering dataset for prototypical common-sense reasoning

    Boratko, M., Li, X., O’Gorman, T., Das, R., Le, D., and Mccallum, A. Protoqa: A question answering dataset for prototypical common-sense reasoning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp.\ 1122--1136, 2020

  7. [15]

    and Prentice, G

    Brady, A. and Prentice, G. Are loot boxes addictive? analyzing participant’s physiological arousal while opening a loot box. Games and Culture, 16 0 (4): 0 419--433, 2021

  8. [16]

    Labor and Monopoly Capital: The Degradation of Work in the Twentieth Century

    Braverman, H. Labor and Monopoly Capital: The Degradation of Work in the Twentieth Century. Monthly Review Press, New York, 1974

  9. [17]

    The rise of ai-generated content in wikipedia

    Brooks, C., Eggert, S., and Peskoff, D. The rise of ai-generated content in wikipedia. In Proceedings of the First Workshop on Advancing Natural Language Processing for Wikipedia, pp.\ 67--79, 2024

  10. [18]

    B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

    Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in Neural Information Processing Systems, 33: 0 1877--1901, 2020

  11. [19]

    and McAfee, A

    Brynjolfsson, E. and McAfee, A. The second machine age: Work, progress, and prosperity in a time of brilliant technologies. WW Norton & Company, 2014

  12. [20]

    P., Bennert, N., Urry, C

    Cardamone, C., Schawinski, K., Sarzi, M., Bamford, S. P., Bennert, N., Urry, C. M., Lintott, C., Keel, W. C., Parejko, J., Nichol, R. C., et al. Galaxy zoo green peas: discovery of a class of compact extremely star-forming galaxies. Monthly Notices of the Royal Astronomical So...

  13. [21]

    data labelers

    Cheng, M. Microsoft, google, and openai are getting questioned about their ai “data labelers”, Sep 2023. URL https://qz.com/tech-companies-ai-data-labelers-congress-1850834407

  14. [22]

    Iconary: A pictionary-based game for testing multimodal communication with drawings and text

    Clark, C., Salvador, J., Schwenk, D., Bonafilia, D., Yatskar, M., Kolve, E., Herrasti, A., Choi, J., Mehta, S., Skjonsberg, S., et al. Iconary: A pictionary-based game for testing multimodal communication with drawings and text. In Proceedings of the 2021 Conference on Empiric...

  15. [23]

    Common Crawl Dataset , 2021

    Common Crawl . Common Crawl Dataset , 2021. URL https://commoncrawl.org/

  16. [24]

    Predicting protein structures with a multiplayer online game

    Cooper, S., Khatib, F., Treuille, A., Barbero, J., Lee, J., Beenen, M., Leaver-Fay, A., Baker, D., Popovi \'c , Z., et al. Predicting protein structures with a multiplayer online game. Nature, 466 0 (7307): 0 756--760, 2010

  17. [25]

    and Joler, V

    Crawford, K. and Joler, V. Anatomy of an ai system. Anatomy of an AI System, 2018

  18. [26]

    Reconsidering the trade-off between expertise and flexibility: A cognitive entrenchment perspective

    Dane, E. Reconsidering the trade-off between expertise and flexibility: A cognitive entrenchment perspective. Academy of Management Review, 35 0 (4): 0 579--603, 2010. doi:10.5465/amr.35.4.zok579

  19. [27]

    Deci, E. L. Effects of externally mediated rewards on intrinsic motivation. Journal of personality and Social Psychology, 18 0 (1): 0 105, 1971

  20. [28]

    L., Olafsen, A

    Deci, E. L., Olafsen, A. H., and Ryan, R. M. Self-determination theory in work organizations: The state of a science. Annual review of organizational psychology and organizational behavior, 4 0 (1): 0 19--43, 2017

  21. [29]

    Imagenet: A large-scale hierarchical image database

    Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pp.\ 248--255. Ieee, 2009

  22. [30]

    Bert: Pre-training of deep bidirectional transformers for language understanding

    Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technolog...

  23. [31]

    Global shift: Mapping the changing contours of the world economy

    Dicken, P. Global shift: Mapping the changing contours of the world economy. SAGE Publications Ltd, 2007

  24. [32]

    D., Ewell, P

    Douglas, B. D., Ewell, P. J., and Brauer, M. Data quality in online human-subjects research: Comparisons between mturk, prolific, cloudresearch, qualtrics, and sona. Plos one, 18 0 (3): 0 e0279720, 2023

  25. [33]

    X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P

    Dubois, Y., Li, C. X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P. S., and Hashimoto, T. B. Alpacafarm: A simulation framework for methods that learn from human feedback. Advances in Neural Information Processing Systems, 36, 2024

  26. [34]

    Real or fake text?: Investigating human ability to detect boundaries between human-written and machine-generated text

    Dugan, L., Ippolito, D., Kirubarajan, A., Shi, S., and Callison-Burch, C. Real or fake text?: Investigating human ability to detect boundaries between human-written and machine-generated text. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pp.\ 12...

  27. [35]

    (FAIR)†, M. F. A. R. D. T., Bakhtin, A., Brown, N., Dinan, E., Farina, G., Flaherty, C., Fried, D., Goff, A., Gray, J., Hu, H., et al. Human-level play in the game of diplomacy by combining language models with strategic reasoning. Science, 378 0 (6624): 0 1067--1074, 2022

  28. [36]

    and Nehmad, E

    Fogel, J. and Nehmad, E. Internet social network communities: Risk taking, trust, and privacy concerns. Computers in human behavior, 25 0 (1): 0 153--160, 2009

  29. [37]

    and Bruckman, A

    Forte, A. and Bruckman, A. Why do people write for wikipedia? incentives to contribute to open-content publishing. In Proceedings of the 2005 GROUP Conference. ACM, 2005

  30. [38]

    Y., Ilharco, G., Fang, A., Hayase, J., Smyrnis, G., Nguyen, T., Marten, R., Wortsman, M., Ghosh, D., Zhang, J., et al

    Gadre, S. Y., Ilharco, G., Fang, A., Hayase, J., Smyrnis, G., Nguyen, T., Marten, R., Wortsman, M., Ghosh, D., Zhang, J., et al. Datacomp: In search of the next generation of multimodal datasets. Advances in Neural Information Processing Systems, 36, 2024

  31. [39]

    Geng, S., Hsieh, C.-Y., Ramanujan, V., Wallingford, M., Li, C.-L., Koh, P. W. W., and Krishna, R. The unmet promise of synthetic training images: Using retrieved real images performs better. Advances in Neural Information Processing Systems, 37: 0 7902--7929, 2024

  32. [40]

    \"U ber-alienated: Powerless and alone in the gig economy

    Glavin, P., Bierman, A., and Schieman, S. \"U ber-alienated: Powerless and alone in the gig economy. Work and Occupations, 48 0 (4): 0 399--431, 2021

  33. [41]

    and Cohen, V

    Gokaslan, A. and Cohen, V. Openwebtext corpus. http://Skylion007.github.io/OpenWebTextCorpus, 2019

  34. [42]

    Goodhart, C. A. and Goodhart, C. Problems of monetary management: the UK experience. Springer, 1984

  35. [43]

    Gray, M. L. and Suri, S. Ghost work: How to stop Silicon Valley from building a new global underclass. Eamon Dolan Books, 2019

  36. [44]

    Google uses esp game to tag images

    Guardian, T. Google uses esp game to tag images. The Guardian, September 2006. URL https://www.theguardian.com/technology/blog/2006/sep/03/googleusesesp

  37. [45]

    ‘mass theft’: Thousands of artists call for ai art auction to be cancelled

    Guardian, T. ‘mass theft’: Thousands of artists call for ai art auction to be cancelled. The Guardian, February 2025. URL https://www.theguardian.com/technology/2025/feb/10/mass-theft-thousands-of-artists-call-for-ai-art-auction-to-be-cancelled

  38. [46]

    Gururangan, S., Card, D., Dreier, S., Gade, E., Wang, L., Wang, Z., Zettlemoyer, L., and Smith, N. A. Whose language counts as high quality? measuring language ideologies in text data selection. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Pro...

  39. [47]

    and Eck, D

    Ha, D. and Eck, D. A neural representation of sketch drawings. arXiv preprint arXiv:1704.03477, 2017

  40. [48]

    The unreasonable effectiveness of data

    Halevy, A., Norvig, P., and Pereira, F. The unreasonable effectiveness of data. IEEE intelligent systems, 24 0 (2): 0 8--12, 2009

  41. [49]

    How the ai industry profits from catastrophe

    Hao, K. How the ai industry profits from catastrophe. MIT Technology Review, 2022. URL https://www.technologyreview.com/2022/04/20/1050392/ai-industry-appen-scale-data-labels/. Accessed: 2025-02-10

  42. [50]

    Stack overflow bans users en masse for rebelling against openai partnership

    Hardware, T. Stack overflow bans users en masse for rebelling against openai partnership. https://www.tomshardware.com/tech-industry/artificial-intelligence/stack-overflow-bans-users-en-masse-for-rebelling-against-openai-partnership-users-banned-for-deleting-answers-to-prevent...

  43. [51]

    Ho, C.-J., Slivkins, A., Suri, S., and Vaughan, J. W. Incentivizing high quality crowdwork. In Proceedings of the 24th International Conference on World Wide Web, pp.\ 419--429, 2015

  44. [52]

    A., Welbl, J., Clark, A., et al

    Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., et al. Training compute-optimal large language models. In Proceedings of the 36th International Conference on Neural Information Processi...

  45. [53]

    Homans, G. C. Social behavior as exchange. American Journal of Sociology, 63 0 (6): 0 597--606, 1958. doi:10.1086/222355

  46. [54]

    From the American system to mass production, 1800-1932: The development of manufacturing technology in the United States

    Hounshell, D. From the American system to mass production, 1800-1932: The development of manufacturing technology in the United States. Baltimore, Md.: Johns Hopkins University Press, 1984

  47. [55]

    and Bernstein, M

    Ikeda, K. and Bernstein, M. S. Pay it backward: Per-task payments on crowdsourcing platforms reduce productivity. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems, pp.\ 4111--4121, 2016

  48. [56]

    Human or not? a gamified approach to the turing test

    Jannai, D., Meron, A., Lenz, B., Levine, Y., and Shoham, Y. Human or not? a gamified approach to the turing test. arXiv preprint arXiv:2305.20010, 2023

  49. [57]

    H., Brown, L., Cheng, J., Khan, M., Gupta, A., Workman, D., Hanna, A., Flowers, J., and Gebru, T

    Jiang, H. H., Brown, L., Cheng, J., Khan, M., Gupta, A., Workman, D., Hanna, A., Flowers, J., and Gebru, T. Ai art and its impact on artists. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, pp.\ 363--374, 2023

  50. [58]

    Types of motivation affect study selection, attention, and dropouts in online experiments

    Jun, E., Hsieh, G., and Reinecke, K. Types of motivation affect study selection, attention, and dropouts in online experiments. Proceedings of the ACM on Human-Computer Interaction, 1 0 (CSCW): 0 1--15, 2017

  51. [59]

    and Tversky, A

    Kahneman, D. and Tversky, A. Prospect theory: An analysis of decision under risk. In Handbook of the fundamentals of financial decision making: Part I, pp.\ 99--127. World Scientific, 2013

  52. [60]

    B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D

    Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020

  53. [61]

    Tesla ai day: Full self-driving and neural network training

    Karpathy, A. Tesla ai day: Full self-driving and neural network training. Tesla AI Day Presentation, 2021. URL https://www.youtube.com/watch?v=j0z4FweCy4M

  54. [62]

    On the folly of rewarding a, while hoping for b

    Kerr, S. On the folly of rewarding a, while hoping for b. Academy of Management Journal, 18 0 (4): 0 769--783, 1975. doi:10.2307/255378

  55. [63]

    C., Group, F

    Khatib, F., DiMaio, F., Group, F. C., Group, F. V. C., Cooper, S., Kazmierczyk, M., Gilski, M., Krzywda, S., Zabranska, H., Pichova, I., et al. Crystal structure of a monomeric retroviral protease solved by protein folding game players. Nature structural & molecular biology, 1...

  56. [64]

    Dynabench: Rethinking benchmarking in nlp

    Kiela, D., Bartolo, M., Nie, Y., Kaushik, D., Geiger, A., Wu, Z., Vidgen, B., Prasad, G., Singh, A., Ringshia, P., et al. Dynabench: Rethinking benchmarking in nlp. arXiv preprint arXiv:2104.14337, 2021

  57. [65]

    Soda: Million-scale dialogue distillation with social commonsense contextualization

    Kim, H., Hessel, J., Jiang, L., West, P., Lu, X., Yu, Y., Zhou, P., Bras, R., Alikhani, M., Kim, G., et al. Soda: Million-scale dialogue distillation with social commonsense contextualization. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Proce...

  58. [66]

    V., Bernstein, M., Gerber, E., Shaw, A., Zimmerman, J., Lease, M., and Horton, J

    Kittur, A., Nickerson, J. V., Bernstein, M., Gerber, E., Shaw, A., Zimmerman, J., Lease, M., and Horton, J. The future of crowd work. In Proceedings of the 2013 conference on Computer supported cooperative work, pp.\ 1301--1318, 2013

  59. [67]

    E., and Gurevych, I

    Klie, J.-C., de Castilho, R. E., and Gurevych, I. Analyzing dataset annotation quality management in the wild. Computational Linguistics, pp.\ 1--48, 2024 a

  60. [68]

    On efficient and statistical quality estimation for data annotation

    Klie, J.-C., Haladjian, J., Kirchner, M., and Nair, R. On efficient and statistical quality estimation for data annotation. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp.\ 15680--15696, 2024 b

  61. [69]

    Ai2-thor: An interactive 3d environment for visual ai

    Kolve, E., Mottaghi, R., Han, W., VanderBilt, E., Weihs, L., Herrasti, A., Deitke, M., Ehsani, K., Gordon, D., Zhu, Y., et al. Ai2-thor: An interactive 3d environment for visual ai. arXiv preprint arXiv:1712.05474, 2017

  62. [70]

    A Theory of Fun for Game Design

    Koster, R. A Theory of Fun for Game Design. Paraglyph Press, Scottsdale, AZ, 2005

  63. [71]

    Kreitmeir, D. H. and Raschky, P. A. The unintended consequences of censoring digital technology--evidence from italy's chatgpt ban. arXiv preprint arXiv:2304.09339, 2023

  64. [72]

    Krizhevsky, A., Sutskever, I., and Hinton, G. E. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25, 2012

  65. [73]

    C., and Bol, N

    Kruikemeier, S., Boerman, S. C., and Bol, N. Breaching the contract? using social contract theory to explain individuals’ online behavior to safeguard privacy. Media Psychology, 23 0 (2): 0 269--292, 2020

  66. [74]

    Motivations to participate in online communities

    Lampe, C., Wash, R., Velasquez, A., and Ozkaya, E. Motivations to participate in online communities. In CHI '10 Extended Abstracts on Human Factors in Computing Systems. ACM, 2010. doi:10.1145/1753846.1753863

  67. [75]

    Improving task instructions for data annotators: How clear rules and higher pay increase performance in data annotation in the ai economy

    Laux, J., Stephany, F., and Liefgreen, A. Improving task instructions for data annotators: How clear rules and higher pay increase performance in data annotation in the ai economy. arXiv preprint arXiv:2312.14565v2, 2024. URL https://arxiv.org/abs/2312.14565v2

  68. [76]

    How eve online players saved real-world scientists 330 years of research on covid-19, May 2021

    LeBlanc, W. How eve online players saved real-world scientists 330 years of research on covid-19, May 2021. URL https://www.ign.com/articles/how-eve-online-players-saved-real-world-scientists-330-years-of-research-on-covid-19

  69. [77]

    Deduplicating training data makes language models better

    Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C., and Carlini, N. Deduplicating training data makes language models better. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp.\ 8424--...

  70. [78]

    Leonard, T. C. Richard h. thaler, cass r. sunstein, nudge: Improving decisions about health, wealth, and happiness: Yale university press, new haven, ct, 2008, 293 pp, 2008

  71. [79]

    overjustification

    Lepper, M. R., Greene, D., and Nisbett, R. E. Undermining children's intrinsic interest with extrinsic reward: A test of the" overjustification" hypothesis. Journal of Personality and social Psychology, 28 0 (1): 0 129, 1973

  72. [80]

    Y., Bansal, H., Guha, E., Keh, S

    Li, J., Fang, A., Smyrnis, G., Ivgi, M., Jordan, M., Gadre, S. Y., Bansal, H., Guha, E., Keh, S. S., Arora, K., et al. Datacomp-lm: In search of the next generation of training sets for language models. Advances in Neural Information Processing Systems, 37: 0 14200--14282, 2024

  73. [81]

    Y.-Y., Lee, A.-H., Lo, K., Chang, J

    Liao, Z., Antoniak, M., Cheong, I., Cheng, E. Y.-Y., Lee, A.-H., Lo, K., Chang, J. C., and Zhang, A. X. Llms as research tools: A large scale survey of researchers' usage and perceptions. arXiv preprint arXiv:2411.05025, 2024

  74. [82]

    J., Schawinski, K., Slosar, A., Land, K., Bamford, S., Thomas, D., Raddick, M

    Lintott, C. J., Schawinski, K., Slosar, A., Land, K., Bamford, S., Thomas, D., Raddick, M. J., Nichol, R. C., Szalay, A., Andreescu, D., et al. Galaxy zoo: morphologies derived from visual inspection of galaxies from the sloan digital sky survey. Monthly Notices of the Royal A...

  75. [83]

    How to calculate worker compensation for amazon mechanical turk, 2024

    Malsburg, T. How to calculate worker compensation for amazon mechanical turk, 2024. URL https://tmalsburg.github.io/mturk-compensation.html. Accessed: 2025-01-19

  76. [84]

    Martin, D., O’Neill, J., Gupta, N., and Hanrahan, B. V. Turking in a global labour market. Computer Supported Cooperative Work (CSCW), 25: 0 39--77, 2016

  77. [85]

    Maslow, A. H. A theory of human motivation. Psychological review, 50 0 (4): 0 370, 1943

  78. [86]

    performance of crowds

    Mason, W. and Watts, D. J. Financial incentives and the" performance of crowds". In Proceedings of the ACM SIGKDD workshop on human computation, pp.\ 77--85, 2009

  79. [87]

    M., and Song, D

    Nair, V., Garrido, G. M., and Song, D. Exploring the unprecedented privacy risks of the metaverse. arXiv preprint arXiv:2207.13176, 2022

  80. [88]

    Collaborative dialogue in minecraft

    Narayan-Chen, A., Jayannavar, P., and Hockenmaier, J. Collaborative dialogue in minecraft. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp.\ 5405--5415, 2019

  81. [89]

    Quality not quantity: On the interaction between dataset design and robustness of clip

    Nguyen, T., Ilharco, G., Wortsman, M., Oh, S., and Schmidt, L. Quality not quantity: On the interaction between dataset design and robustness of clip. Advances in Neural Information Processing Systems, 35: 0 21455--21469, 2022

  82. [90]

    Training ai to be loyal

    Oh, S., Tyagi, H., and Viswanath, P. Training ai to be loyal. arXiv preprint arXiv:2502.15720, 2025

  83. [91]

    Chatgpt: Openai language model

    OpenAI. Chatgpt: Openai language model. https://chat.openai.com, 2023. Accessed: January 26, 2025

  84. [92]

    Captcha if you can: how you’ve been training ai for years without realising it, 2018

    O’Malley, J. Captcha if you can: how you’ve been training ai for years without realising it, 2018. URL https://www.techradar.com/news/captcha-if-you-can-how-youve-been-training-ai-for-years-without-realising-it

  85. [93]

    S., Popowski, L., Cai, C., Morris, M

    Park, J. S., Popowski, L., Cai, C., Morris, M. R., Liang, P., and Bernstein, M. S. Social simulacra: Creating populated prototypes for social computing systems. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology, pp.\ 1--18, 2022

  86. [94]

    S., O'Brien, J., Cai, C

    Park, J. S., O'Brien, J., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, pp.\ 1--22, 2023

  87. [95]

    Exclusive: Openai used kenyan workers on less than \ 2 per hour to make chatgpt less toxic

    Perrigo, B. Exclusive: Openai used kenyan workers on less than \ 2 per hour to make chatgpt less toxic. Time Magazine, 18: 0 2023, 2023

  88. [96]

    Data scarcity: When will ai hit a wall?, 2025

    Pieces. Data scarcity: When will ai hit a wall?, 2025. URL https://pieces.app/blog/data-scarcity-when-will-ai-hit-a-wall. Accessed: 2025-01-19

  89. [97]

    Prolific vs

    Prolific. Prolific vs. mturk, 2024. URL https://www.prolific.com/prolific-vs-mturk. Accessed: 2025-01-19

  90. [98]

    D., Yang, T.-Y., Partsey, R., Desai, R., Clegg, A

    Puig, X., Undersander, E., Szot, A., Cote, M. D., Yang, T.-Y., Partsey, R., Desai, R., Clegg, A. W., Hlavac, M., Min, S. Y., et al. Habitat 3.0: A co-habitat for humans, avatars and robots. arXiv preprint arXiv:2310.13724, 2023

  91. [99]

    X., Reif, E., Simon, G., Hussein, N., Clement, N., Wexler, J., Cai, C

    Qian, C., Liu, M. X., Reif, E., Simon, G., Hussein, N., Clement, N., Wexler, J., Cai, C. J., Terry, M., and Kahng, M. The evolution of llm adoption in industry data curation practices. arXiv preprint arXiv:2412.16089, 2024

  92. [100]

    and Bonds-Raacke, J

    Raacke, J. and Bonds-Raacke, J. Myspace and facebook: Applying the uses and gratifications theory to exploring friend-networking sites. Cyberpsychology & behavior, 11 0 (2): 0 169--174, 2008

  93. [101]

    Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21 0 (140): 0 1--67, 2020

  94. [102]

    and Gajos, K

    Reinecke, K. and Gajos, K. Z. Labinthewild: Conducting large-scale online experiments with uncompensated samples. In Proceedings of the 18th ACM conference on computer supported cooperative work & social computing, pp.\ 1364--1378, 2015

  95. [103]

    Ruggiero, T. E. Uses and gratifications theory in the 21st century. Mass communication & society, 3 0 (1): 0 3--37, 2000

  96. [104]

    An empirical study & evaluation of modern \ CAPTCHAs \

    Searles, A., Nakatsuka, Y., Ozturk, E., Paverd, A., Tsudik, G., and Enkoji, A. An empirical study & evaluation of modern \ CAPTCHAs \ . In 32nd usenix security symposium (usenix security 23), pp.\ 3081--3097, 2023 a

  97. [105]

    T., and Tsudik, G

    Searles, A., Prapty, R. T., and Tsudik, G. Dazed & confused: A large-scale real-world user study of recaptchav2. arXiv preprint arXiv:2311.10911, 2023 b

  98. [106]

    Shah, N. B. and Zhou, D. Double or nothing: Multiplicative incentive mechanisms for crowdsourcing. Journal of Machine Learning Research, 17 0 (165): 0 1--52, 2016

  99. [107]

    Ai models collapse when trained on recursively generated data

    Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., and Gal, Y. Ai models collapse when trained on recursively generated data. Nature, 631 0 (8022): 0 755--759, 2024

  100. [108]

    Der Kobra-Effekt: Wie man Irrwege der Wirtschaftspolitik vermeidet

    Siebert, H. Der Kobra-Effekt: Wie man Irrwege der Wirtschaftspolitik vermeidet. Deutsche Verlags-Anstalt, Stuttgart, Germany, 2001

  101. [109]

    and Sutton, R

    Silver, D. and Sutton, R. S. The era of experience. https://storage.googleapis.com/deepmind-media/Era-of-Experience

  102. [110]

    J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al

    Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. Mastering the game of go with deep neural networks and tree search. Nature, 529 0 (7587): 0 484--489, 2016

  103. [111]

    Dolma: an open corpus of three trillion tokens for language model pretraining research

    Soldaini, L., Kinney, R., Bhagia, A., Schwenk, D., Atkinson, D., Authur, R., Bogin, B., Chandu, K., Dumas, J., Elazar, Y., et al. Dolma: an open corpus of three trillion tokens for language model pretraining research. In Proceedings of the 62nd Annual Meeting of the Associatio...

  104. [112]

    and Kuiper, K

    Stafford, L. and Kuiper, K. Social exchange theories: Calculating the rewards and costs of personal relationships. In Engaging theories in interpersonal communication, pp.\ 379--390. Routledge, 2021

  105. [113]

    Scalability in perception for autonomous driving: Waymo open dataset

    Sun, P., Kretzschmar, H., Dotiwalla, X., Chouard, A., Patnaik, V., Tsui, P., Guo, J., Zhou, Y., Chai, Y., Caine, B., et al. Scalability in perception for autonomous driving: Waymo open dataset. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognitio...

  106. [114]

    The bitter lesson

    Sutton, R. The bitter lesson. Incomplete Ideas (blog), 13 0 (1): 0 38, 2019

  107. [115]

    L., Bhagavatula, C., Goldberg, Y., Choi, Y., and Berant, J

    Talmor, A., Yoran, O., Bras, R. L., Bhagavatula, C., Goldberg, Y., Choi, Y., and Berant, J. Commonsenseqa 2.0: Exposing the limits of ai through gamification. arXiv preprint arXiv:2201.05320, 2022

  108. [116]

    and Hashimoto, T

    Taori, R. and Hashimoto, T. Data feedback loops: Model-driven amplification of dataset biases. In International Conference on Machine Learning, pp.\ 33883--33920. PMLR, 2023

  109. [117]

    Stack overflow users sabotage their posts after openai deal

    Technica, A. Stack overflow users sabotage their posts after openai deal. https://arstechnica.com/information-technology/2024/05/stack-overflow-users-sabotage-their-posts-after-openai-deal/, 2024. Accessed: 2025-01-03

  110. [118]

    Tesla's approach to autonomous driving: Autopilot ai and full self-driving

    Tesla, I. Tesla's approach to autonomous driving: Autopilot ai and full self-driving. Online, 2021. URL https://www.tesla.com/autopilotAI

  111. [119]

    Chatgpt is fun, but not an author

    Thorp, H. Chatgpt is fun, but not an author. Science, 379 0 (6630): 0 313, 2023. doi:10.1126/science.adg7879. URL https://www.science.org/doi/10.1126/science.adg7879

  112. [120]

    Upwork vs

    Upwork. Upwork vs. fiverr: An in-depth comparison, 2024. URL https://www.upwork.com/resources/upwork-vs-fiverr. Accessed: 2025-01-19

  113. [121]

    H., and West, R

    Veselovsky, V., Ribeiro, M. H., and West, R. Artificial artificial artificial intelligence: Crowd workers widely use large language models for text production tasks. arXiv preprint arXiv:2306.07899, 2023

  114. [122]

    J., Gordon, A., Rothschild, D., and West, R

    Veselovsky, V., Horta Ribeiro, M., Cozzolino, P. J., Gordon, A., Rothschild, D., and West, R. Prevalence and prevention of large language model use in crowd work. Communications of the ACM, 68 0 (3): 0 42--47, 2025

  115. [123]

    M., Mathieu, M., Dudzik, A., Chung, J., Choi, D

    Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al. Grandmaster level in starcraft ii using multi-agent reinforcement learning. Nature, 575 0 (7782): 0 350--354, 2019

  116. [124]

    and Zou, J

    Vodrahalli, K. and Zou, J. Artwhisperer: A dataset for characterizing human-ai interactions in artistic creations. arXiv preprint arXiv:2306.08141, 2023

  117. [125]

    Games with a purpose

    Von Ahn, L. Games with a purpose. Computer, 39 0 (6): 0 92--94, 2006

  118. [126]

    and Dabbish, L

    Von Ahn, L. and Dabbish, L. Esp: Labeling images with a computer game. In AAAI spring symposium: Knowledge collection from volunteer contributors, volume 2, pp.\ 1, 2005

  119. [127]

    Verbosity: a game for collecting common-sense facts

    Von Ahn, L., Kedia, M., and Blum, M. Verbosity: a game for collecting common-sense facts. In Proceedings of the SIGCHI conference on Human Factors in computing systems, pp.\ 75--78, 2006 a

  120. [128]

    Peekaboom: a game for locating objects in images

    Von Ahn, L., Liu, R., and Blum, M. Peekaboom: a game for locating objects in images. In Proceedings of the SIGCHI conference on Human Factors in computing systems, pp.\ 55--64, 2006 b

  121. [129]

    The exploited labor behind artificial intelligence

    Williams, A., Miceli, M., and Gebru, T. The exploited labor behind artificial intelligence. Noema Magazine, 22, 2022

  122. [130]

    Characteristics of gamers who purchase loot box: A systematic literature review

    Yokomitsu, K., Irie, T., Shinkawa, H., and Tanaka, M. Characteristics of gamers who purchase loot box: A systematic literature review. Current Addiction Reports, 8: 0 481--493, 2021

  123. [131]

    Lima: Less is more for alignment

    Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., et al. Lima: Less is more for alignment. Advances in Neural Information Processing Systems, 36, 2024

  124. [132]

    Aligning books and movies: Towards story-like visual explanations by watching movies and reading books

    Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In Proceedings of the IEEE international conference on computer vision, pp.\ 19--27, 2015

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.