REVIEW 5 minor 36 references
Artificial Effort
T0 review · 0 major / 5 minor · reviewed 2026-07-12 · grok-4.5
Pith's one-line read Most real-effort tasks used in economics experiments can already be solved by LLMs accurately and cheaply, so unsupervised scores may no longer measure human effort.
desk verdict Clean empirical boundary condition: most canonical real-effort tasks are already automatable by mid-tier multimodal LLMs at negligible cost, with null incentive effects. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
A standardized screenshot-to-API pipeline that feeds each oTree task instance (instructions plus rendered image) to 23 multimodal LLMs under fixed token and time limits, scoring exact-match accuracy over 20 runs per model–task pair and converting token use into dollar cost.
What would settle it
Measure actual substitution rates and detection rates when real online participants are free to use commercial LLMs or browser agents on the same eight tasks, and check whether the high accuracy and positive net-gain numbers survive under those conditions.
Extended reading notes
Core claim
Most of the eight canonical real-effort tasks can be solved accurately by current LLMs at a cost far below typical online piece rates, while verbal monetary incentives and human-persona prompts leave LLM accuracy unchanged; therefore, when participants can cheaply outsource task completion, observed performance may no longer measure genuine human effort.
Load-bearing premise
That the screenshot-plus-API pipeline and its cost figures are a realistic lower bound on what a profit-seeking online participant can actually achieve under real platform constraints.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper evaluates whether eight canonical real-effort tasks used in experimental economics remain valid measures of human effort when multimodal LLMs can complete them. Using 23 models from OpenAI, Google, and Anthropic (accessed via OpenRouter), 20 independent runs per model–task under uniform token/time limits, and exact-match scoring against oTree ground truth, the authors show that Addition, Pair Summation, Letter Decoding, and Sequence Completion are near-ceiling for most models, while Sudoku, Counting Zeros, and String Entry remain harder. Accuracy rises with model generation and mid-tier models close the gap with frontier ones. API costs leave a large positive net gain relative to a $0.25-per-correct piece-rate benchmark; an “optimal participant” who cherry-picks the best model per task reaches ~93% average accuracy. A 2×2 verbal-incentive × human-persona design (T0–T4), analyzed with a binomial GLMM with model and task fixed effects, finds no detectable effect of incentive language or persona framing. The authors conclude that, in unsupervised settings, observed performance may no longer reflect genuine human effort and offer practical recommendations (prefer perceptual tasks, use incentive responsiveness and within-session dynamics as diagnostics).
Significance. If the results hold, the paper supplies a clear, timely boundary condition for a workhorse tool in experimental economics: most of the tested real-effort tasks are already automatable at high accuracy and negligible cost, and verbal incentives—central to the logic of real-effort designs—do not move LLM accuracy. Strengths include preregistration, public code and data, a transparent cost accounting, a multi-provider multi-tier design that documents rapid mid-tier catch-up, and a properly specified null-incentive analysis. The contribution is empirical and methodological rather than theoretical, but it is directly actionable for online and unsupervised experiments and for the design of future effort tasks.
minor comments (5)
- Section 2.2 and footnote 5 correctly note that full browser automation is feasible and would not raise token cost, but the manuscript could more explicitly flag that measured accuracy/cost is a lower bound on what a sophisticated participant could achieve (and that prevalence and detection risk are left for future work).
- Figure 1 pools all 23 models; a short note that the ranking of task difficulty is stable when restricted to top-tier models would help readers who care only about frontier capability.
- The $0.25 piece-rate benchmark is described as conservative; a one-sentence citation or range from recent Prolific/MTurk real-effort studies would make the external benchmark fully transparent.
- Minor presentation: a few model names in the heatmaps (e.g., “gemini-3.1-flash-lite”) and the chronological ordering within tiers could be cross-checked against Table 5 for consistency of release dates and labels.
- The discussion of stationary accuracy as a diagnostic is useful; a brief caveat that temperature/sampling settings or multi-agent wrappers could reintroduce variance would avoid overclaiming uniqueness of the flat profile.
Circularity Check
No circularity: direct empirical measurement of LLM accuracy and API cost on fixed tasks; no fitted parameters re-presented as predictions.
full rationale
The paper's central claims rest on direct measurement: 20 independent runs per model–task under exact-match scoring, token counts and OpenRouter prices used to compute API cost, and a binomial GLM with model/task fixed effects for the five prompt treatments. Accuracy is the fraction of exact matches; cost is the sum of priced input and output tokens; net gain is that cost subtracted from an external $0.25 piece-rate benchmark taken from typical online-experiment rates, not estimated from the LLM data. Verbal-incentive null results are likewise measured contrasts (T0–T4), not derived from a fitted behavioral parameter. There is no self-definitional loop, no parameter fitted on a subset and then called a prediction, no uniqueness theorem imported from the authors' prior work, and no ansatz smuggled via self-citation. The derivation chain is therefore self-contained empirical reporting; circularity score is zero.
Assumptions & free parameters
free parameters (2)
- piece_rate_benchmark =
$0.25 per correct answer
- max_output_tokens_and_timeout =
2048 tokens / 120 s
assumptions (4)
- domain assumption Real-effort tasks measure costly human cognitive exertion only when a human actually performs them.
- domain assumption Screenshot-plus-API (or equivalent browser automation) is a feasible and low-cost way for an online participant to outsource a task.
- ad hoc to paper Exact string match to ground truth is the appropriate correctness criterion (no partial credit).
- domain assumption API calls produce i.i.d. responses with no within-session learning or fatigue.
invented entities (1)
-
optimal participant (cherry-picking best model per task)
Cite this review
Pith. "Pith review of Artificial Effort." pith.science (2026). https://pith.science/paper/E2HVE3H6
@misc{pith2026260523920,
author = {Pith},
title = {Pith review of: Artificial Effort},
year = {2026},
howpublished = {\url{https://pith.science/paper/E2HVE3H6}},
note = {Machine review of arXiv:2605.23920}
}
read the original abstract
Real-effort tasks, in which participants perform cognitively costly activities whose outcomes depend on actual performance, are widely used in experimental economics. Their validity, however, rests on the assumption that a human performs them. We study whether this assumption still holds in the era of Artificial Intelligence (AI) and Large Language Models (LLMs). Using 8 canonical real-effort tasks and 23 LLMs from three major providers, we show that most tasks can now be solved accurately and at a negligible cost, while only a few resist automation. Performance improves with each model generation, and midtier models are rapidly closing the gap with frontier ones, broadening the set of widely accessible models that can automate these tasks. Additionally, we show that verbally offering monetary incentives has no effect on LLM performance. Our findings establish a boundary condition for the use of real-effort tasks in unsupervised settings: when participants can cheaply outsource task completion to an LLM, observed performance may no longer reflect genuine human effort.
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[1]
Abeler, J., A. Falk, L. Goette, and D. Huffman (2011): Reference points and effort provision, American Economic Review, 101, 470--492
2011
-
[2]
Anthropic (2023): Introducing Claude , https://www.anthropic.com/news/introducing-claude, accessed: 2026-04-13
2023
-
[3]
Niederle, and C
Augenblick, N., M. Niederle, and C. Sprenger (2015): Working over time: Dynamic inconsistency in real effort tasks, The Quarterly Journal of Economics, 130, 1067--1115
2015
-
[4]
Wiegmann, E
Bevendorff, J., M. Wiegmann, E. Richter, M. Potthast, and B. Stein (2025): The Two Paradigms of LLM Detection: Authorship Attribution vs. Authorship Verification, in Findings of the Association for Computational Linguistics: ACL 2025, ed. by W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, Vienna, Austria: Association for Computational Linguistics, 3762--3787
2025
-
[5]
Bracha, A. and C. Fershtman (2013): Competitive Incentives: Working Harder or Working Smarter? Management Science, 59, 771--781
2013
-
[6]
Franke, and P
Calsamiglia, C., J. Franke, and P. Rey-Biel (2013): The incentive effects of affirmative action in a real-effort tournament, Journal of Public Economics, 98, 15--31
2013
-
[7]
Carpenter, J. and E. Huet-Vaughn (2019): 19. Real-effort tasks, Handbook of research methods and applications in experimental economics, 368
2019
-
[8]
Exley, S
Celebi, C., C. Exley, S. Harrs, H. Kivimaki, M. Serra-Garcia, and J. Yusof (2026): Mission Possible: The Collection of High-Quality Data,
2026
Show all 36 references
-
[9]
Gneezy, and A
Charness, G., U. Gneezy, and A. Henderson (2018): Experimental methods: Measuring effort in economics experiments, Journal of Economic Behavior & Organization, 149, 74--87
2018
-
[10]
Chen, D. L., M. Schonger, and C. Wickens (2016): oTree—An open-source platform for laboratory, online, and field experiments, Journal of Behavioral and Experimental Finance, 9, 88--97
2016
-
[11]
Tworek, H
Chen, M., J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al. (2021): Evaluating large language models trained on code, arXiv preprint arXiv:2107.03374
2021 arXiv
-
[12]
Cowhey, O
Clark, P., I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord (2018): Think you have solved question answering? Try ARC, the AI2 reasoning challenge, arXiv preprint arXiv:1803.05457
2018 arXiv
-
[13]
Kosaraju, M
Cobbe, K., V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al. (2021): Training verifiers to solve math word problems, arXiv preprint arXiv:2110.14168
2021 arXiv
-
[14]
Ding, Z., G. Deng, Y. Liu, J. Ding, J. Chen, Y. Sui, and Y. Li (2025): IllusionCAPTCHA: A CAPTCHA based on Visual Illusion,
2025
-
[15]
Gangadharan, and N
Erkal, N., L. Gangadharan, and N. Nikiforakis (2011): Relative earnings and giving in a real-effort experiment, American Economic Review, 101, 3330--3348
2011
-
[16]
Ferrando, J
Fu, T., R. Ferrando, J. Conde, C. Arriaga, and P. Reviriego (2024): Why Do Large Language Models ( LLMs ) Struggle to Count Letters? arXiv preprint arXiv:2412.18626
2024 arXiv
-
[17]
Burns, S
Hendrycks, D., C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt (2021 a ): Measuring Massive Multitask Language Understanding, in International Conference on Learning Representations
2021
-
[18]
Burns, S
Hendrycks, D., C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt (2021 b ): Measuring Mathematical Problem Solving With the MATH Dataset, in Thirty-Fifth Conference on Neural Information Processing Systems
2021
-
[19]
Jimenez, C. E., J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan (2024): SWE-bench: Can Language Models Resolve Real-World GitHub Issues? in The Twelfth International Conference on Learning Representations
2024
-
[20]
Joshi, M., E. Choi, D. S. Weld, and L. Zettlemoyer (2017): TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension, in Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, 1601--1611
2017
-
[21]
Kessler, J. B. and M. I. Norton (2016): Tax aversion in labor supply, Journal of Economic Behavior & Organization, 124, 15--28
2016
-
[22]
Palomaki, O
Kwiatkowski, T., J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, et al. (2019): Natural questions: a benchmark for question answering research, Transactions of the Association for Computational Linguistics, 7, 453--466
2019
-
[23]
Bottou, Y
LeCun, Y., L. Bottou, Y. Bengio, and P. Haffner (2002): Gradient-based learning applied to document recognition, Proceedings of the IEEE, 86, 2278--2324
2002
-
[24]
Alghamdi, M
Maiya, A., R. Alghamdi, M. L. Pacheco, A. Trivedi, and F. Somenzi (2025): Explaining Puzzle Solutions in Natural Language: An Exploratory Study on 6x6 Sudoku, in Findings of the Association for Computational Linguistics: ACL 2025, ed. by W. Che, J. Nabende, E. Shutova, and M. ...
2025
-
[25]
Niederle, M. and L. Vesterlund (2007): Do women shy away from competition? Do men compete too much? The quarterly journal of economics, 122, 1067--1101
2007
-
[26]
Jiang, and J
Nogueira, R., Z. Jiang, and J. Lin (2021): Investigating the Limitations of Transformers with Simple Arithmetic Tasks, arXiv preprint arXiv:2102.13019
2021 arXiv
-
[27]
OpenAI (2022): Introducing ChatGPT , https://openai.com/blog/chatgpt, accessed: 2026-04-13
2022
-
[28]
Pichai, S. and D. Hassabis (2023): Introducing Gemini : Our Largest and Most Capable AI Model, https://blog.google/technology/ai/google-gemini-ai/, accessed: 2026-04-13
2023
-
[29]
Slonim, and M
Rosaz, J., R. Slonim, and M. C. Villeval (2016): Quitting and peer effects at work, Labour Economics, 39, 55--67
2016
-
[30]
Srivastava, A. et al. (2022): Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models, arXiv preprint arXiv:2206.04615
2022 arXiv
-
[31]
Team, G. et al. (2023): Gemini: a family of highly capable multimodal models, arXiv preprint arXiv:2312.11805
2023 arXiv
-
[32]
Wang, J., C. Zhu, Y. Zhou, L. Li, X. He, and J. Xiong (2025): COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA Solvers,
2025
-
[33]
Wei, J., X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. H. Chi, Q. V. Le, and D. Zhou (2022): Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, in Advances in Neural Information Processing Systems, vol. 35, 24824--24837
2022
-
[34]
Kaplan, G
Yehudai, G., H. Kaplan, G. Dar, R. Rassin, A. Ghandeharioun, M. Geva, and A. Globerson (2024): When Can Transformers Count to n ? arXiv preprint arXiv:2407.15160
2024
-
[35]
Holtzman, Y
Zellers, R., A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi (2019): Hellaswag: Can a machine really finish your sentence? in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 3472--3482
2019
-
[36]
Zhang, J., Z. Zhou, X. Ji, S. Liu, and Z. Zhao (2025): CAPTURE: A Benchmark and Evaluation for LVLMs in CAPTCHA Resolving,
2025
Reviewed July 12, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.