Pith. sign in

REVIEW 5 minor 36 references

Artificial Effort

T0 review · 0 major / 5 minor · reviewed 2026-07-12 · grok-4.5

Pith's one-line read Most real-effort tasks used in economics experiments can already be solved by LLMs accurately and cheaply, so unsupervised scores may no longer measure human effort.

desk verdict Clean empirical boundary condition: most canonical real-effort tasks are already automatable by mid-tier multimodal LLMs at negligible cost, with null incentive effects. read the letter →

arxiv 2605.23920 v1 pith:E2HVE3H6 submitted 2026-04-17 cs.CY cs.AI

classification cs.CYcs.AI
keywords LargeLanguageModelsReal-efforttasksExperimentaleconomicsAutomationOnlineexperimentsIncentivesConstructvalidity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Real-effort tasks are meant to measure costly human performance in experimental economics. This paper tests whether that assumption still holds once large language models are available. It runs eight standard tasks against 23 multimodal models from three major providers and finds that four of the tasks are already solved near-perfectly at negligible cost, while only a few remain hard. Newer and mid-tier models keep closing the gap with frontier ones, and verbally offered money leaves model accuracy unchanged. The practical upshot is a boundary condition: in unsupervised online settings, a participant who can outsource the task to an LLM can produce high scores that no longer reflect genuine human effort.

What carries the argument

A standardized screenshot-to-API pipeline that feeds each oTree task instance (instructions plus rendered image) to 23 multimodal LLMs under fixed token and time limits, scoring exact-match accuracy over 20 runs per model–task pair and converting token use into dollar cost.

What would settle it

Measure actual substitution rates and detection rates when real online participants are free to use commercial LLMs or browser agents on the same eight tasks, and check whether the high accuracy and positive net-gain numbers survive under those conditions.

Watch

Extended reading notes

Core claim

Most of the eight canonical real-effort tasks can be solved accurately by current LLMs at a cost far below typical online piece rates, while verbal monetary incentives and human-persona prompts leave LLM accuracy unchanged; therefore, when participants can cheaply outsource task completion, observed performance may no longer measure genuine human effort.

Load-bearing premise

That the screenshot-plus-API pipeline and its cost figures are a realistic lower bound on what a profit-seeking online participant can actually achieve under real platform constraints.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 5 minor

Summary. The paper evaluates whether eight canonical real-effort tasks used in experimental economics remain valid measures of human effort when multimodal LLMs can complete them. Using 23 models from OpenAI, Google, and Anthropic (accessed via OpenRouter), 20 independent runs per model–task under uniform token/time limits, and exact-match scoring against oTree ground truth, the authors show that Addition, Pair Summation, Letter Decoding, and Sequence Completion are near-ceiling for most models, while Sudoku, Counting Zeros, and String Entry remain harder. Accuracy rises with model generation and mid-tier models close the gap with frontier ones. API costs leave a large positive net gain relative to a $0.25-per-correct piece-rate benchmark; an “optimal participant” who cherry-picks the best model per task reaches ~93% average accuracy. A 2×2 verbal-incentive × human-persona design (T0–T4), analyzed with a binomial GLMM with model and task fixed effects, finds no detectable effect of incentive language or persona framing. The authors conclude that, in unsupervised settings, observed performance may no longer reflect genuine human effort and offer practical recommendations (prefer perceptual tasks, use incentive responsiveness and within-session dynamics as diagnostics).

Significance. If the results hold, the paper supplies a clear, timely boundary condition for a workhorse tool in experimental economics: most of the tested real-effort tasks are already automatable at high accuracy and negligible cost, and verbal incentives—central to the logic of real-effort designs—do not move LLM accuracy. Strengths include preregistration, public code and data, a transparent cost accounting, a multi-provider multi-tier design that documents rapid mid-tier catch-up, and a properly specified null-incentive analysis. The contribution is empirical and methodological rather than theoretical, but it is directly actionable for online and unsupervised experiments and for the design of future effort tasks.

minor comments (5)
  1. Section 2.2 and footnote 5 correctly note that full browser automation is feasible and would not raise token cost, but the manuscript could more explicitly flag that measured accuracy/cost is a lower bound on what a sophisticated participant could achieve (and that prevalence and detection risk are left for future work).
  2. Figure 1 pools all 23 models; a short note that the ranking of task difficulty is stable when restricted to top-tier models would help readers who care only about frontier capability.
  3. The $0.25 piece-rate benchmark is described as conservative; a one-sentence citation or range from recent Prolific/MTurk real-effort studies would make the external benchmark fully transparent.
  4. Minor presentation: a few model names in the heatmaps (e.g., “gemini-3.1-flash-lite”) and the chronological ordering within tiers could be cross-checked against Table 5 for consistency of release dates and labels.
  5. The discussion of stationary accuracy as a diagnostic is useful; a brief caveat that temperature/sampling settings or multi-agent wrappers could reintroduce variance would avoid overclaiming uniqueness of the flat profile.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: direct empirical measurement of LLM accuracy and API cost on fixed tasks; no fitted parameters re-presented as predictions.

full rationale

The paper's central claims rest on direct measurement: 20 independent runs per model–task under exact-match scoring, token counts and OpenRouter prices used to compute API cost, and a binomial GLM with model/task fixed effects for the five prompt treatments. Accuracy is the fraction of exact matches; cost is the sum of priced input and output tokens; net gain is that cost subtracted from an external $0.25 piece-rate benchmark taken from typical online-experiment rates, not estimated from the LLM data. Verbal-incentive null results are likewise measured contrasts (T0–T4), not derived from a fitted behavioral parameter. There is no self-definitional loop, no parameter fitted on a subset and then called a prediction, no uniqueness theorem imported from the authors' prior work, and no ansatz smuggled via self-citation. The derivation chain is therefore self-contained empirical reporting; circularity score is zero.

Assumptions & free parameters 2 free parameters · 4 assumptions · 1 invented entities

This is an empirical benchmarking paper, not a formal derivation. The central claim rests on a small set of domain and experimental-design assumptions rather than free parameters or invented physical entities. The only numerical choices that affect the economic-substitution conclusion are the external piece-rate benchmark and the uniform API limits.

free parameters (2)
  • piece_rate_benchmark = $0.25 per correct answer
    Net-gain calculations use a fixed $0.25 per correct answer taken as a conservative online-experiment rate; the qualitative conclusion (positive net gain for every model) is robust to moderate changes but the exact dollar figures are not.
  • max_output_tokens_and_timeout = 2048 tokens / 120 s
    Uniform cap of 2048 output tokens and 120 s; responses exceeding either are scored incorrect. These are experimental controls, not fitted to the accuracy data, but they bound measured performance.
assumptions (4)
  • domain assumption Real-effort tasks measure costly human cognitive exertion only when a human actually performs them.
    Stated in the introduction as the construct-validity premise of the entire literature the paper critiques.
  • domain assumption Screenshot-plus-API (or equivalent browser automation) is a feasible and low-cost way for an online participant to outsource a task.
    Section 2.2 and footnote 5; the cost and accuracy numbers are interpreted as lower bounds under this premise.
  • ad hoc to paper Exact string match to ground truth is the appropriate correctness criterion (no partial credit).
    Scoring rule in Section 2.2; it is strict and may understate partial competence on String Entry and Distorted Text.
  • domain assumption API calls produce i.i.d. responses with no within-session learning or fatigue.
    Used in the Discussion as a potential diagnostic of automation; standard for current LLM APIs.
invented entities (1)
  • optimal participant (cherry-picking best model per task)
    purpose: Upper-bound the automation threat by allowing a strategic agent to select the strongest model for each task, raising average accuracy from ~82% to 93%.
    A constructed counterfactual, not a new physical or cognitive entity; independent_evidence is false because it is defined inside the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Artificial Effort." pith.science (2026). https://pith.science/paper/E2HVE3H6

@misc{pith2026260523920,
  author       = {Pith},
  title        = {Pith review of: Artificial Effort},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/E2HVE3H6}},
  note         = {Machine review of arXiv:2605.23920}
}
read the original abstract

Real-effort tasks, in which participants perform cognitively costly activities whose outcomes depend on actual performance, are widely used in experimental economics. Their validity, however, rests on the assumption that a human performs them. We study whether this assumption still holds in the era of Artificial Intelligence (AI) and Large Language Models (LLMs). Using 8 canonical real-effort tasks and 23 LLMs from three major providers, we show that most tasks can now be solved accurately and at a negligible cost, while only a few resist automation. Performance improves with each model generation, and midtier models are rapidly closing the gap with frontier ones, broadening the set of widely accessible models that can automate these tasks. Additionally, we show that verbally offering monetary incentives has no effect on LLM performance. Our findings establish a boundary condition for the use of real-effort tasks in unsupervised settings: when participants can cheaply outsource task completion to an LLM, observed performance may no longer reflect genuine human effort.

Figures

Figures reproduced from arXiv: 2605.23920 by the authors.

Figure 1
Figure 1. Average accuracy per task under the control treatment (T0), pooled across [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Average accuracy by LLM within each provider family under the control [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Accuracy (%) per LLM and task under the control treatment (T0). LLMs [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: Cost–accuracy–tokens trade-off under the control treatment (T0). Each [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Net gain per model ($0.25 per correct answer minus API cost) under the [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: Incentive effect on accuracy (percentage points) for each model–task pair [PITH_FULL_IMAGE:figures/full_fig_p018_6.png]
Figure 7
Figure 7. Figure 7: Incentive effect on accuracy (percentage points) for each model–task pair [PITH_FULL_IMAGE:figures/full_fig_p018_7.png]
Figure 8
Figure 8. Figure 8: Accuracy (%) per model and task—T1: Standard, No Incentive. [PITH_FULL_IMAGE:figures/full_fig_p020_8.png]
Figure 9
Figure 9. Figure 9: Accuracy (%) per model and task—T2: Standard, Incentive. [PITH_FULL_IMAGE:figures/full_fig_p021_9.png]
Figure 10
Figure 10. Figure 10: Accuracy (%) per model and task—T3: Human, No Incentive. [PITH_FULL_IMAGE:figures/full_fig_p021_10.png]
Figure 11
Figure 11. Figure 11: Accuracy (%) per model and task—T4: Human, Incentive. [PITH_FULL_IMAGE:figures/full_fig_p022_11.png]
Figure 12
Figure 12. Figure 12: Average API cost per task ($) C¯m by model under the control treatment (T0), computed as in the text above. Models are grouped by provider and tier; darker shades indicate newer releases within the same tier. Error bars denote standard errors. 24 [PITH_FULL_IMAGE:fig…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

36 extracted references · 7 linked inside Pith

  1. [1]

    Abeler, J., A. Falk, L. Goette, and D. Huffman (2011): Reference points and effort provision, American Economic Review, 101, 470--492

  2. [2]

    Anthropic (2023): Introducing Claude , https://www.anthropic.com/news/introducing-claude, accessed: 2026-04-13

  3. [3]

    Niederle, and C

    Augenblick, N., M. Niederle, and C. Sprenger (2015): Working over time: Dynamic inconsistency in real effort tasks, The Quarterly Journal of Economics, 130, 1067--1115

  4. [4]

    Wiegmann, E

    Bevendorff, J., M. Wiegmann, E. Richter, M. Potthast, and B. Stein (2025): The Two Paradigms of LLM Detection: Authorship Attribution vs. Authorship Verification, in Findings of the Association for Computational Linguistics: ACL 2025, ed. by W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, Vienna, Austria: Association for Computational Linguistics, 3762--3787

  5. [5]

    Bracha, A. and C. Fershtman (2013): Competitive Incentives: Working Harder or Working Smarter? Management Science, 59, 771--781

  6. [6]

    Franke, and P

    Calsamiglia, C., J. Franke, and P. Rey-Biel (2013): The incentive effects of affirmative action in a real-effort tournament, Journal of Public Economics, 98, 15--31

  7. [7]

    Carpenter, J. and E. Huet-Vaughn (2019): 19. Real-effort tasks, Handbook of research methods and applications in experimental economics, 368

  8. [8]

    Exley, S

    Celebi, C., C. Exley, S. Harrs, H. Kivimaki, M. Serra-Garcia, and J. Yusof (2026): Mission Possible: The Collection of High-Quality Data,

Show all 36 references
  1. [9]

    Gneezy, and A

    Charness, G., U. Gneezy, and A. Henderson (2018): Experimental methods: Measuring effort in economics experiments, Journal of Economic Behavior & Organization, 149, 74--87

  2. [10]

    Chen, D. L., M. Schonger, and C. Wickens (2016): oTree—An open-source platform for laboratory, online, and field experiments, Journal of Behavioral and Experimental Finance, 9, 88--97

  3. [11]

    Tworek, H

    Chen, M., J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al. (2021): Evaluating large language models trained on code, arXiv preprint arXiv:2107.03374

  4. [12]

    Cowhey, O

    Clark, P., I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord (2018): Think you have solved question answering? Try ARC, the AI2 reasoning challenge, arXiv preprint arXiv:1803.05457

  5. [13]

    Kosaraju, M

    Cobbe, K., V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al. (2021): Training verifiers to solve math word problems, arXiv preprint arXiv:2110.14168

  6. [14]

    Ding, Z., G. Deng, Y. Liu, J. Ding, J. Chen, Y. Sui, and Y. Li (2025): IllusionCAPTCHA: A CAPTCHA based on Visual Illusion,

  7. [15]

    Gangadharan, and N

    Erkal, N., L. Gangadharan, and N. Nikiforakis (2011): Relative earnings and giving in a real-effort experiment, American Economic Review, 101, 3330--3348

  8. [16]

    Ferrando, J

    Fu, T., R. Ferrando, J. Conde, C. Arriaga, and P. Reviriego (2024): Why Do Large Language Models ( LLMs ) Struggle to Count Letters? arXiv preprint arXiv:2412.18626

  9. [17]

    Burns, S

    Hendrycks, D., C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt (2021 a ): Measuring Massive Multitask Language Understanding, in International Conference on Learning Representations

  10. [18]

    Burns, S

    Hendrycks, D., C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt (2021 b ): Measuring Mathematical Problem Solving With the MATH Dataset, in Thirty-Fifth Conference on Neural Information Processing Systems

  11. [19]

    Jimenez, C. E., J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan (2024): SWE-bench: Can Language Models Resolve Real-World GitHub Issues? in The Twelfth International Conference on Learning Representations

  12. [20]

    Joshi, M., E. Choi, D. S. Weld, and L. Zettlemoyer (2017): TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension, in Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, 1601--1611

  13. [21]

    Kessler, J. B. and M. I. Norton (2016): Tax aversion in labor supply, Journal of Economic Behavior & Organization, 124, 15--28

  14. [22]

    Palomaki, O

    Kwiatkowski, T., J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, et al. (2019): Natural questions: a benchmark for question answering research, Transactions of the Association for Computational Linguistics, 7, 453--466

  15. [23]

    Bottou, Y

    LeCun, Y., L. Bottou, Y. Bengio, and P. Haffner (2002): Gradient-based learning applied to document recognition, Proceedings of the IEEE, 86, 2278--2324

  16. [24]

    Alghamdi, M

    Maiya, A., R. Alghamdi, M. L. Pacheco, A. Trivedi, and F. Somenzi (2025): Explaining Puzzle Solutions in Natural Language: An Exploratory Study on 6x6 Sudoku, in Findings of the Association for Computational Linguistics: ACL 2025, ed. by W. Che, J. Nabende, E. Shutova, and M. ...

  17. [25]

    Niederle, M. and L. Vesterlund (2007): Do women shy away from competition? Do men compete too much? The quarterly journal of economics, 122, 1067--1101

  18. [26]

    Jiang, and J

    Nogueira, R., Z. Jiang, and J. Lin (2021): Investigating the Limitations of Transformers with Simple Arithmetic Tasks, arXiv preprint arXiv:2102.13019

  19. [27]

    OpenAI (2022): Introducing ChatGPT , https://openai.com/blog/chatgpt, accessed: 2026-04-13

  20. [28]

    Pichai, S. and D. Hassabis (2023): Introducing Gemini : Our Largest and Most Capable AI Model, https://blog.google/technology/ai/google-gemini-ai/, accessed: 2026-04-13

  21. [29]

    Slonim, and M

    Rosaz, J., R. Slonim, and M. C. Villeval (2016): Quitting and peer effects at work, Labour Economics, 39, 55--67

  22. [30]

    Srivastava, A. et al. (2022): Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models, arXiv preprint arXiv:2206.04615

  23. [31]

    Team, G. et al. (2023): Gemini: a family of highly capable multimodal models, arXiv preprint arXiv:2312.11805

  24. [32]

    Wang, J., C. Zhu, Y. Zhou, L. Li, X. He, and J. Xiong (2025): COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA Solvers,

  25. [33]

    Wei, J., X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. H. Chi, Q. V. Le, and D. Zhou (2022): Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, in Advances in Neural Information Processing Systems, vol. 35, 24824--24837

  26. [34]

    Kaplan, G

    Yehudai, G., H. Kaplan, G. Dar, R. Rassin, A. Ghandeharioun, M. Geva, and A. Globerson (2024): When Can Transformers Count to n ? arXiv preprint arXiv:2407.15160

  27. [35]

    Holtzman, Y

    Zellers, R., A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi (2019): Hellaswag: Can a machine really finish your sentence? in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 3472--3482

  28. [36]

    Zhang, J., Z. Zhou, X. Ji, S. Liu, and Z. Zhao (2025): CAPTURE: A Benchmark and Evaluation for LVLMs in CAPTCHA Resolving,

Pith tools

Reviewed July 12, 2026 · model on record in the stance chip above.