{"id":"ae06c160-5b16-4b61-819a-d85d2a94ad73","arxiv_id":"2605.23920","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"low","formal_verification":"none","parameter_count":2,"one_line_summary":"Twenty-three LLMs solve most of eight standard real-effort tasks accurately and cheaply, with no response to verbal incentives, so unsupervised real-effort measures may no longer capture human effort.","lead":"Most canonical real-effort tasks in experimental economics can now be solved accurately by current LLMs at a tiny fraction of typical human piece rates. This creates a construct-validity problem for unsupervised online experiments: measured performance may reflect AI access rather than human effort.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The reader correctly isolates the strongest claim and the main scope limitation. The paper’s evidence (Figures 1–5, Tables 3–4, Appendix heatmaps) directly supports capability, generational improvement, cost dominance, and incentive nulls; the screenshot-plus-API design is explicitly framed as a lower bound that already yields positive net gains for every model. Because the claim is conditional (“when participants can cheaply outsource \to performance may no longer reflect human effort”) rather than a prevalence claim, the missing field measurement of actual cheating does not undermine the reported results. No stronger technical soft spot (scoring rule, token accounting, treatment design, or model selection) rises to load-bearing status. Verdict therefore remains ACCEPT; the concrete test above is only a useful robustness check, not a required fix.","tokens_in":22483,"tokens_out":544,"duration_ms":16921,"concrete_test":"Re-execute the public replication pipeline (https://github.com/belerico/llm-real-effort) for the three hardest tasks (Sudoku, Counting Zeros, String Entry) under the control prompt, using only the cheapest mid-tier models that remain freely or near-freely accessible today (e.g., Gemini Flash-Lite / GPT-5-mini class); if average accuracy stays above ~60 % and net gain remains positive at $0.25/correct, the economic-dominance half of the claim is robust even under tighter access constraints.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is a carefully scoped capability-and-economics result: most of the eight canonical real-effort tasks are already solved by current multimodal LLMs at high accuracy and at API costs that leave a large positive net gain relative to typical online piece rates ($0.25/correct), while verbal incentive language leaves accuracy unchanged (T0–T4, binomial model with model/task FEs). That evidence is measured directly (20 runs per model–task, exact-match scoring, public code/data) and is sufficient for the stated boundary condition. The reader’s noted gap—absence of measured human cheating rates or platform detection risk under real constraints—is a genuine scope limitation (Section 2.2, footnote 5) but is not required for the claim as written; the paper never asserts prevalence, only that outsourcing is already feasible and dominant when it occurs. No internal inconsistency, circular derivation, or unsupported leap appears in the accuracy, cost, or null-incentive results.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper evaluates whether eight canonical real-effort tasks used in experimental economics remain valid measures of human effort when multimodal LLMs can complete them. Using 23 models from OpenAI, Google, and Anthropic (accessed via OpenRouter), 20 independent runs per model–task under uniform token/time limits, and exact-match scoring against oTree ground truth, the authors show that Addition, Pair Summation, Letter Decoding, and Sequence Completion are near-ceiling for most models, while Sudoku, Counting Zeros, and String Entry remain harder. Accuracy rises with model generation and mid-tier models close the gap with frontier ones. API costs leave a large positive net gain relative to a $0.25-per-correct piece-rate benchmark; an “optimal participant” who cherry-picks the best model per task reaches ~93% average accuracy. A 2×2 verbal-incentive × human-persona design (T0–T4), analyzed with a binomial GLMM with model and task fixed effects, finds no detectable effect of incentive language or persona framing. The authors conclude that, in unsupervised settings, observed performance may no longer reflect genuine human effort and offer practical recommendations (prefer perceptual tasks, use incentive responsiveness and within-session dynamics as diagnostics).","tokens_in":22751,"tokens_out":793,"duration_ms":7512,"significance":"If the results hold, the paper supplies a clear, timely boundary condition for a workhorse tool in experimental economics: most of the tested real-effort tasks are already automatable at high accuracy and negligible cost, and verbal incentives—central to the logic of real-effort designs—do not move LLM accuracy. Strengths include preregistration, public code and data, a transparent cost accounting, a multi-provider multi-tier design that documents rapid mid-tier catch-up, and a properly specified null-incentive analysis. The contribution is empirical and methodological rather than theoretical, but it is directly actionable for online and unsupervised experiments and for the design of future effort tasks.","major_comments":[],"minor_comments":[{"comment":"Section 2.2 and footnote 5 correctly note that full browser automation is feasible and would not raise token cost, but the manuscript could more explicitly flag that measured accuracy/cost is a lower bound on what a sophisticated participant could achieve (and that prevalence and detection risk are left for future work).","section":null},{"comment":"Figure 1 pools all 23 models; a short note that the ranking of task difficulty is stable when restricted to top-tier models would help readers who care only about frontier capability.","section":null},{"comment":"The $0.25 piece-rate benchmark is described as conservative; a one-sentence citation or range from recent Prolific/MTurk real-effort studies would make the external benchmark fully transparent.","section":null},{"comment":"Minor presentation: a few model names in the heatmaps (e.g., “gemini-3.1-flash-lite”) and the chronological ordering within tiers could be cross-checked against Table 5 for consistency of release dates and labels.","section":null},{"comment":"The discussion of stationary accuracy as a diagnostic is useful; a brief caveat that temperature/sampling settings or multi-agent wrappers could reintroduce variance would avoid overclaiming uniqueness of the flat profile.","section":null}],"recommendation":"accept","confidential_remarks":"The manuscript is a clean, well-executed empirical contribution that fits a methods/experimental-economics or computational-social-science venue. The reader’s and skeptic’s assessments align with mine: the absence of measured human cheating rates is a scope limitation, not a load-bearing flaw for the stated claim. I see no reason to delay acceptance for further experiments."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a straightforward capability-and-cost paper that experimental economists running unsupervised online studies should actually read. The new result is concrete: eight oTree-style real-effort tasks, 23 multimodal models across OpenAI/Google/Anthropic, 20 exact-match runs each, plus a transparent API-cost comparison against a $0.25 piece rate and a 2x2 incentive/persona prompt test. Most tasks (addition, pair summation, letter decoding, sequence completion) are already near-ceiling; Sudoku, counting zeros, and string entry still resist single models but fall under cherry-picking. Mid-tier models are closing the gap fast, every model shows positive net gain, and verbal incentives do nothing (binomial GLMM with model/task FEs, all contrasts null).\n\nWhat it does well is the design hygiene: preregistration, public code/data, uniform token/time limits, heatmaps by model-task, and cost accounting that is easy to audit. The null incentive result is useful precisely because it is null—it shows the usual experimental language is inert for LLMs and can therefore serve as a diagnostic. The generational improvement pattern is also cleanly documented.\n\nSoft spots are real but scoped. The pipeline is screenshot-plus-API rather than full browser automation; the authors note the latter is feasible and would only lower cost further, so their numbers are a conservative lower bound on the threat, not an overstatement. They do not measure observed cheating rates or platform detection risk; that is a genuine scope limit, not a hole in the capability claim they actually make. The $0.25 piece-rate benchmark is external and conventional, not estimated from the LLM data. No circularity, no load-bearing free parameters that drive the conclusions.\n\nThis is for methodologists and anyone designing online real-effort experiments. It does not rewrite theory; it forces a practical redesign of unsupervised settings. I would send it to peer review without hesitation. Cite it if you run or review online effort tasks; bring it to reading group if the group cares about experimental validity under AI.","headline":"Clean empirical boundary condition: most canonical real-effort tasks are already automatable by mid-tier multimodal LLMs at negligible cost, with null incentive effects.","tokens_in":23306,"tokens_out":521,"would_cite":true,"duration_ms":6285,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Most real-effort tasks used in economics experiments can already be solved by LLMs accurately and cheaply, so unsupervised scores may no longer measure human effort.","keywords":["Large Language Models","Real-effort tasks","Experimental economics","Automation","Online experiments","Incentives","Construct validity"],"falsifier":"Measure actual substitution rates and detection rates when real online participants are free to use commercial LLMs or browser agents on the same eight tasks, and check whether the high accuracy and positive net-gain numbers survive under those conditions.","tokens_in":23409,"feed_emoji":"🤖","tokens_out":528,"duration_ms":4753,"temperature":0.7,"pith_summary":"Real-effort tasks are meant to measure costly human performance in experimental economics. This paper tests whether that assumption still holds once large language models are available. It runs eight standard tasks against 23 multimodal models from three major providers and finds that four of the tasks are already solved near-perfectly at negligible cost, while only a few remain hard. Newer and mid-tier models keep closing the gap with frontier ones, and verbally offered money leaves model accuracy unchanged. The practical upshot is a boundary condition: in unsupervised online settings, a participant who can outsource the task to an LLM can produce high scores that no longer reflect genuine human effort.","feed_headline":"LLMs already solve most real-effort tasks used in economics","feed_subtitle":"At typical online piece rates, outsourcing beats human effort; money talk does not change model accuracy.","key_machinery":"A standardized screenshot-to-API pipeline that feeds each oTree task instance (instructions plus rendered image) to 23 multimodal LLMs under fixed token and time limits, scoring exact-match accuracy over 20 runs per model–task pair and converting token use into dollar cost.","core_discovery":"Most of the eight canonical real-effort tasks can be solved accurately by current LLMs at a cost far below typical online piece rates, while verbal monetary incentives and human-persona prompts leave LLM accuracy unchanged; therefore, when participants can cheaply outsource task completion, observed performance may no longer measure genuine human effort.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["LLMs solve most real-effort tasks used in econ experiments","Cheap LLMs automate 8 canonical real-effort tasks","Most real-effort tasks fall to current LLMs at near-zero cost","Verbal pay incentives leave LLM task accuracy unchanged","Outsourcing to LLMs can erase genuine human effort measures"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That the screenshot-plus-API pipeline and its cost figures are a realistic lower bound on what a profit-seeking online participant can actually achieve under real platform constraints.","fun_headline_variants_meta":{"raw":{"variants":["LLMs solve most real-effort tasks used in econ experiments","Cheap LLMs automate 8 canonical real-effort tasks","Most real-effort tasks fall to current LLMs at near-zero cost","Verbal pay incentives leave LLM task accuracy unchanged","Outsourcing to LLMs can erase genuine human effort measures"]},"model":"grok-4.5","effort":"low","cost_usd":0.004564,"raw_usage":{"total_tokens":1288,"prompt_tokens":695,"num_sources_used":0,"completion_tokens":66,"cost_in_usd_ticks":45640000,"prompt_tokens_details":{"text_tokens":695,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":527,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":695,"tokens_out":66,"duration_ms":5956,"temperature":1.0,"reasoning_tokens":527,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T19:33:14.174334+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Measure actual substitution rates and detection rates when real online participants are free to use commercial LLMs or browser agents on the same eight tasks, and check whether the high accuracy and positive net-gain numbers survive under those conditions.","supporting_citations":[],"review_version":2}