{"id":"b0c7ee58-187e-437a-bbfe-a76319405fc8","arxiv_id":"2507.00181","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a 40-person randomized experiment, students allowed to use ChatGPT reported significantly lower cognitive engagement during an academic writing task than students who wrote without AI.","lead":"A small randomized experiment found that students who used ChatGPT while writing an argumentative essay reported lower mental effort, attention, and deep thinking on a new four-item self-report scale than students who wrote without AI. The result adds to evidence that AI assistance may encourage cognitive offloading, though the study's measurement and sample size limit how far the conclusion can be generalized.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The CES-AI is an unvalidated, transparent self-report scale; without a behavioral or social-desirability check, the reported group difference may reflect demand characteristics rather than actual cognitive engagement.","rationale":"The reader's weakest-assumption analysis identifies the same load-bearing issue: the CES-AI is an unvalidated, manipulation-transparent self-report scale, so the between-group difference may reflect demand characteristics rather than actual cognitive engagement. I agree with that assessment and do not see a more fundamental flaw. The statistical result is internally consistent, the design is randomized, and the authors acknowledge the self-report limitation. The concern is serious but addressable through validation and objective measurement, which is exactly why the reader's CONDITIONAL verdict is appropriate. I would not move the verdict to REJECT because the paper is transparent about its limitations and the reported effect is plausible; I would not move it to ACCEPT because the central construct has no external validity evidence. The concrete test I propose would settle whether the concern lands by comparing self-reported engagement against objective behavioral traces in the same experimental setup.","tokens_in":6027,"tokens_out":3757,"duration_ms":50970,"concrete_test":"Re-run the experiment with a pre-registered protocol that adds objective behavioral engagement indicators to the existing design: log keystrokes and edit actions, time spent writing before the first ChatGPT query, number of self-generated arguments before consulting the tool, and post-task open-ended justifications, plus a validated social-desirability scale. If the ChatGPT group shows equivalent or higher objective engagement while CES-AI scores remain lower, the self-report effect is an artifact of the instrument; if objective engagement is also lower, the concern is mitigated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim depends entirely on the CES-AI, a four-item self-report measure built for this study and supported only by internal consistency (alpha = .88). The items are transparently close to the manipulation: participants in the ChatGPT condition were told they could use the tool for ideas, phrasing, or argument development, then asked whether they \"put effort into thinking through the problem myself\" (Item 2) and \"explored different ways to solve the problem or approach the task\" (Item 4). Because the hypothesis is obvious from the items and the procedure, the ChatGPT group may simply report lower effort to match expected behavior, while the control group, monitored via TeamViewer with cameras on, has additional social-desirability pressure to claim high effort. The study has no pre-test, no behavioral measure of engagement, no social-desirability check, and no external validation of the CES-AI; the authors' own limitations paragraph concedes that self-report \"may not fully capture the actual depth or quality of cognitive engagement.\" The reported F(1,38) = 19.2 is arithmetically consistent with the group means and standard deviations, so the internal statistics are not the problem. The load-bearing premise is that the CES-AI measures cognitive engagement rather than perceived demand or self-presentation, and that premise is currently unsecured.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a randomized experiment (N=40) comparing cognitive engagement during an argumentative writing task between a ChatGPT-assisted condition and a no-assistance control condition. Engagement was measured with a newly developed four-item self-report scale, the CES-AI, after the task. A one-way ANOVA found significantly lower CES-AI scores in the ChatGPT group (M=2.95, SD=1.18) than in the control group (M=4.19, SD=0.45), F(1,38)=19.2, p<0.001. The authors interpret this as evidence of cognitive offloading and reduced deep thinking when students use AI tools, and they draw pedagogical implications.","tokens_in":6278,"tokens_out":2361,"duration_ms":30243,"significance":"If the result is valid, the finding is relevant to the growing literature on AI in education and to debates about cognitive offloading. The study has strengths: random assignment, a clearly described experimental protocol with monitoring to prevent contamination, and statistical results that are arithmetically consistent with the reported means and standard deviations. However, the central inference depends entirely on an unvalidated, transparent self-report scale whose items closely overlap with the experimental manipulation. The paper provides no behavioral measure, no external validation of the CES-AI, no effect size or confidence interval, and no data or code for verification. These issues substantially temper the strength of the conclusion, but the core research question is important and the design is a reasonable starting point for a more thorough investigation.","major_comments":[{"comment":"The load-bearing premise of the study is that the CES-AI validly measures cognitive engagement, but this is not established. Table 1 shows that items 2 ('I put effort into thinking through the problem myself') and 4 ('I explored different ways to solve the problem or approach the task') are so close to the experimental manipulation (using ChatGPT for ideas, phrasing, or argument development) that the group difference may reflect demand characteristics or self-presentation rather than actual engagement. Internal consistency alone (Cronbach's alpha = 0.88) does not demonstrate construct validity. The paper should provide convergent/divergent validity evidence, a social-desirability check, a manipulation check, or triangulation with a non-self-report outcome; the authors themselves acknowledge this gap in the Conclusions, but it is not merely a limitation if the central claim rests on it.","section":"Instrument (Table 1)"},{"comment":"The paper reports only F and p, omitting effect size and confidence intervals. Given the small sample and the large difference in variances between conditions (SD = 0.45 control vs. 1.18 experimental, a ratio of 2.6), the standard ANOVA assumption of homogeneity of variance is questionable; the authors should report Levene's test or use a Welch correction. Reporting a standardized effect size (e.g., partial eta-squared, which can be computed as approximately 0.336) and a 95% confidence interval for the mean difference would help readers judge the practical significance and precision of the result.","section":"Results"},{"comment":"The mechanism of 'cognitive offloading' is not directly supported by the data. The Procedure states that participants in the ChatGPT condition 'were allowed to use ChatGPT' and 'could consult the AI tool,' but the paper reports no information about whether or how extensively participants actually used the tool, what they used it for, or how the writing products differed. Without any behavioral trace of ChatGPT use, the lower CES-AI scores could be driven by many factors other than offloading (e.g., perceived legitimacy of using the tool, task interpretation, or the wording of the instruction to 'engage actively'). The Discussion should temper the causal language or include an analysis that links actual tool use to the outcome.","section":"Procedure and Discussion"}],"minor_comments":[{"comment":"In the sentence 'This finding indicated that the controlled group exhibited significantly higher cognitive engagement scores compared to the experimental group,' the word 'controlled' should be 'control'.","section":"Results"},{"comment":"Figure 1 would be more informative with individual data points or boxplots and error bars, especially given the small sample and the high variance in the experimental group.","section":"Figure 1"},{"comment":"No sample-size justification or power analysis is reported; given N=40, the study may be underpowered to detect small or moderate effects, and this should be acknowledged.","section":"Methodology / Participants"},{"comment":"The in-text citation 'Lin et al., 2023' does not match the reference list entry 'Lin, T. J. (2023)' which appears to be a single-author work; please correct the citation style.","section":"Introduction and References"},{"comment":"The title's phrase 'lazy thinkers' is an interpretive gloss that goes beyond the self-report measure; the title and abstract would be more accurate if they referred to 'lower self-reported cognitive engagement' rather than making an essentialist claim about thinking dispositions.","section":"Title and Abstract"}],"recommendation":"major_revision","confidential_remarks":"The manuscript addresses a timely topic, but the present form has a central measurement-validity issue that must be resolved before publication. The authors may need to provide additional validation data, add non-self-report outcome measures, or substantially soften the causal claims. I would also encourage the editor to consider whether the journal's standards require data and code availability; the absence of both makes it difficult to verify the reported ANOVA results beyond the arithmetic check."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a clean, honest little experiment whose load-bearing claim rests on a scale that is too transparent to carry the weight. I agree with the reader's conditional verdict.\n\nWhat's actually new: a randomized comparison (n=40) with a purpose-built 4-item self-report scale, CES-AI, showing ChatGPT-assisted writers report lower cognitive engagement (M=2.95 vs 4.19, F(1,38)=19.2, p<0.001). That's a concrete, if small, addition to the existing EEG and systematic-review evidence (Kosmyna et al., Lo et al.), which the paper cites correctly. The design is decent: random assignment, equivalent prior AI use, both groups monitored via TeamViewer to prevent cheating. The ANOVA is appropriate and the arithmetic checks out. The limitations paragraph is unusually candid for a paper with a sensational title.\n\nThe soft spot is exactly where the stress-test puts it: the CES-AI. Items like 'I put effort into thinking through the problem myself' and 'I explored different ways to solve the problem' are almost the manipulation check. Any participant who just used ChatGPT knows they didn't do those things, and any control participant under camera knows they should say they did. The scale has no validation beyond alpha=.88, which only tells you the items hang together, not that they measure engagement rather than perceived demand or self-presentation. No pre-test, no behavioral measure, no social-desirability check, no data or code. The authors concede the self-report limitation themselves, but that concession doesn't rescue the central inference.\n\nMinor: the paper overclaims in the title and discussion with 'lazy thinkers', which is a value judgment not supported by the measure. Some effect size and confidence intervals would help. The sample is homogeneous (Greek-speaking, linguistics background, avg age 35), which limits generalizability.\n\nBottom line: the paper is a sincere, transparent pilot study, not a scam. But right now it documents that people who used ChatGPT report lower engagement on a scale that practically asks them to say so. That's worth knowing, but it's not yet evidence of a cognitive decline. I'd send it to peer review rather than desk-reject — referees could demand validation, the data, and a more measured title, and the result might then be a useful contribution to the AI-offloading literature. I would not cite it in its current form.","headline":"A plausible but unsecured result: the effect may be real, but the unvalidated CES-AI and the monitored control condition leave demand characteristics as a live alternative explanation.","tokens_in":6769,"tokens_out":2253,"would_cite":false,"duration_ms":25579,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that students allowed to use ChatGPT during an argumentative writing task reported substantially lower mental effort, attention, deep processing, and strategic thinking than students writing without it, and interprets…","keywords":["cognitive engagement","ChatGPT","large language models","cognitive offloading","CES-AI","argumentative writing","self-report scale","AI in education"],"falsifier":"Run the same 40-person design but add an unannounced free-recall test of the essay's arguments and a count of planning notes people produce before writing. If the ChatGPT group recalls as much and plans as much while still rating their engagement lower, the CES-AI is picking up demand characteristics rather than actual engagement; if they recall less and plan less, the offloading reading is supported.","tokens_in":5840,"feed_emoji":"🤖","tokens_out":10838,"duration_ms":108661,"temperature":0.7,"pith_summary":"The paper reports a controlled experiment asking whether letting students use ChatGPT during an academic writing task changes how deeply they engage with the task. Forty students were randomly assigned to write a 300-word argumentative essay either with ChatGPT 3.5 available or without any external help, then rated their engagement on a four-item scale built for the study. The ChatGPT group averaged 2.95, well below the control group's 4.19, a difference the statistical test reports as F(1,38)=19.2, p<0.001. The paper's conclusion is that AI assistance produces cognitive offloading: students lean on the model for ideas and argument development, report less effort, attention, deep processing, and strategic flexibility, and in that sense become lazier thinkers. If this is right, it matters because it challenges the optimistic view of generative AI as a cognitive scaffold and urges educators to design tasks that require reflective engagement with AI output.","feed_headline":"ChatGPT use tied to lower cognitive engagement in essay task","feed_subtitle":"Randomized trial: ChatGPT-assisted writers scored 2.95 vs 4.19 on a four-item engagement scale.","key_machinery":"The load-bearing object is the CES-AI, a four-item Likert scale (1–5) constructed for this study to measure cognitive engagement as mental effort, sustained attention, deep processing, and strategic thinking; the paper reports a Cronbach's alpha of 0.88 for the scale. The contrast it works with is the experimental manipulation: one group writes a timed 300-word argument about integrating AI into academic practice with ChatGPT 3.5 available, while the other group writes the same prompt with no external help. Because the scale is the only outcome measure, the argument depends on these four self-report items tracking genuine differences in cognitive engagement rather than participants' guesses about what the study expects.","core_discovery":"The central discovery, on the paper's terms, is a clean group difference on the CES-AI, a four-item self-report measure of cognitive engagement. Participants who could consult ChatGPT during the writing task reported lower scores on every facet the scale samples—deep understanding, effortful thinking, sustained attention, and strategy exploration—with a mean of 2.95 (SD=1.18) versus 4.19 (SD=0.45) in the control group; the one-way ANOVA gave F(1,38)=19.2, p<0.001. The paper reads this as evidence that relying on a generative AI for argumentative writing produces cognitive offloading rather than scaffolding: the tool supplies mental steps students would otherwise take, so they invest less. It presents the result as an extension of prior work showing neural engagement drops with LLM use and as a caution for educators integrating chatbots into writing instruction.","pith_inferences":["A per-item analysis would be a cheap test of the demand-characteristics reading: if the gap is concentrated in the items that directly describe doing the work oneself, the scale may be measuring compliance with the experimental setup rather than felt engagement.","Repeating the study with a behavioral dependent variable, such as revision count, planning notes, eye tracking, or EEG, would show whether lower self-reported engagement corresponds to measurably shallower processing; without such convergence, the offloading interpretation remains one plausible account among others.","A variant that varies the quality of the AI assistance, from generic prompts to polished model text, could separate whether the act of consulting AI or the usefulness of its output drives the drop in reported engagement."],"forward_implications":["Under the paper's account, using ChatGPT during drafting reduces the mental work of planning and evaluating arguments, which is exactly the work educators want students to practice.","Writing assignments that permit open access to ChatGPT should include built-in critical evaluation of AI output, since the paper argues that unaided reflection is not what happens.","Cognitive-engagement theories may need to treat AI tools as effort-replacing resources in some contexts rather than only as scaffolds that extend thinking.","The result justifies building AI-specific engagement measures, because the paper finds that generic instruments were unavailable before the CES-AI."],"supporting_citations":[{"why":"Provides the EEG-based evidence that LLM use reduces neural engagement and memory, which the present study extends to a self-report measure.","marker":"Kosmyna et al. (2025)"},{"why":"Systematic review of ChatGPT's influence on engagement whose weak, mixed cognitive-engagement evidence motivates the new scale.","marker":"Lo et al. (2024)"},{"why":"Supplies the tripartite engagement framework and defines cognitive engagement as mental investment, grounding item 1.","marker":"Fredricks et al. (2004)"},{"why":"Defends self-report measurement of cognitive engagement, used to justify the CES-AI method.","marker":"Greene (2015)"},{"why":"Documents the absence of a standard cognitive-engagement scale, justifying construction of the CES-AI.","marker":"Li (2021)"},{"why":"Sets the criterion for acceptable Cronbach's alpha used to claim the CES-AI is internally consistent.","marker":"Taber (2018)"},{"why":"Identifies attention regulation as part of engagement, grounding item 3 on sustained focus.","marker":"Skinner et al. (2009)"},{"why":"Provides the strategic-thinking and metacognition framework underlying item 4 on exploring multiple strategies.","marker":"Pintrich (2004)"}],"fun_headline_variants":["ChatGPT use in essays linked to lower cognitive effort","AI writing tools may promote lazy thinking: study","Randomized trial: ChatGPT cuts cognitive engagement","ChatGPT writers score lower on deep thinking scale","Study: ChatGPT users offload thinking during writing"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole result hangs on the four CES-AI items being a valid measure of cognitive engagement, even though two items—'I put effort into thinking through the problem myself' and 'I explored different ways to solve the problem or approach the task'—come very close to restating the experimental difference, so the group contrast may partly reflect what participants think the researcher wants to hear.","fun_headline_variants_meta":{"raw":{"variants":["ChatGPT use in essays linked to lower cognitive effort","AI writing tools may promote lazy thinking: study","Randomized trial: ChatGPT cuts cognitive engagement","ChatGPT writers score lower on deep thinking scale","Study: ChatGPT users offload thinking during writing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000189,"raw_usage":{"total_tokens":1318,"prompt_tokens":912,"completion_tokens":406,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":528,"completion_tokens_details":{"reasoning_tokens":335}},"tokens_in":528,"tokens_out":406,"duration_ms":4978,"temperature":1.0,"reasoning_tokens":335,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:21:25.533457+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 40-person design but add an unannounced free-recall test of the essay's arguments and a count of planning notes people produce before writing. If the ChatGPT group recalls as much and plans as much while still rating their engagement lower, the CES-AI is picking up demand characteristics rather than actual engagement; if they recall less and plan less, the offloading reading is supported.","supporting_citations":[],"review_version":1}