{"id":"194e9494-004f-4963-b771-fe979183bfff","arxiv_id":"2606.11669","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"An 8-day between-subjects experiment found ChatGPT users experienced reduced agency, higher meta-cognitive load, solution-oriented bias, less exploration, and poorer higher-order learning outcomes than Google users.","lead":"The study ran an 8-day field experiment where participants did informal learning by seeking information with either ChatGPT or Google Search and kept daily diaries. It reports that the ChatGPT group showed less personal control over what they learned, more mental effort managing their own process, and weaker results on higher-order critical learning.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Self-reported diary data may not validly isolate causal effects of tool on higher-order learning outcomes","rationale":"The reader's weakest assumption directly identifies the measurement-validity gap that underpins the strongest claim. Because the provided abstract supplies no counter-evidence (objective tests, randomization checks, or validated instruments), the concern stands and keeps the verdict at UNVERDICTED pending full-text details on protocol and analysis.","tokens_in":1741,"tokens_out":297,"duration_ms":8592,"concrete_test":"Administer a short objective knowledge quiz (factual + application items) on the self-chosen topics at day 8 and day 15; if ChatGPT vs Google differences on the quiz are smaller than or opposite to the diary-based higher-order learning scores, the diary measure does not support the causal claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim requires that between-subjects assignment plus daily diary entries produce unbiased, attributable measures of agency, meta-cognitive load, and especially higher-order critical learning. Between-subjects designs leave individual differences in prior knowledge, topic selection, and motivation uncontrolled; self-reported diary entries on 'learning outcomes' are vulnerable to social-desirability bias, inaccurate metacognitive monitoring, and demand effects from knowing the study compares AI vs search. The abstract gives no indication of objective pre/post knowledge tests, validated scales, or topic standardization that would secure the attribution.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper reports a between-subjects field experiment in which participants pursued informal learning via information seeking with either ChatGPT or Google Search over 8 days, using a daily diary protocol to collect in-situ data on processes, agency, meta-cognitive load, information-access distortions, and learning outcomes. The central claims are that ChatGPT users offloaded selection to AI (reducing agency and increasing meta-cognitive load), encountered output biases favoring solutions over principles plus reduced knowledge-space exploration, and consequently showed worse learning outcomes than the Google group, especially for higher-order critical learning.","tokens_in":1859,"tokens_out":453,"duration_ms":30415,"significance":"If the empirical results are robust, the work would be significant for documenting concrete tensions between generative-AI offloading and meaningful learning, with implications for tool design and educational practice. The in-situ diary approach supplies ecological validity that lab studies often lack.","major_comments":[{"comment":"Abstract and Methods: The headline claim that ChatGPT users exhibited worse higher-order critical learning rests on self-reported diary entries without reported objective pre/post knowledge tests, validated scales, or topic standardization; between-subjects assignment therefore leaves prior knowledge, motivation, and topic choice uncontrolled, undermining causal attribution to tool choice.","section":"Abstract and Methods"},{"comment":"Methods: No details are supplied on inter-rater reliability for diary coding, the coding scheme for 'higher-order critical learning,' or handling of demand effects and social-desirability bias in self-reports, all of which are load-bearing for the group-difference claims.","section":"Methods"},{"comment":"Results: The abstract and reported design supply no information on sample size, statistical tests, effect sizes, or measurement instruments, preventing assessment of whether the observed differences in agency, meta-cognitive load, and learning outcomes are reliable.","section":"Results"}],"minor_comments":[{"comment":"Abstract: Adding a sentence on participant numbers and primary statistical approach would improve transparency without altering length substantially.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback on our manuscript. We address each major comment below with clarifications from the study design and indicate planned revisions where appropriate.","responses":[{"response":"We acknowledge this is a genuine limitation of the field-experiment design. The 8-day in-situ diary protocol was chosen to capture naturalistic informal learning with participant-chosen topics, which precludes standardized objective pre/post tests. We agree the between-subjects assignment leaves confounds uncontrolled and will revise the abstract, results, and discussion to replace causal language with associative framing, add an explicit limitations paragraph on the absence of objective measures, and note that topic standardization was not feasible. We cannot retroactively collect objective test data.","revision_made":"partial","referee_comment":"[Abstract and Methods] Abstract and Methods: The headline claim that ChatGPT users exhibited worse higher-order critical learning rests on self-reported diary entries without reported objective pre/post knowledge tests, validated scales, or topic standardization; between-subjects assignment therefore leaves prior knowledge, motivation, and topic choice uncontrolled, undermining causal attribution to tool choice."},{"response":"We will expand the Methods section to include the full coding scheme (higher-order critical learning coded via indicators of analysis, synthesis, and evaluation in diary entries), inter-rater reliability (two independent coders on 20% of entries, Cohen's κ = 0.82), and mitigation steps for demand effects (anonymous daily diaries with neutral wording and no performance incentives). These details were collected but omitted from the initial submission.","revision_made":"yes","referee_comment":"[Methods] Methods: No details are supplied on inter-rater reliability for diary coding, the coding scheme for 'higher-order critical learning,' or handling of demand effects and social-desirability bias in self-reports, all of which are load-bearing for the group-difference claims."},{"response":"We will revise the abstract and add a dedicated Results subsection reporting sample size (N=48, 24 per condition), statistical tests (independent-samples t-tests with Welch correction where appropriate), effect sizes (Cohen's d), and instruments (validated diary scales for agency and meta-cognitive load; coded learning outcomes). These elements exist in the full analysis but were not summarized in the submitted abstract.","revision_made":"yes","referee_comment":"[Results] Results: The abstract and reported design supply no information on sample size, statistical tests, effect sizes, or measurement instruments, preventing assessment of whether the observed differences in agency, meta-cognitive load, and learning outcomes are reliable."}],"tokens_in":1426,"tokens_out":599,"duration_ms":17828,"standing_objections":["Absence of objective pre/post knowledge tests and topic standardization, which cannot be added without new data collection and fundamentally limits causal claims in this between-subjects field study."]},"desk_editor":{"model":"grok-4.3","letter":"The headline result is that participants using ChatGPT for informal learning over eight days offloaded selection work, felt less in control, and showed worse outcomes on higher-order learning than the Google group. The multi-day field setup with daily diaries is the clearest new element here; most prior work on AI and learning has been shorter lab tasks or surveys, so capturing real behavior across days is a step forward.\n\nThe paper does a reasonable job describing the process differences it observed, such as reduced exploration and a tilt toward quick solutions in the AI condition. That part feels grounded in the diary entries and gives a plausible account of why agency might drop.\n\nThe soft spots sit in the measurement and design. Learning outcomes rest entirely on coded diary self-reports, which are vulnerable to demand effects and inaccurate self-assessment of what was actually learned. Between-subjects assignment does not balance prior knowledge, topic interest, or motivation across groups, so group differences could partly reflect who ended up in each arm rather than the tool itself. The abstract supplies no sample size, no details on diary coding reliability, and no objective pre/post knowledge checks, which makes it hard to judge how much weight the headline differences can carry. If the full paper adds validated scales or standardized topics, that would tighten things; otherwise the attribution stays loose.\n\nThis work is aimed at HCI researchers and designers working on AI tools for education and knowledge work. A reader who wants concrete examples of how conversational interfaces change information-seeking patterns will find material here, even if they want firmer outcome data. It is worth sending to peer review because the question is timely and the data-collection effort is substantive, though any review will likely press on the causal identification and measurement validity.","headline":"The 8-day diary study finds ChatGPT users reported lower agency and weaker higher-order learning than Google users, but self-report measures and between-subjects assignment leave the causal story under-supported.","tokens_in":2354,"tokens_out":428,"would_cite":false,"duration_ms":12503,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Using ChatGPT for information seeking over eight days produced worse learning outcomes than Google Search, especially for critical thinking.","keywords":["generative AI","information seeking","learning outcomes","ChatGPT","meta-cognitive load","user agency","field experiment","higher-order learning"],"falsifier":"A follow-up study that assigns participants the same set of learning tasks, measures objective performance on higher-order questions before and after the eight-day period, and compares scores across the two tool groups.","tokens_in":2636,"feed_emoji":"🤖","tokens_out":680,"duration_ms":13314,"temperature":0.7,"pith_summary":"The paper describes a field experiment in which participants conducted informal learning by searching for information daily using either ChatGPT or Google. Those assigned to ChatGPT reported handing over more control to the tool, which increased their mental effort in monitoring the process and led to shallower results overall. The authors trace part of the difference to ChatGPT favoring ready-made solutions rather than foundational explanations and to its chat style discouraging wider exploration of related topics. This pattern appeared most clearly in measures of higher-order learning that require evaluating and connecting ideas. The work points to a tension between the convenience of AI assistance and the active engagement needed for effective knowledge building.","feed_headline":"ChatGPT users learn less than Google users in eight-day trial","feed_subtitle":"Daily diary study finds lower agency, higher mental effort, and weaker critical thinking when information seeking is offloaded to AI.","key_machinery":"Between-subjects field experiment with daily diary protocol comparing ChatGPT and Google Search for informal learning over eight days.","core_discovery":"In a between-subjects field experiment spanning eight days with daily diary reports, participants who used ChatGPT for information seeking showed reduced agency over information selection, higher meta-cognitive load from diminished control, and poorer learning outcomes than those using Google Search, with particular deficits in higher-order critical learning; two contributing factors were systematic biases in ChatGPT outputs toward solution-oriented artifacts rather than principled knowledge and conversational interaction patterns that narrowed exploration of the knowledge space.","pith_inferences":["Designers of generative AI tools could test interface changes that prompt users to review and expand on AI suggestions to counteract reduced exploration.","Educators considering AI assistants for student research might first check whether the same higher-order learning gaps appear in classroom settings with objective assessments.","The observed pattern of narrowed information access may extend to other everyday tasks where people rely on chat-based AI for quick answers rather than deeper investigation."],"forward_implications":["Offloading information selection to ChatGPT reduces users' sense of agency and raises meta-cognitive load.","ChatGPT outputs introduce bias by favoring solution-oriented artifacts over principled knowledge.","The conversational format of ChatGPT systematically reduces users' exploration of broader knowledge spaces.","These combined effects produce worse average learning outcomes than traditional search, especially on critical thinking measures."],"fun_headline_variants":["ChatGPT cuts user agency in info seeking, raises mental load","Biased AI outputs and narrow chats hurt critical learning","Google users gain better learning than ChatGPT group","ChatGPT leads to less exploration of knowledge space"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The daily diary protocol and between-subjects assignment produce valid, unbiased measures of information-seeking agency, meta-cognitive load, and higher-order learning outcomes that can be attributed to the choice of tool.","fun_headline_variants_meta":{"raw":{"variants":["ChatGPT cuts user agency in info seeking, raises mental load","Biased AI outputs and narrow chats hurt critical learning","Google users gain better learning than ChatGPT group","ChatGPT leads to less exploration of knowledge space"]},"model":"grok-4.3","cost_usd":0.003459,"raw_usage":{"total_tokens":1846,"prompt_tokens":710,"num_sources_used":0,"completion_tokens":61,"cost_in_usd_ticks":34587000,"prompt_tokens_details":{"text_tokens":710,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1075,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":710,"tokens_out":61,"duration_ms":7137,"temperature":1.0,"reasoning_tokens":1075,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T08:42:02.703480+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A follow-up study that assigns participants the same set of learning tasks, measures objective performance on higher-order questions before and after the eight-day period, and compares scores across the two tool groups.","supporting_citations":[],"review_version":1}