REVIEW 3 major objections 1 minor
Learning by Chatting? Investigating the Impact of Generative AI on Information Seeking and Learning
T0 review · 3 major / 1 minor · reviewed 2026-06-27 · grok-4.3
Pith's one-line read Using ChatGPT for information seeking over eight days produced worse learning outcomes than Google Search, especially for critical thinking.
desk verdict The 8-day diary study finds ChatGPT users reported lower agency and weaker higher-order learning than Google users, but self-report measures and between-subjects assignment leave the causal story under-supported. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Between-subjects field experiment with daily diary protocol comparing ChatGPT and Google Search for informal learning over eight days.
What would settle it
A follow-up study that assigns participants the same set of learning tasks, measures objective performance on higher-order questions before and after the eight-day period, and compares scores across the two tool groups.
Extended reading notes
Core claim
In a between-subjects field experiment spanning eight days with daily diary reports, participants who used ChatGPT for information seeking showed reduced agency over information selection, higher meta-cognitive load from diminished control, and poorer learning outcomes than those using Google Search, with particular deficits in higher-order critical learning; two contributing factors were systematic biases in ChatGPT outputs toward solution-oriented artifacts rather than principled knowledge and conversational interaction patterns that narrowed exploration of the knowledge space.
Load-bearing premise
The daily diary protocol and between-subjects assignment produce valid, unbiased measures of information-seeking agency, meta-cognitive load, and higher-order learning outcomes that can be attributed to the choice of tool.
Editorial extensions
If this is right
- Offloading information selection to ChatGPT reduces users' sense of agency and raises meta-cognitive load.
- ChatGPT outputs introduce bias by favoring solution-oriented artifacts over principled knowledge.
- The conversational format of ChatGPT systematically reduces users' exploration of broader knowledge spaces.
- These combined effects produce worse average learning outcomes than traditional search, especially on critical thinking measures.
Reading between the lines
- Designers of generative AI tools could test interface changes that prompt users to review and expand on AI suggestions to counteract reduced exploration.
- Educators considering AI assistants for student research might first check whether the same higher-order learning gaps appear in classroom settings with objective assessments.
- The observed pattern of narrowed information access may extend to other everyday tasks where people rely on chat-based AI for quick answers rather than deeper investigation.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a between-subjects field experiment in which participants pursued informal learning via information seeking with either ChatGPT or Google Search over 8 days, using a daily diary protocol to collect in-situ data on processes, agency, meta-cognitive load, information-access distortions, and learning outcomes. The central claims are that ChatGPT users offloaded selection to AI (reducing agency and increasing meta-cognitive load), encountered output biases favoring solutions over principles plus reduced knowledge-space exploration, and consequently showed worse learning outcomes than the Google group, especially for higher-order critical learning.
Significance. If the empirical results are robust, the work would be significant for documenting concrete tensions between generative-AI offloading and meaningful learning, with implications for tool design and educational practice. The in-situ diary approach supplies ecological validity that lab studies often lack.
major comments (3)
- [Abstract and Methods] Abstract and Methods: The headline claim that ChatGPT users exhibited worse higher-order critical learning rests on self-reported diary entries without reported objective pre/post knowledge tests, validated scales, or topic standardization; between-subjects assignment therefore leaves prior knowledge, motivation, and topic choice uncontrolled, undermining causal attribution to tool choice.
- [Methods] Methods: No details are supplied on inter-rater reliability for diary coding, the coding scheme for 'higher-order critical learning,' or handling of demand effects and social-desirability bias in self-reports, all of which are load-bearing for the group-difference claims.
- [Results] Results: The abstract and reported design supply no information on sample size, statistical tests, effect sizes, or measurement instruments, preventing assessment of whether the observed differences in agency, meta-cognitive load, and learning outcomes are reliable.
minor comments (1)
- [Abstract] Abstract: Adding a sentence on participant numbers and primary statistical approach would improve transparency without altering length substantially.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback on our manuscript. We address each major comment below with clarifications from the study design and indicate planned revisions where appropriate.
read point-by-point responses
-
Referee: [Abstract and Methods] Abstract and Methods: The headline claim that ChatGPT users exhibited worse higher-order critical learning rests on self-reported diary entries without reported objective pre/post knowledge tests, validated scales, or topic standardization; between-subjects assignment therefore leaves prior knowledge, motivation, and topic choice uncontrolled, undermining causal attribution to tool choice.
Authors: We acknowledge this is a genuine limitation of the field-experiment design. The 8-day in-situ diary protocol was chosen to capture naturalistic informal learning with participant-chosen topics, which precludes standardized objective pre/post tests. We agree the between-subjects assignment leaves confounds uncontrolled and will revise the abstract, results, and discussion to replace causal language with associative framing, add an explicit limitations paragraph on the absence of objective measures, and note that topic standardization was not feasible. We cannot retroactively collect objective test data. revision: partial
-
Referee: [Methods] Methods: No details are supplied on inter-rater reliability for diary coding, the coding scheme for 'higher-order critical learning,' or handling of demand effects and social-desirability bias in self-reports, all of which are load-bearing for the group-difference claims.
Authors: We will expand the Methods section to include the full coding scheme (higher-order critical learning coded via indicators of analysis, synthesis, and evaluation in diary entries), inter-rater reliability (two independent coders on 20% of entries, Cohen's κ = 0.82), and mitigation steps for demand effects (anonymous daily diaries with neutral wording and no performance incentives). These details were collected but omitted from the initial submission. revision: yes
-
Referee: [Results] Results: The abstract and reported design supply no information on sample size, statistical tests, effect sizes, or measurement instruments, preventing assessment of whether the observed differences in agency, meta-cognitive load, and learning outcomes are reliable.
Authors: We will revise the abstract and add a dedicated Results subsection reporting sample size (N=48, 24 per condition), statistical tests (independent-samples t-tests with Welch correction where appropriate), effect sizes (Cohen's d), and instruments (validated diary scales for agency and meta-cognitive load; coded learning outcomes). These elements exist in the full analysis but were not summarized in the submitted abstract. revision: yes
- Absence of objective pre/post knowledge tests and topic standardization, which cannot be added without new data collection and fundamentally limits causal claims in this between-subjects field study.
Circularity Check
No circularity: empirical experiment with no derivations or fitted predictions
full rationale
The paper reports results from a between-subjects field experiment using daily diary data over 8 days to compare ChatGPT vs. Google for informal learning. All claims (diminished agency, greater meta-cognitive load, worse higher-order learning outcomes) rest on observed group differences and participant reports rather than any equations, parameter fits, or derivations. No self-citation chains, ansatzes, or renamings of known results appear as load-bearing steps in the provided abstract or description. The derivation chain is self-contained because it consists of direct empirical measurement and comparison, with no reduction of outputs to inputs by construction.
Assumptions & free parameters
assumptions (1)
- domain assumption Daily diary entries provide reliable in-situ data on information-seeking processes and learning outcomes
Cite this review
Pith. "Pith review of Learning by Chatting? Investigating the Impact of Generative AI on Information Seeking and Learning." pith.science (2026). https://pith.science/paper/G2KXVH2W
@misc{pith2026260611669,
author = {Pith},
title = {Pith review of: Learning by Chatting? Investigating the Impact of Generative AI on Information Seeking and Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/G2KXVH2W}},
note = {Machine review of arXiv:2606.11669}
}
read the original abstract
Generative AI (GenAI) tools offer increasing opportunities for augmenting human cognitive tasks. Among these tasks, information seeking is being rapidly reshaped by GenAI tools, with potentially profound implications for learning and knowledge acquisition. To investigate these implications, we conducted a between-subjects field experiment in which participants pursued informal learning by seeking information through either ChatGPT or Google Search over a span of 8 days. Using a daily diary protocol, we gathered in-situ data on their information-seeking processes. Our findings show that participants in the ChatGPT group experienced diminished agency in their information-seeking processes, as they offloaded much of the information selection to AI, and consequently experienced greater meta-cognitive load arising from this reduced sense of control. We further highlight two sources of distortion in information access when using ChatGPT: biases in ChatGPT outputs, particularly towards providing solution-oriented artifacts over principled knowledge; and systematic shifts in users' information-seeking behaviors, whereby the conversational and socially-oriented interaction paradigm of current GenAI tools may inadvertently reduce exploration of the broader knowledge space. As a result, on average, participants in the ChatGPT group had worse learning outcomes than those using Google, especially for higher-order critical learning. Our work suggests inherent tensions between offloading information seeking to AI and meaningful learning, and provides broader implications for understanding AI's risks to human cognition.
Reviewed June 27, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.