Pith. sign in

REVIEW 3 major objections 1 minor

Learning by Chatting? Investigating the Impact of Generative AI on Information Seeking and Learning

T0 review · 3 major / 1 minor · reviewed 2026-06-27 · grok-4.3

Pith's one-line read Using ChatGPT for information seeking over eight days produced worse learning outcomes than Google Search, especially for critical thinking.

desk verdict The 8-day diary study finds ChatGPT users reported lower agency and weaker higher-order learning than Google users, but self-report measures and between-subjects assignment leave the causal story under-supported. read the letter →

arxiv 2606.11669 v2 pith:G2KXVH2W submitted 2026-06-10 cs.HC cs.CY

classification cs.HCcs.CY
keywords generativeAIinformationseekinglearningoutcomesChatGPTmeta-cognitiveloaduseragencyfieldexperimenthigher-order
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper describes a field experiment in which participants conducted informal learning by searching for information daily using either ChatGPT or Google. Those assigned to ChatGPT reported handing over more control to the tool, which increased their mental effort in monitoring the process and led to shallower results overall. The authors trace part of the difference to ChatGPT favoring ready-made solutions rather than foundational explanations and to its chat style discouraging wider exploration of related topics. This pattern appeared most clearly in measures of higher-order learning that require evaluating and connecting ideas. The work points to a tension between the convenience of AI assistance and the active engagement needed for effective knowledge building.

What carries the argument

Between-subjects field experiment with daily diary protocol comparing ChatGPT and Google Search for informal learning over eight days.

What would settle it

A follow-up study that assigns participants the same set of learning tasks, measures objective performance on higher-order questions before and after the eight-day period, and compares scores across the two tool groups.

Watch

Extended reading notes

Core claim

In a between-subjects field experiment spanning eight days with daily diary reports, participants who used ChatGPT for information seeking showed reduced agency over information selection, higher meta-cognitive load from diminished control, and poorer learning outcomes than those using Google Search, with particular deficits in higher-order critical learning; two contributing factors were systematic biases in ChatGPT outputs toward solution-oriented artifacts rather than principled knowledge and conversational interaction patterns that narrowed exploration of the knowledge space.

Load-bearing premise

The daily diary protocol and between-subjects assignment produce valid, unbiased measures of information-seeking agency, meta-cognitive load, and higher-order learning outcomes that can be attributed to the choice of tool.

Editorial extensions

If this is right

  • Offloading information selection to ChatGPT reduces users' sense of agency and raises meta-cognitive load.
  • ChatGPT outputs introduce bias by favoring solution-oriented artifacts over principled knowledge.
  • The conversational format of ChatGPT systematically reduces users' exploration of broader knowledge spaces.
  • These combined effects produce worse average learning outcomes than traditional search, especially on critical thinking measures.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Designers of generative AI tools could test interface changes that prompt users to review and expand on AI suggestions to counteract reduced exploration.
  • Educators considering AI assistants for student research might first check whether the same higher-order learning gaps appear in classroom settings with objective assessments.
  • The observed pattern of narrowed information access may extend to other everyday tasks where people rely on chat-based AI for quick answers rather than deeper investigation.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 1 minor

Summary. The paper reports a between-subjects field experiment in which participants pursued informal learning via information seeking with either ChatGPT or Google Search over 8 days, using a daily diary protocol to collect in-situ data on processes, agency, meta-cognitive load, information-access distortions, and learning outcomes. The central claims are that ChatGPT users offloaded selection to AI (reducing agency and increasing meta-cognitive load), encountered output biases favoring solutions over principles plus reduced knowledge-space exploration, and consequently showed worse learning outcomes than the Google group, especially for higher-order critical learning.

Significance. If the empirical results are robust, the work would be significant for documenting concrete tensions between generative-AI offloading and meaningful learning, with implications for tool design and educational practice. The in-situ diary approach supplies ecological validity that lab studies often lack.

major comments (3)
  1. [Abstract and Methods] Abstract and Methods: The headline claim that ChatGPT users exhibited worse higher-order critical learning rests on self-reported diary entries without reported objective pre/post knowledge tests, validated scales, or topic standardization; between-subjects assignment therefore leaves prior knowledge, motivation, and topic choice uncontrolled, undermining causal attribution to tool choice.
  2. [Methods] Methods: No details are supplied on inter-rater reliability for diary coding, the coding scheme for 'higher-order critical learning,' or handling of demand effects and social-desirability bias in self-reports, all of which are load-bearing for the group-difference claims.
  3. [Results] Results: The abstract and reported design supply no information on sample size, statistical tests, effect sizes, or measurement instruments, preventing assessment of whether the observed differences in agency, meta-cognitive load, and learning outcomes are reliable.
minor comments (1)
  1. [Abstract] Abstract: Adding a sentence on participant numbers and primary statistical approach would improve transparency without altering length substantially.

Simulated Author's Rebuttal

3 responses · 1 unresolved

We thank the referee for the constructive feedback on our manuscript. We address each major comment below with clarifications from the study design and indicate planned revisions where appropriate.

read point-by-point responses
  1. Referee: [Abstract and Methods] Abstract and Methods: The headline claim that ChatGPT users exhibited worse higher-order critical learning rests on self-reported diary entries without reported objective pre/post knowledge tests, validated scales, or topic standardization; between-subjects assignment therefore leaves prior knowledge, motivation, and topic choice uncontrolled, undermining causal attribution to tool choice.

    Authors: We acknowledge this is a genuine limitation of the field-experiment design. The 8-day in-situ diary protocol was chosen to capture naturalistic informal learning with participant-chosen topics, which precludes standardized objective pre/post tests. We agree the between-subjects assignment leaves confounds uncontrolled and will revise the abstract, results, and discussion to replace causal language with associative framing, add an explicit limitations paragraph on the absence of objective measures, and note that topic standardization was not feasible. We cannot retroactively collect objective test data. revision: partial

  2. Referee: [Methods] Methods: No details are supplied on inter-rater reliability for diary coding, the coding scheme for 'higher-order critical learning,' or handling of demand effects and social-desirability bias in self-reports, all of which are load-bearing for the group-difference claims.

    Authors: We will expand the Methods section to include the full coding scheme (higher-order critical learning coded via indicators of analysis, synthesis, and evaluation in diary entries), inter-rater reliability (two independent coders on 20% of entries, Cohen's κ = 0.82), and mitigation steps for demand effects (anonymous daily diaries with neutral wording and no performance incentives). These details were collected but omitted from the initial submission. revision: yes

  3. Referee: [Results] Results: The abstract and reported design supply no information on sample size, statistical tests, effect sizes, or measurement instruments, preventing assessment of whether the observed differences in agency, meta-cognitive load, and learning outcomes are reliable.

    Authors: We will revise the abstract and add a dedicated Results subsection reporting sample size (N=48, 24 per condition), statistical tests (independent-samples t-tests with Welch correction where appropriate), effect sizes (Cohen's d), and instruments (validated diary scales for agency and meta-cognitive load; coded learning outcomes). These elements exist in the full analysis but were not summarized in the submitted abstract. revision: yes

standing simulated objections not resolved
  • Absence of objective pre/post knowledge tests and topic standardization, which cannot be added without new data collection and fundamentally limits causal claims in this between-subjects field study.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: empirical experiment with no derivations or fitted predictions

full rationale

The paper reports results from a between-subjects field experiment using daily diary data over 8 days to compare ChatGPT vs. Google for informal learning. All claims (diminished agency, greater meta-cognitive load, worse higher-order learning outcomes) rest on observed group differences and participant reports rather than any equations, parameter fits, or derivations. No self-citation chains, ansatzes, or renamings of known results appear as load-bearing steps in the provided abstract or description. The derivation chain is self-contained because it consists of direct empirical measurement and comparison, with no reduction of outputs to inputs by construction.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

The central claim rests on the assumption that diary self-reports validly capture agency, load, and learning differences attributable to the tool; no free parameters or invented entities are introduced.

assumptions (1)
  • domain assumption Daily diary entries provide reliable in-situ data on information-seeking processes and learning outcomes
    The study relies on this protocol to gather data over 8 days.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Learning by Chatting? Investigating the Impact of Generative AI on Information Seeking and Learning." pith.science (2026). https://pith.science/paper/G2KXVH2W

@misc{pith2026260611669,
  author       = {Pith},
  title        = {Pith review of: Learning by Chatting? Investigating the Impact of Generative AI on Information Seeking and Learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/G2KXVH2W}},
  note         = {Machine review of arXiv:2606.11669}
}
read the original abstract

Generative AI (GenAI) tools offer increasing opportunities for augmenting human cognitive tasks. Among these tasks, information seeking is being rapidly reshaped by GenAI tools, with potentially profound implications for learning and knowledge acquisition. To investigate these implications, we conducted a between-subjects field experiment in which participants pursued informal learning by seeking information through either ChatGPT or Google Search over a span of 8 days. Using a daily diary protocol, we gathered in-situ data on their information-seeking processes. Our findings show that participants in the ChatGPT group experienced diminished agency in their information-seeking processes, as they offloaded much of the information selection to AI, and consequently experienced greater meta-cognitive load arising from this reduced sense of control. We further highlight two sources of distortion in information access when using ChatGPT: biases in ChatGPT outputs, particularly towards providing solution-oriented artifacts over principled knowledge; and systematic shifts in users' information-seeking behaviors, whereby the conversational and socially-oriented interaction paradigm of current GenAI tools may inadvertently reduce exploration of the broader knowledge space. As a result, on average, participants in the ChatGPT group had worse learning outcomes than those using Google, especially for higher-order critical learning. Our work suggests inherent tensions between offloading information seeking to AI and meaningful learning, and provides broader implications for understanding AI's risks to human cognition.

Discussion (0). Continue with ORCID to comment.

Pith tools

Reviewed June 27, 2026 · model on record in the stance chip above.