REVIEW 4 major objections 5 minor 42 references
Pilot Study on Generative AI and Critical Thinking in Higher Education Classrooms
T0 review · 4 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read This pilot study reports that a 15-minute video lesson on evaluating generative-AI output improved students' scores on a related critical-thinking assignment in one introductory data science course, with p = 0.0319 on a Kruskal-Wallis test
desk verdict A transparent but statistically fragile pilot study that honestly labels its own weak evidence; useful as a template, not as a demonstration of effectiveness. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The intervention is a 15-minute prerecorded video introducing large language models, hallucinations, and two worked examples of flawed ChatGPT answers, plus companion exercises; the control condition is the same four-question written assignment without the video. Because the samples are small and deviate from normality, the comparison uses the Kruskal-Wallis nonparametric test, with assignment grades as the response variable and video exposure as a binary treatment indicator. The null hypothesis is that the treatment and control distributions are the same, and the reported p-value is the load-bearing numerical result of the paper.
What would settle it
Re-run the CDS 101 comparison after excluding students who earned 0 (did not submit), as the authors themselves did; the p-value rises above 0.05. A pre-registered larger study with students randomly assigned within the same section that also finds no treatment difference would settle the central claim. Alternatively, compare Session A and Session C students on their other coursework to check whether the groups were unequal before the video.
Extended reading notes
Core claim
The paper claims that students in the Summer 2025 CDS 101 course who received a 15-minute prerecorded lesson on how to evaluate generative-AI output scored higher on a four-prompt critical-thinking assignment than control-section students who did not receive the lesson (Kruskal-Wallis p = 0.0319, treatment n = 6 versus control n = 4). The authors state that 'we have evidence that the video lessons do have a positive impact on student outcomes for the related assignment,' while immediately adding that removing students who earned 0 by not submitting eliminates statistical significance, and that assigning treatment and control to different summer sessions is an assumption rather than a proven
Load-bearing premise
The treatment and control groups are assumed to be comparable even though they came from different summer sessions, with voluntary participation and no pre-test or background data, so differences in student ability or motivation—not the video—could explain the higher average score.
Editorial extensions
If this is right
- If the video lesson is genuinely effective, a low-cost, scalable 15-minute intervention can improve students' ability to analyze, critique, and revise AI-generated answers in introductory data science courses.
- The observed voluntary participation rate of roughly 25–30% informs the logistics and recruitment plans for larger multi-section studies.
- The positive result gives a 'weak prior' that justifies continuing and expanding the experiment to more sections, courses, and semesters before drawing strong conclusions.
- Because the effect disappears when non-submitting students are excluded, the current evidence supports designing better-controlled replications rather than immediate adoption at scale.
Reading between the lines
- The significant p-value could reflect session-level differences—Session A and Session C students may differ in ability, motivation, or grading conditions—so the most defensible reading is that this study demonstrates feasibility, not a proven causal effect.
- A sharper test would randomly assign students within a single section or collect a pre-test; if the treatment effect disappears when non-submitters are excluded, the video may boost assignment submission or engagement rather than critical-thinking skill itself.
- The mechanism worth testing is whether students learned a repeatable evaluation rubric (such as checking for hallucinations and walking through examples) or merely imitated the worked examples; a transfer task using novel AI outputs would separate these explanations.
- Because CDS 130 had too few participants for analysis, pooling data across multiple courses in future semesters could reveal whether the effect generalizes beyond one course and one instructor pairing.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This pilot study reports a quasi-experiment in two introductory computational/data science courses (CDS 101 and CDS 130) at George Mason University. Students in treatment sections watched a prerecorded 15-minute video on evaluating generative AI outputs and then completed a written assignment; control sections completed the same assignment without the video. The main quantitative analysis is a Kruskal-Wallis (KW) test comparing CDS 101 treatment (n=6, Summer Session A) and control (n=4, Summer Session C) assignment scores, yielding p = 0.0319. The authors describe this as 'weak evidence' that the video has a positive impact, while acknowledging that removing students who scored 0 (non-submitters) removes statistical significance and that assigning treatment and control to different sessions is an assumption. Participation rates were about 28% overall. The paper also reports pilot logistics (e.g., voluntary participation, section enrollment) and concludes that 'we have some evidence that the lesson plan is effective,' calling for larger future studies.
Significance. If the result were credible, this would be a useful contribution to the literature on GAI and critical thinking, an area with limited empirical work. The manuscript has notable strengths: a transparent description of a simple, nonparametric analysis; explicit acknowledgement of key limitations; open data and code via GitHub; and an honest presentation of the result as a 'weak prior.' The choice of KW over ANOVA for small, non-normal samples is appropriate. However, the central claim rests on a single unadjusted p-value from very small, non-equivalent groups. The paper's own sensitivity disclosure—significance disappears when non-submitters are excluded—undermines the causal interpretation. As a pilot, it is useful for planning future studies, but the current evidence does not support the stated conclusion that the lesson plan is effective.
major comments (4)
- [Section III.B / Section IV.B] The KW test in Equations (10)–(11) is a two-sided omnibus test of whether the two distributions differ; it does not establish the direction of the difference. The manuscript states that 'the video lesson impacts performance' and later claims a 'positive impact,' but no group medians, means, effect size, or confidence interval are reported. Please report the direction and magnitude (e.g., group medians, rank-biserial correlation or Cliff's delta, and a bootstrap/permutation CI) so the reader can assess whether the difference is substantively positive.
- [Section IV.B] The treatment (CDS 101 A01, Summer Session A) and control (CDS 101 C01, Summer Session C) are drawn from different sessions with different voluntary participation rates (35% vs. 25%) and no pre-test or covariate data. The statement 'we are assuming that assigning a treatment and control in different sessions adequately satisfies the experimental design' is load-bearing: any session-specific factor (student ability, motivation, prior AI exposure, grading timing, workload) is a plausible confound that could explain the observed p-value. The causal language in Sections III.B and V should be softened to describe an associational pilot result, or the authors should provide a sensitivity analysis demonstrating robustness to plausible confounds (e.g., permutation tests, covariate adjustment if any demographic data exist).
- [Section IV.B] The paper acknowledges that removing students who scored 0 (non-submitters) removes statistical significance, but this analysis is not reported in the main text. This is a critical robustness check: if the significant KW result is driven by differential submission behavior rather than by critical-thinking performance, the central claim collapses. Please report the KW test on non-zero scores with the exact p-value and effect size, and if possible analyze submission/non-submission rates as an outcome. This should be a primary result, not a post hoc caveat.
- [Table II / Section II.D] Participation is voluntary and the analysis includes only students who agreed to participate. If participation is correlated with student characteristics (e.g., conscientiousness or prior interest in AI), the treatment and control groups are not exchangeable even within a session. The paper provides no information about nonparticipants. At minimum, the authors should state this explicitly as a self-selection confound and avoid causal claims; ideally, they could compare participants to the full enrolled population on available characteristics (e.g., final course grade, GPA).
minor comments (5)
- [Section II.B] This section is a textbook review of z-tests, t-tests, OLS, ANOVA, and the normal distribution (Equations 1–9). It is not necessary for the pilot analysis and could be condensed to one paragraph, focusing only on the KW test and its assumptions.
- [Throughout] There are multiple typos: 'Kruskall-Wallis' should be 'Kruskal-Wallis'; 'treament' in Section IV.B; 'useage' in Section I; 'instramental' in Section VI; 'Anderson' etc. Please proofread. Also, 'ANOV A' has a spacing issue.
- [Table I] The assignment name 'W A 13.5' for CDS 130 is unclear; please clarify whether this is 'WA 13.5' or another course-specific identifier.
- [Figure 2] The QQ-plot shows tail deviations from normality, but the text does not state which normality test (if any) was used or how strongly the tails deviate. A Shapiro-Wilk or Anderson-Darling test would be more informative than visual inspection alone.
- [Section II.A / Reference [37]] The definitions of AI and GAI are adopted from the first author's prior work [37]; citing one's own work for definitions is acceptable, but the definitions are unusual ('creation of content with a spatial component' for GAI). Please clarify the source's context or cite additional primary definitions.
Circularity Check
No circularity: the treatment/control comparison is empirical and self-contained; the only self-citations are for definitions and a routine statistical assumption, neither of which feeds the result.
full rationale
This is an empirical pilot study, not a derivation. The central claim in Section IV.B — that the video lessons have a positive impact on assignment outcomes — rests on a Kruskal-Wallis test comparing assignment scores of students who did and did not receive the video. The outcome variable (assignment score) is measured independently of the treatment, and no parameter is fitted from the outcome and then renamed as a prediction. The only self-citations are [37] (definitions of AI and GAI) and [40] (a normality assumption for OLS residuals); neither is load-bearing for the statistical result, and both could be replaced by standard references without changing the analysis. The paper's own caveats about small samples, voluntary participation, and different summer sessions are threats to causal validity, not evidence of circularity. Therefore the derivation chain is self-contained, and any weaknesses are design/statistical concerns rather than circular reasoning.
Assumptions & free parameters
assumptions (4)
- domain assumption The assignment rubric measures critical thinking about GAI output.
- domain assumption Treatment and control groups are exchangeable despite different summer sessions.
- domain assumption Voluntary participation does not introduce selection bias.
- standard math Kruskal-Wallis test assumptions are met (independent samples, ordinal outcome, no extreme ties).
Cite this review
Pith. "Pith review of Pilot Study on Generative AI and Critical Thinking in Higher Education Classrooms." pith.science (2026). https://pith.science/paper/DNLYVVIK
@misc{pith2026250900167,
author = {Pith},
title = {Pith review of: Pilot Study on Generative AI and Critical Thinking in Higher Education Classrooms},
year = {2026},
howpublished = {\url{https://pith.science/paper/DNLYVVIK}},
note = {Machine review of arXiv:2509.00167}
}
read the original abstract
Generative AI (GAI) tools have seen rapid adoption in educational settings, yet their role in fostering critical thinking remains underexplored. While previous studies have examined GAI as a tutor for specific lessons or as a tool for completing assignments, few have addressed how students critically evaluate the accuracy and appropriateness of GAI-generated responses. This pilot study investigates students' ability to apply structured critical thinking when assessing Generative AI outputs in introductory Computational and Data Science courses. Given that GAI tools often produce contextually flawed or factually incorrect answers, we designed learning activities that require students to analyze, critique, and revise AI-generated solutions. Our findings offer initial insights into students' ability to engage critically with GAI content and lay the groundwork for more comprehensive studies in future semesters.
Figures
Reference graph
Works this paper leans on
-
[1]
The intertwined histories of artificial intelligence and education,
S. Doroudi, “The intertwined histories of artificial intelligence and education,” International Journal of Artificial Intelligence in Education , vol. 33, no. 4, pp. 885–928, 2023
work page 2023
-
[2]
O. Zawacki-Richter, V . I. Mar ´ın, M. Bond, and F. Gouverneur, “Sys- tematic review of research on artificial intelligence applications in higher education–where are the educators?” International journal of educational technology in higher education , vol. 16, no. 1, pp. 1–27, 2019
work page 2019
-
[3]
T. K. Chiu, Q. Xia, X. Zhou, C. S. Chai, and M. Cheng, “Systematic literature review on opportunities, challenges, and future research rec- ommendations of artificial intelligence in education,” Computers and Education: Artificial Intelligence , vol. 4, p. 100118, 2023
work page 2023
-
[4]
Artificial Intelligence Index Report 2024,
N. Maslej, L. Fattorini, R. Perrault, V . Parli, A. Reuel, E. Brynjolfsson, J. Etchemendy, K. Ligett, T. Lyons, J. Manyika, J. C. Niebles, Y . Shoham, R. Wald, and J. Clark, “Artificial Intelligence Index Report 2024,” May 2024, arXiv:2405.19522. [Online]. Available: http://arxiv.org/abs/2405.19522
arXiv 2024
-
[5]
P. Kitcharoen, S. Howimanporn, and S. Chookaew, “Enhancing teachers’ ai competencies through artificial intelligence of things professional development training.” International Journal of Interactive Mobile Tech- nologies, vol. 18, no. 2, 2024
work page 2024
-
[6]
Openai and microsoft bankroll new a.i. training for teachers,
N. Singer, “Openai and microsoft bankroll new a.i. training for teachers,” The New York Times , 2025
work page 2025
-
[7]
Microsoft, “AFT to launch National Academy for AI Instruction with Microsoft, OpenAI, Anthropic and United Federation of Teachers,” textscurl: urlhttps://news.microsoft.com/source/2025/07/08/aft-to-launch-national- academy-for-ai-instruction-with-microsoft-openai-anthropic-and-united- federation-of-teachers/, jul 2025
work page 2025
-
[8]
B. Plato, C. Emlyn-Jones, and W. Preddy, “Lysis; symposium; phaedrus,” (No Title), 2022
work page 2022
Show all 42 references
-
[9]
University of Notre Dame Press, 1994
Discours de La Methode/Discourse on the Method: A Bilingual Edition with an Interpretive Essay . University of Notre Dame Press, 1994. [Online]. Available: http://www.jstor.org/stable/j.ctv1bvnf2j
1994
-
[10]
God’s machines: Descartes on the mechanization of mind,
O. Holland, P. Husbands, and M. Wheeler, “God’s machines: Descartes on the mechanization of mind,” in The Mechanical Mind in History . United States: MIT Press, 2008
2008
-
[11]
V ANHAELEN, AUTOMATA: Activating Human Behavior
A. V ANHAELEN, AUTOMATA: Activating Human Behavior . Penn State University Press, 2022, pp. 77–96. [Online]. Available: http: //www.jstor.org/stable/10.5325/jj.5233096.9
2022 doi
-
[12]
How the u.s. public and ai experts view artificial intelligence,
P. R. Center, “How the u.s. public and ai experts view artificial intelligence,” Pew Research Center, 2025, accessed: [July 14, 2025]. [Online]. Available: https://www.pewresearch.org/internet/2025/04/03/ how-the-us-public-and-ai-experts-view-artificial-intelligence/
2025
-
[13]
Impact of genera- tive ai on critical thinking skills in undergraduates: A systematic review,
P. Premkumar, M. Yatigammana, and S. Kannangara, “Impact of genera- tive ai on critical thinking skills in undergraduates: A systematic review,” Journal of Desk Research Review and Analysis , vol. 2, no. 1, 2024
2024
-
[14]
Is it harmful or helpful? examining the causes and consequences of generative ai usage among university students,
M. Abbas, F. A. Jam, and T. I. Khan, “Is it harmful or helpful? examining the causes and consequences of generative ai usage among university students,” International journal of educational technology in higher education, vol. 21, no. 1, p. 10, 2024
2024
-
[15]
Generative artificial intelligence amplifies the role of critical thinking skills and reduces reliance on prior knowledge while promoting in-depth learning,
G. Zhao, H. Sheng, Y . Wang, X. Cai, and T. Long, “Generative artificial intelligence amplifies the role of critical thinking skills and reduces reliance on prior knowledge while promoting in-depth learning,” Education Sciences, vol. 15, no. 5, p. 554, 2025
2025
-
[16]
Critical thinking in the age of generative ai: Effects of a short-term experiential learning intervention on efl learners,
N. Cong-Lem, T. Tat Nguyen, and K. Nhat Hoang Nguyen, “Critical thinking in the age of generative ai: Effects of a short-term experiential learning intervention on efl learners,” 2025
2025
-
[17]
Harnessing the power of ai to education,
K. Srinivasa, M. Kurni, and K. Saritha, “Harnessing the power of ai to education,” in Learning, teaching, and assessment methods for contemporary learners: pedagogy for the digital generation . Springer, 2022, pp. 311–342
2022
-
[18]
Factors affecting university students’ generative ai literacy: Evidence and eval- uation in the uk and hong kong contexts,
X. O’Dea, D. Tsz Kit Ng, M. O’Dea, and V . Shkuratskyy, “Factors affecting university students’ generative ai literacy: Evidence and eval- uation in the uk and hong kong contexts,” Policy Futures in Education , p. 14782103241287401, 2024
2024
-
[19]
Chatgpt for good? on opportunities and challenges of large language models for education,
E. Kasneci, K. Seßler, S. K ¨uchemann, M. Bannert, D. Dementieva, F. Fischer, U. Gasser, G. Groh, S. G ¨unnemann, E. H ¨ullermeier et al. , “Chatgpt for good? on opportunities and challenges of large language models for education,” Learning and individual differences , vol. 10...
2023
-
[20]
Can automated feedback improve teachers’ uptake of student ideas? evidence from a randomized controlled trial in a large-scale online course,
D. Demszky, J. Liu, H. C. Hill, D. Jurafsky, and C. Piech, “Can automated feedback improve teachers’ uptake of student ideas? evidence from a randomized controlled trial in a large-scale online course,” Educational Evaluation and Policy Analysis , vol. 46, no. 3, pp. 483– 505, 2024
2024
-
[21]
Reshaping curriculum adap- tation in the age of artificial intelligence: Mapping teachers’ ai-driven curriculum adaptation patterns,
F. Karatas ¸, B. Eric ¸ok, and L. Tanrikulu, “Reshaping curriculum adap- tation in the age of artificial intelligence: Mapping teachers’ ai-driven curriculum adaptation patterns,” British Educational Research Journal , vol. 51, no. 1, pp. 154–180, 2025
2025
-
[22]
Grading exams using large language models: A comparison between human and ai grading of exams in higher education using chatgpt,
J. Flod ´en, “Grading exams using large language models: A comparison between human and ai grading of exams in higher education using chatgpt,” British educational research journal , vol. 51, no. 1, pp. 201– 224, 2025
2025
-
[23]
The future of grading programming assignments in education: the role of chatgpt in automating the assessment and feedback process. think skills creativity. 2024; 52: 101522
M. Jukiewicz, “The future of grading programming assignments in education: the role of chatgpt in automating the assessment and feedback process. think skills creativity. 2024; 52: 101522.”
2024
-
[24]
How do physics students evaluate artificial intelligence responses on comprehension questions? a study on the perceived scientific accuracy and linguistic quality of chatgpt,
M. N. Dahlkemper, S. Z. Lahme, and P. Klein, “How do physics students evaluate artificial intelligence responses on comprehension questions? a study on the perceived scientific accuracy and linguistic quality of chatgpt,” Physical Review Physics Education Research , vol. 19, n...
2023
-
[25]
Learning to fake it: limited responses and fabricated references provided by chatgpt for medical questions,
J. Gravel, M. D’Amours-Gravel, and E. Osmanlliu, “Learning to fake it: limited responses and fabricated references provided by chatgpt for medical questions,” Mayo Clinic Proceedings: Digital Health , vol. 1, no. 3, pp. 226–234, 2023
2023
-
[26]
Your brain on chatgpt: Accu- mulation of cognitive debt when using an ai assistant for essay writing task,
N. Kosmyna, E. Hauptmann, Y . T. Yuan, J. Situ, X.-H. Liao, A. V . Beresnitzky, I. Braunstein, and P. Maes, “Your brain on chatgpt: Accu- mulation of cognitive debt when using an ai assistant for essay writing task,” arXiv preprint arXiv:2506.08872 , 2025
2025 arXiv
-
[27]
Are skepticism and moderation dominating attitudes toward ai-based technologies?
S.-V . Oprea, I. Nica, A. B ˆara, and I.-A. Georgescu, “Are skepticism and moderation dominating attitudes toward ai-based technologies?” American Journal of Economics and Sociology , vol. 83, no. 3, pp. 567– 607, 2024
2024
-
[28]
The impact of generative ai on critical thinking: Self- reported reductions in cognitive effort and confidence effects from a survey of knowledge workers,
H.-P. Lee, A. Sarkar, L. Tankelevitch, I. Drosos, S. Rintel, R. Banks, and N. Wilson, “The impact of generative ai on critical thinking: Self- reported reductions in cognitive effort and confidence effects from a survey of knowledge workers,” in Proceedings of the 2025 CHI con...
2025
-
[29]
Chatgpt as an educational tool: Opportunities, challenges, and recommendations for communication, business writing, and composition courses,
M. A. AlAfnan, S. Dishari, M. Jovic, and K. Lomidze, “Chatgpt as an educational tool: Opportunities, challenges, and recommendations for communication, business writing, and composition courses,” Journal of Artificial Intelligence and Technology , vol. 3, no. 2, pp. 60–68, 2023
2023
-
[30]
Chatting and cheating: Ensuring academic integrity in the era of chatgpt,
D. R. Cotton, P. A. Cotton, and J. R. Shipway, “Chatting and cheating: Ensuring academic integrity in the era of chatgpt,” Innovations in education and teaching international, vol. 61, no. 2, pp. 228–239, 2024
2024
-
[31]
Can GPT-3 write an academic paper on itself, with minimal human input?
G. Generative Pretrained Transformer, A. O. Thunstr ¨om, and S. Steingrimsson, “Can GPT-3 write an academic paper on itself, with minimal human input?” Jun. 2022, working paper or preprint. [Online]. Available: https://hal.science/hal-03701250
2022
-
[32]
Google users are less likely to click on links when an ai summary appears in the results,
P. R. Center, “Google users are less likely to click on links when an ai summary appears in the results,” Pew Research Center, 2025, accessed: [July 24, 2025]
2025
-
[33]
34% of u.s. adults have used chatgpt, about double the share in 2023,
——, “34% of u.s. adults have used chatgpt, about double the share in 2023,” Pew Research Center, 2025, accessed: [July 24, 2025]. [Online]. Available: https://www.pewresearch.org/short-reads/2025/06/ 25/34-of-us-adults-have-used-chatgpt-about-double-the-share-in-2023/
2023
-
[34]
Global Public Opinion on Artificial Intelligence (GPO-AI),
P. J. Loewen, B. Lee-Whiting, Arai, T. Bergeron, T. Galipeau, I. Gazendam, H. Needham, L. Slinger, and S. Yusypovych, “Global Public Opinion on Artificial Intelligence (GPO-AI),” Tech. Rep., 2024. [Online]. Available: https://srinstitute.utoronto.ca/public-opinion-ai
2024
-
[35]
Students’ voices on generative ai: Per- ceptions, benefits, and challenges in higher education,
C. K. Y . Chan and W. Hu, “Students’ voices on generative ai: Per- ceptions, benefits, and challenges in higher education,” International Journal of Educational Technology in Higher Education , vol. 20, no. 1, p. 43, 2023
2023
-
[36]
Artificial Intelligence and Public Policy,
A. D. Thierer, A. Castillo O’Sullivan, and R. Russell, “Artificial Intelligence and Public Policy,” Rochester, NY , Aug. 2017. [Online]. Available: https://papers.ssrn.com/abstract=3021135
2017
-
[37]
Artificial Intelligence Policy Framework for Institutions,
W. F. Lamberti, “Artificial Intelligence Policy Framework for Institutions,” Dec. 2024, arXiv:2412.02834 [cs]. [Online]. Available: http://arxiv.org/abs/2412.02834
2024 arXiv
-
[38]
G. K. Bhattacharyya and R. A. Johnson, Statistical Concepts and Methods, 1st ed. Wiley, 1977. [Online]. Available: https://www.wiley. com/en-us/Statistical+Concepts+and+Methods-p-9780471072041
1977
-
[39]
Mendenhall and T
W. Mendenhall and T. T. Sincich, A Second Course in Statistics: Regression Analysis, 7th ed. Boston, MA: Pearson, Jan. 2011
2011
-
[40]
An Overview of Explainable and Interpretable Artificial Intelligence,
W. F. Lamberti, “An Overview of Explainable and Interpretable Artificial Intelligence,” 2022
2022
-
[41]
N. R. Draper and H. Smith, Applied Regression Analysis , third edi- tion ed. New York: Wiley-Interscience, Apr. 1998
1998
-
[42]
G. M. University. (1999) Mason core. [Online]. Available: https: //catalog.gmu.edu/mason-core/ APPENDIX IRB D ETAILS The research plan for the institutional review board (IRB) was conducted at GMU via the Office of Research Integrity and Assurance. The review, feedback, and ap...
1999
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.