REVIEW 4 major objections 4 minor 6 references
From Novelty to Normalisation: Tracking Changing Perceptions of AI in Higher Education, 2024-2026
T0 review · 4 major / 4 minor · reviewed 2026-08-02 · deepseek-v4-flash
Pith's one-line read AI perceptions in higher education changed rapidly between 2024 and 2026: students normalised generative-AI use, staff caution persisted, and the gap between them widened.
desk verdict Useful longitudinal survey of AI perceptions, but the central 'widening gap' finding is stranded by an internal contradiction and a missing inferential table. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is a repeated cross-sectional longitudinal survey design: three annual waves (2024, 2025, 2026) using a stable core of Likert-scaled statements about AI familiarity, use, training, coursework, and institutional policy, administered to undergraduates, doctoral researchers, teaching staff, and non-teaching staff (total n=1,665). Analysis relies on a linear mixed-effects regression with rating as the dependent variable and population, year, and question as fixed effects, followed by posthoc pairwise comparisons between time points to locate where and for whom perceptions changed.
What would settle it
A direct test would reanalyse the data using only respondents who participated for the first time in each wave, and would also run the survey with identical wording across waves; if the significant increases in familiarity and integration and the widening restriction gap disappear under either condition, the trends are artefacts of measurement or sampling rather than real change.
Extended reading notes
Core claim
The paper claims that between 2024 and 2026, students at Ulster University normalised generative AI use—familiarity, experience, and reported integration all rose significantly—while teaching staff continued to express caution about academic integrity, assessment design, and critical thinking. The central longitudinal finding is that the student–staff gap widened over the three waves, driven largely by opposite movements on whether students should be restricted from using AI: students became less supportive of restrictions, while teaching staff became more supportive. The paper also reports a counterintuitive trend: as students gained more experience with AI, their ratings of its value in ed
Load-bearing premise
All longitudinal conclusions assume that the three waves are directly comparable—that minor wording changes in survey framing and the recruitment of a new self-selected sample each year (with repeat participation allowed but untracked) did not shift mean ratings or alter sample composition in ways that produce the observed changes.
Editorial extensions
If this is right
- Institutional AI policy must be treated as time-varying: guidance written for 2023 is already outdated by 2025, so universities need adaptive frameworks reviewed on a short cycle.
- AI literacy efforts should shift from introductory awareness-raising to critical engagement and discipline-specific application, since familiarity is now baseline for most students.
- Training demand diverges by group: students' interest in further training declined as experience grew, while staff interest stayed steady, implying different training pathways are needed.
- Trust in AI for marking is uniformly low and stable across all groups and years, so policies should separate AI as a learning aid from AI as an evaluative authority.
- The widening student–staff gap on restriction suggests assessment design must account for routine student AI use while preserving expectations of independent thought.
Reading between the lines
- If the normalisation trend holds beyond this single institution, then cross-sectional studies of AI perceptions published in any given year are likely to be stale by the time they appear, and meta-analyses that pool studies across years may confound temporal change with population differences.
- The paper's finding that exposure did not increase enthusiasm suggests a possible 'familiarity–scepticism' curve: early adoption is followed by critical appraisal as users encounter limitations; this could be tested by tracking whether the later waves show stabilisation or further decline.
- The opposite-direction results between sections 4.4.1 and 5.1.4 on the restriction item—where the results text reports students increasing and PhDs decreasing, while the discussion frames it as students versus teaching staff—point to a need for the authors to clarify which pairwise comparisons drive the 'widening gap' claim.
- Because repeat participation was permitted but not tracked, a reanalysis isolating first-time respondents in each wave would test whether the observed trends reflect genuine population change or a cohort of returning, more AI-engaged participants.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript reports a three-wave longitudinal survey (2024, 2025, 2026) of four university populations (undergraduate students, doctoral researchers, teaching staff, non-teaching staff) at Ulster University, with a stated total N=1,665. The central claims are that familiarity, experience, and integration of generative AI rose across waves; that trust in AI marking remained low and flat; and that a student–staff gap widened, particularly on whether students should be restricted from using AI for coursework. The paper interprets these trends through the authors' earlier qualitative framework and draws policy implications about adaptive institutional guidance.
Significance. The study addresses a genuine gap: most AI-perception research is cross-sectional, and few datasets include doctoral and professional-services staff alongside undergraduates and teaching staff. The repeated-measures design, even if repeated cross-sectional, is valuable, and the descriptive pattern—rising familiarity without rising enthusiasm—is interesting and partly falsifiable. However, the quantitative evidence for the headline claims is not presently auditable: the only inferential table is empty, no test statistics are reported in the text, and the restriction-item result is described in opposite directions in two sections. If corrected and fully reported, the dataset could make a useful contribution; in its current form the manuscript does not support its main conclusions.
major comments (4)
- [§4, Table 3] Table 3 is empty in the submitted manuscript. The text repeatedly refers to 'significant differences' and β values, but no coefficients, standard errors, p-values, confidence intervals, or effect sizes are given. All significance claims in Sections 4.1–4.6 and 5.1 are therefore unauditable. Please provide the complete Table 3 and report the model specification and post-hoc test statistics.
- [§4.4.1 vs §5.1.4/§5.2.1.2] The paper contradicts itself on the direction of the restriction item. §4.4.1 states that students' ratings for 'Students should be restricted from using AI for coursework' increased and PhD ratings decreased. §5.1.4 states that students became less likely to support restriction while teaching staff became more likely, and §5.2.1.2 repeats this. The abstract's 'widening gap' claim depends on this item. The authors must identify which description is correct and correct the other; as written, the central result is internally inconsistent.
- [§3.1–3.2] All longitudinal conclusions assume wave-to-wave comparability. The manuscript states that 'minor changes to the survey framing' were made in waves 2 and 3 but does not specify them, and that each wave used a new self-selected sample (with £10 voucher) in which repeats were allowed but not tracked. Without evidence that rewording did not shift means and that samples are compositionally comparable, the trend estimates are confounded. Please document the exact wording changes and provide demographic comparisons across waves; if possible, analyse the repeat-participant subsample separately.
- [§3.3] Likert responses are treated as interval (1–5) in linear mixed models without robustness checks. Given the heavily skewed distributions described (e.g., AI-marking near floor), results should be verified with ordinal models or cumulative-link mixed models, or at least with sensitivity analyses using different numeric codings.
minor comments (4)
- [§3.1, Table 1] The abstract states total N=1,665, but Table 1's year totals sum to 2023 (780+599+644) and the population totals also sum to 2023. Please reconcile this discrepancy and confirm the correct sample size.
- [Figures throughout] Many figures are referenced but not visible in the manuscript text. Please ensure all figures are inserted and legible, with axis labels and legend text readable.
- [§5.2.1] The qualitative framework from Gerard et al. (2026) is used to interpret quantitative results, but the manuscript does not describe the sample or method of that qualitative study beyond theme names. A brief description or a pointer to the published paper would help readers assess the fit.
- [Data availability statement] The data availability statement says data 'will be made available' but provides no repository or access link. Given the auditability concerns, please provide a permanent repository or explicit sharing plan.
Circularity Check
No significant circularity: the survey comparisons are independent measurements; self-citations are interpretive, not load-bearing.
full rationale
This paper is an observational longitudinal survey, not a derivation chain. The Likert responses across waves and populations are independent measurements, and the paper's trend claims are comparisons of means over time/groups rather than quantities that are equal to their inputs by construction. No parameter is fitted to a subset of the data and then used to 'predict' a closely related quantity; the analysis plan (regression + posthoc pairwise comparisons) tests whether means differ, and the reported differences are empirical outcomes. The self-citations are not load-bearing in a circular sense: Gerard et al. (2025) is cited for the repeated-survey design and Gerard et al. (2026) is used as a qualitative interpretive framework, but the quantitative results stand on the survey items themselves, which were adapted from Petricini et al. (2024). The manuscript's internal contradiction about the direction of the restriction-item change (§4.4.1 reports students' ratings increased, while §5.1.4 reports students became less likely to support restriction) and the absence of Table 3 are serious correctness and auditability problems, not circularity. They indicate that the reported 'widening gap' finding may be unreliable or internally inconsistent, but they do not show that any conclusion is forced by definition, by a fitted input, or by a self-citation chain. No circular step can be exhibited with the paper's own equations or construction, so the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (2)
- Ordinal-to-interval Likert mapping (1–5)
- Unspecified wave-2/wave-3 survey-framing changes
assumptions (4)
- domain assumption Ordinal Likert categories can be treated as interval-scale measurements for parametric mixed-effects modeling
- domain assumption Wave-to-wave comparability of measures despite unfixed sampling and undisclosed item rewording
- domain assumption Each wave's volunteer sample represents its underlying Ulster University population
- domain assumption Self-reported familiarity and use track actual engagement with AI
Cite this review
Pith. "Pith review of From Novelty to Normalisation: Tracking Changing Perceptions of AI in Higher Education, 2024-2026." pith.science (2026). https://pith.science/paper/HSWJFKV2
@misc{pith2026260716223,
author = {Pith},
title = {Pith review of: From Novelty to Normalisation: Tracking Changing Perceptions of AI in Higher Education, 2024-2026},
year = {2026},
howpublished = {\url{https://pith.science/paper/HSWJFKV2}},
note = {Machine review of arXiv:2607.16223}
}
read the original abstract
The rapid integration of generative artificial intelligence (AI) has reshaped the landscape of higher education. Students have embraced tools such as ChatGPT with striking speed, while teaching staff and institutions have responded with greater caution. Existing research on AI perceptions has mainly been cross-sectional, providing single-point snapshots that view attitudes as stable rather than evolving. This paper presents a longitudinal study of AI perceptions in higher education, tracking undergraduates, doctoral researchers, teaching staff and non-teaching staff at Ulster University across three survey waves between 2024 and 2026 (n=1,665). A quantitative survey design measured familiarity, reported use and perceived risk; results show that students rapidly normalised AI use over the period, moving from tentative experimentation to routine engagement, while staff expressed persistent concerns about academic integrity, assessment design, and critical thinking. Doctoral and non-teaching staff occupied intermediate positions, reflecting both pragmatic adoption and institutional caution. The student-staff gap widened as institutional policy struggled to keep pace with actual practice. By tracking these shifts directly rather than reconstructing them from disconnected studies, the paper moves beyond descriptive accounts of AI attitudes and demonstrates the importance of capturing perceptions in real time. The findings carry significant implications for adaptive institutional policy, AI literacy initiatives, and targeted staff training.
Figures
Reference graph
Works this paper leans on
-
[1]
Alshamy, A., Al-Harthi, A. S. A., & Abdullah, S. (2025). Perceptions of generative AI tools in higher education: Insights from students and academics at Sultan Qaboos University. Education Sciences, 15(4), Article
2025
-
[43]
https://doi.org/10.1186/s41239-023-00411-8 Contractor, Z., & Reyes, G. (2025). Generative AI in higher education: Evidence from an elite college (arXiv:2508.00717). arXiv. https://doi.org/10.48550/arXiv.2508.00717 Department for Education. (2023). Generative artificial intelligence in education call for evidence: Summary of responses. https://www.gov.uk/g...
work page Pith review arXiv doi:10.48550/arxiv.2508.00717 2025
-
[87]
https://doi.org/10.26209/td2024vol17iss21825 Quality Assurance Agency for Higher Education. (2023a). Maintaining quality and standards in the ChatGPT era: QAA advice on the opportunities and challenges posed by generative artificial intelligence. https://www.qaa.ac.uk/sector-resources/generative-artificial- intelligence/qaa-advice-and-resources Quality As...
-
[501]
https://doi.org/10.3390/educsci15040501 15 Bates, D., Mächler, M., Bolker, B., & Walker, S. (2015). Fitting linear mixed-effects models using lme4. Journal of Statistical Software, 67(1), 1–48. https://doi.org/10.18637/jss.v067.i01 Chan, C. K. Y., & Hu, W. (2023). Students’ voices on generative AI: Perceptions, benefits, and challenges in higher education...
-
[1039]
https://doi.org/10.3390/educsci15081039 JISC. (2023a). A generative AI primer. National Centre for AI in Tertiary Education. https://nationalcentreforai.jiscinvolve.org JISC. (2023b). Student perceptions of generative AI. https://www.jisc.ac.uk/reports/student- perceptions-of-generative-ai JISC. (2025). Student perceptions of AI
-
[2025]
https://www.jisc.ac.uk/reports/student- perceptions-of-ai-2025 Lenth, R. V. (2025). emmeans: Estimated marginal means, aka least-squares means [R package]. https://CRAN.R-project.org/package=emmeans Petricini, T., Zipf, S., & Wu, C. (2024). Perceptions about generative AI and ChatGPT use by faculty and students. Transformative Dialogues: Teaching and Lear...
2025
Reviewed August 2, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.