{"id":"3b74b68e-439d-479b-a44e-3ce1061831b4","arxiv_id":"2607.16223","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"Between 2024 and 2026, students at Ulster University normalised generative-AI use while staff caution persisted, with the paper claiming a widening student-staff gap.","lead":"A three-year survey at Ulster University tracked how students, doctoral researchers, teaching staff, and non-teaching staff perceive generative AI, finding reported familiarity and use climbed sharply while trust in AI for marking stayed very low. The study offers universities time-series evidence that AI attitudes shift quickly and that policy written for one year can be outdated by the next.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's central 'widening gap' claim is internally contradicted: §4.4.1 and §5.1.4 report opposite directions for the restriction item, and Table 3—the only inferential evidence—is missing from the manuscript.","rationale":"The reader's weakest_assumption focused on cross-wave comparability, which is a legitimate and important threat to any longitudinal trend claim. However, the more decisive, load-bearing problem is internal inconsistency: the paper's two result sections report opposite directions for the precise item that is said to define the widening student–staff gap. Because Table 3 is absent, the contradiction cannot be resolved from the manuscript, and the central claim is therefore not merely fragile but unverifiable in its current form. This reinforces the reader's REJECT verdict, but for a somewhat different reason than the stated weakest assumption. I still view the paper as potentially rehabilitable: if the authors supply Table 3 and the data and correct the contradictory prose, the descriptive trends (rising familiarity/use, low trust in AI marking) may survive. The concrete test targets the single item whose sign decides the headline narrative, so it directly settles whether the concern lands.","tokens_in":12065,"tokens_out":3340,"duration_ms":32898,"concrete_test":"Obtain Table 3 or the raw data and compute the estimated marginal mean change (β or difference) for the item 'Students should be restricted from using AI for coursework' by population and year, specifically 2024 vs 2026 and 2024 vs 2025, using the described lmer/emmeans pipeline. If the student β is positive (increased support for restriction) and the teaching-staff β is zero or negative, then §5.1.4's widening-gap claim is contradicted by the paper's own results and the central finding collapses. If the student β is negative and teaching-staff β positive, then §4.4.1 is wrong and needs correction. The check can be run by requesting the missing Table 3 from the authors or by reanalyzing the released data with the stated mixed-effects model.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract and §5.1.4 rest on a specific longitudinal result: students became less supportive of restricting AI for coursework while teaching staff became more supportive, producing 'the widening gap at the heart of this study.' But §4.4.1 reports the opposite for the same item: 'students’ ratings increased' on 'Students should be restricted from using AI for coursework,' PhD ratings decreased, and teaching/non-teaching staff showed no change. §5.2.1.2 repeats the §5.1.4 version ('students' views on restricting AI in coursework lessened over time, while teaching staff felt that they could increase'). These are mutually exclusive descriptions of the same analysis. The only place the actual inferential results appear is Table 3, which is not printed in the manuscript; no p-values, confidence intervals, or effect sizes are reported in the text. Therefore the reader cannot determine which direction is correct, and the headline gap-widening finding is unauditable and internally inconsistent. This is not a dispute with external consensus or a subtle design limitation; it is a direct contradiction within the paper's own results. If §4.4.1 is accurate, the student–staff gap did not widen on the very item used to define it; if §5.1.4 is accurate, §4.4.1 must be retracted. Either way, the central claim as written is not supported. Cross-wave instrument/sample comparability (§3.1–3.2) is a further concern, but it is secondary: even with perfect comparability, the sign contradiction remains unresolved.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports a three-wave longitudinal survey (2024, 2025, 2026) of four university populations (undergraduate students, doctoral researchers, teaching staff, non-teaching staff) at Ulster University, with a stated total N=1,665. The central claims are that familiarity, experience, and integration of generative AI rose across waves; that trust in AI marking remained low and flat; and that a student–staff gap widened, particularly on whether students should be restricted from using AI for coursework. The paper interprets these trends through the authors' earlier qualitative framework and draws policy implications about adaptive institutional guidance.","tokens_in":12313,"tokens_out":3404,"duration_ms":31412,"significance":"The study addresses a genuine gap: most AI-perception research is cross-sectional, and few datasets include doctoral and professional-services staff alongside undergraduates and teaching staff. The repeated-measures design, even if repeated cross-sectional, is valuable, and the descriptive pattern—rising familiarity without rising enthusiasm—is interesting and partly falsifiable. However, the quantitative evidence for the headline claims is not presently auditable: the only inferential table is empty, no test statistics are reported in the text, and the restriction-item result is described in opposite directions in two sections. If corrected and fully reported, the dataset could make a useful contribution; in its current form the manuscript does not support its main conclusions.","major_comments":[{"comment":"Table 3 is empty in the submitted manuscript. The text repeatedly refers to 'significant differences' and β values, but no coefficients, standard errors, p-values, confidence intervals, or effect sizes are given. All significance claims in Sections 4.1–4.6 and 5.1 are therefore unauditable. Please provide the complete Table 3 and report the model specification and post-hoc test statistics.","section":"§4, Table 3"},{"comment":"The paper contradicts itself on the direction of the restriction item. §4.4.1 states that students' ratings for 'Students should be restricted from using AI for coursework' increased and PhD ratings decreased. §5.1.4 states that students became less likely to support restriction while teaching staff became more likely, and §5.2.1.2 repeats this. The abstract's 'widening gap' claim depends on this item. The authors must identify which description is correct and correct the other; as written, the central result is internally inconsistent.","section":"§4.4.1 vs §5.1.4/§5.2.1.2"},{"comment":"All longitudinal conclusions assume wave-to-wave comparability. The manuscript states that 'minor changes to the survey framing' were made in waves 2 and 3 but does not specify them, and that each wave used a new self-selected sample (with £10 voucher) in which repeats were allowed but not tracked. Without evidence that rewording did not shift means and that samples are compositionally comparable, the trend estimates are confounded. Please document the exact wording changes and provide demographic comparisons across waves; if possible, analyse the repeat-participant subsample separately.","section":"§3.1–3.2"},{"comment":"Likert responses are treated as interval (1–5) in linear mixed models without robustness checks. Given the heavily skewed distributions described (e.g., AI-marking near floor), results should be verified with ordinal models or cumulative-link mixed models, or at least with sensitivity analyses using different numeric codings.","section":"§3.3"}],"minor_comments":[{"comment":"The abstract states total N=1,665, but Table 1's year totals sum to 2023 (780+599+644) and the population totals also sum to 2023. Please reconcile this discrepancy and confirm the correct sample size.","section":"§3.1, Table 1"},{"comment":"Many figures are referenced but not visible in the manuscript text. Please ensure all figures are inserted and legible, with axis labels and legend text readable.","section":"Figures throughout"},{"comment":"The qualitative framework from Gerard et al. (2026) is used to interpret quantitative results, but the manuscript does not describe the sample or method of that qualitative study beyond theme names. A brief description or a pointer to the published paper would help readers assess the fit.","section":"§5.2.1"},{"comment":"The data availability statement says data 'will be made available' but provides no repository or access link. Given the auditability concerns, please provide a permanent repository or explicit sharing plan.","section":"Data availability statement"}],"recommendation":"major_revision","confidential_remarks":"The missing Table 3 and the sign contradiction on the restriction item are serious, not cosmetic. If the authors can supply the full inferential results and correct the inconsistency, the paper may become publishable; I would be willing to review a revised version. The sample-size inconsistency in Table 1/abstract should also be fixed before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere's my take on Gerard et al. This paper is worth knowing about, but not as published. It reports a three-wave (2024–2026) survey of undergraduates, PhD researchers, teaching staff, and non-teaching staff at Ulster University on AI perceptions, with n=1,665. That design genuinely isn't in the cited literature: the closest, JISC's successive reports, lack the same-instrument, multi-population structure. The descriptive core is credible and useful—familiarity and experience with AI rose across all groups, use was increasingly 'addressed' in teaching, training participation rose, and trust in AI for marking stayed flat and low. The policy implications about adaptive training are common-sense and well argued.\n\nBut the paper has a load-bearing internal contradiction. §4.4.1 reports that on the item 'Students should be restricted from using AI for coursework,' students' ratings increased, PhD ratings decreased, and teaching/non-teaching staff were unchanged. §5.1.4 and §5.2.1.2 report the opposite: students became less supportive, teaching staff more supportive, and this divergence is called 'the widening gap at the heart of this study.' You cannot audit which is right, because the only inferential table—Table 3, with β values—is blank in the manuscript. No p-values, confidence intervals, or effect sizes appear anywhere. So the central claim is unauditable and internally inconsistent. That's not a minor slip; it's the finding the abstract rests on.\n\nThe design also weakens trend claims: each wave recruited a new sample, repeats were permitted but untracked, and the survey framing got 'minor' unspecified changes in waves 2 and 3. Likert means are treated as interval without comment. Data are promised but not shipped. None of these are fatal alone, but they mean every longitudinal conclusion rests on a comparability the paper never demonstrates. The interpretation through the authors' own qualitative framework (§5.2.1) is fine, but it is not independent confirmation.\n\nIf §4.4.1 is accurate, the gap didn't widen on the defining item. If §5.1.4 is accurate, §4.4.1 must be retracted. Either way, the paper as written doesn't establish its headline result. It is rehabilitable, though: the descriptive findings are probably real, the design is novel, and the authors need to correct the contradiction, publish the actual inferential table with effect sizes, and be honest about the repeated cross-sectional limits.\n\nWould I send it to review? Yes—the design deserves referee time, and the issues are fixable. But I wouldn't cite it until the data and corrected results are available.\n\nBest,\n[Your name]","headline":"Useful longitudinal survey of AI perceptions, but the central 'widening gap' finding is stranded by an internal contradiction and a missing inferential table.","tokens_in":12942,"tokens_out":2683,"would_cite":false,"duration_ms":23642,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI perceptions in higher education changed rapidly between 2024 and 2026: students normalised generative-AI use, staff caution persisted, and the gap between them widened.","keywords":["generative AI","higher education","longitudinal survey","student perceptions","staff perceptions","normalisation","academic integrity","AI policy"],"falsifier":"A direct test would reanalyse the data using only respondents who participated for the first time in each wave, and would also run the survey with identical wording across waves; if the significant increases in familiarity and integration and the widening restriction gap disappear under either condition, the trends are artefacts of measurement or sampling rather than real change.","tokens_in":1228,"feed_emoji":"🎓","tokens_out":4161,"duration_ms":44178,"temperature":0.7,"pith_summary":"This paper tries to establish that perceptions of generative AI in higher education are not static but change measurably over time, and different academic populations change at different rates. Using three annual survey waves at one university, it argues that students moved from tentative experimentation to routine integration of AI tools, while teaching staff maintained persistent concerns about integrity and assessment, so the student–staff divide grew. The authors also find that increased familiarity did not produce increased enthusiasm: students' perceived value of AI declined even as their use rose. If correct, the paper shows that institutional AI policies must be treated as time-sensitive and population-specific, because guidance written for one year can be outdated the next.","feed_headline":"Students normalised AI by 2026; staff caution didn't","feed_subtitle":"Three annual waves at one university show the AI gap widening as familiarity breeds scepticism.","key_machinery":"The central mechanism is a repeated cross-sectional longitudinal survey design: three annual waves (2024, 2025, 2026) using a stable core of Likert-scaled statements about AI familiarity, use, training, coursework, and institutional policy, administered to undergraduates, doctoral researchers, teaching staff, and non-teaching staff (total n=1,665). Analysis relies on a linear mixed-effects regression with rating as the dependent variable and population, year, and question as fixed effects, followed by posthoc pairwise comparisons between time points to locate where and for whom perceptions changed.","core_discovery":"The paper claims that between 2024 and 2026, students at Ulster University normalised generative AI use—familiarity, experience, and reported integration all rose significantly—while teaching staff continued to express caution about academic integrity, assessment design, and critical thinking. The central longitudinal finding is that the student–staff gap widened over the three waves, driven largely by opposite movements on whether students should be restricted from using AI: students became less supportive of restrictions, while teaching staff became more supportive. The paper also reports a counterintuitive trend: as students gained more experience with AI, their ratings of its value in ed","pith_inferences":["If the normalisation trend holds beyond this single institution, then cross-sectional studies of AI perceptions published in any given year are likely to be stale by the time they appear, and meta-analyses that pool studies across years may confound temporal change with population differences.","The paper's finding that exposure did not increase enthusiasm suggests a possible 'familiarity–scepticism' curve: early adoption is followed by critical appraisal as users encounter limitations; this could be tested by tracking whether the later waves show stabilisation or further decline.","The opposite-direction results between sections 4.4.1 and 5.1.4 on the restriction item—where the results text reports students increasing and PhDs decreasing, while the discussion frames it as students versus teaching staff—point to a need for the authors to clarify which pairwise comparisons drive the 'widening gap' claim.","Because repeat participation was permitted but not tracked, a reanalysis isolating first-time respondents in each wave would test whether the observed trends reflect genuine population change or a cohort of returning, more AI-engaged participants."],"forward_implications":["Institutional AI policy must be treated as time-varying: guidance written for 2023 is already outdated by 2025, so universities need adaptive frameworks reviewed on a short cycle.","AI literacy efforts should shift from introductory awareness-raising to critical engagement and discipline-specific application, since familiarity is now baseline for most students.","Training demand diverges by group: students' interest in further training declined as experience grew, while staff interest stayed steady, implying different training pathways are needed.","Trust in AI for marking is uniformly low and stable across all groups and years, so policies should separate AI as a learning aid from AI as an evaluative authority.","The widening student–staff gap on restriction suggests assessment design must account for routine student AI use while preserving expectations of independent thought."],"fun_headline_variants":["Students rush to AI, staff hold back: gap widens in 3-year study","AI normalised by students, not staff: 3-year shift","From novelty to routine: students outpace staff on AI","The AI normalisation gap: students embrace, staff resist over 3 years","AI attitudes diverge: students ahead, staff wary across 3 waves"],"cache_read_input_tokens":14080,"weakest_assumption_plain":"All longitudinal conclusions assume that the three waves are directly comparable—that minor wording changes in survey framing and the recruitment of a new self-selected sample each year (with repeat participation allowed but untracked) did not shift mean ratings or alter sample composition in ways that produce the observed changes.","fun_headline_variants_meta":{"raw":{"variants":["Students rush to AI, staff hold back: gap widens in 3-year study","AI normalised by students, not staff: 3-year shift","From novelty to routine: students outpace staff on AI","The AI normalisation gap: students embrace, staff resist over 3 years","AI attitudes diverge: students ahead, staff wary across 3 waves"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000822,"raw_usage":{"total_tokens":3431,"prompt_tokens":741,"completion_tokens":2690,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":485,"completion_tokens_details":{"reasoning_tokens":2594}},"tokens_in":485,"tokens_out":2690,"duration_ms":16988,"temperature":1.0,"reasoning_tokens":2594,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T10:46:29.342633+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct test would reanalyse the data using only respondents who participated for the first time in each wave, and would also run the survey with identical wording across waves; if the significant increases in familiarity and integration and the widening restriction gap disappear under either condition, the trends are artefacts of measurement or sampling rather than real change.","supporting_citations":[],"review_version":1}