{"id":"539fb7ba-e86f-4af1-a88a-f34f0dbb3659","arxiv_id":"2505.00100","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Across two semesters and four courses, a structured GenAI literacy lab shifted students' self-reported comfort and openness toward AI without increasing self-reported use on graded work.","lead":"This paper evaluates AI-Lab, a short classroom intervention that teaches undergraduates how to use generative AI tools responsibly. Pre/post surveys and focus groups indicate students became more comfortable and open to using AI for concepts and debugging, while self-reported use on graded homework stayed flat.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The causal claim is untested: the pre/post design has no control group, and the paper itself (Section 4.2) documents a policy shift and stark semester differences in baseline usage that could explain the observed attitude changes.","rationale":"The reader's weakest_assumption is exactly the concern I would raise: no control group and a documented policy shift. I agree. I considered two other candidate concerns. First, the reported rank-biserial effect sizes are suspiciously large given small raw shifts on 5-point items, likely inflated by tied discordant pairs; however, even if effect sizes were corrected, the significance tests would probably survive, so this is not the load-bearing issue. Second, the paper lists 'openness to debugging' as both significant (Section 3.2 list item 2) and non-significant (Table 13 p=0.6592; Section 4.3), but the abstract does not claim openness to debugging changed, so this internal inconsistency, while it needs correction, does not undermine the central claim. The central claim is causal and the design cannot support causal inference without a non-intervention comparison or explicit modeling of semester/policy effects. I am not objecting to the plausibility of the intervention; focus groups and stable graded-work frequency are useful supportive evidence. My recommendation is therefore not to reject the paper but to keep it conditional on a causal-identification check. Since the reader's verdict is already CONDITIONAL, no adjustment is needed.","tokens_in":16402,"tokens_out":4992,"duration_ms":55168,"concrete_test":"Run the paired analyses separately for S24 and F24 and fit an ordinal mixed-effects model for the three headline perception items (openness/conceptual, comfort/conceptual, comfort/homework) with time, semester, and time x semester as fixed effects and a random intercept per student (or per student-course if duplicates exist). If the within-semester time effect is absent or much smaller in F24, or if the time x semester interaction is significant, then the policy/familiarity confound remains plausible and the causal claim fails. If the estimated time effect is comparable and significant within both semesters, the policy confound is weakened, though a true control condition would still be needed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is causal: a short AI-Lab can shift students' self-reported comfort with and willingness to use GenAI and shape their engagement strategies. The supporting evidence is a paired pre/post design with no non-intervention comparison group. The paper's own Section 4.2 documents a major confound: between Spring and Fall 2024, Purdue moved from recommending to requiring AI statements in course syllabi, and Figure 5 shows baseline homework-use frequency differs sharply by semester (low-use categories: 28.49% in S24 vs 52.30% in F24). The combined-semester Wilcoxon tests therefore cannot separate the intervention effect from concurrent policy change, maturation, or increasing familiarity with GenAI over the semester. The qualitative focus-group themes are retrospective self-reports and are consistent with demand characteristics; they do not supply the missing counterfactual. The absence of a significant increase in self-reported use on graded work is a null result and cannot establish that the intervention prevented increases. A well-executed paired study with large N is a real strength, but it does not address causal identification. If the central claim is to stand as 'can shift', the design must rule out or model the semester/policy confound.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a mixed-methods evaluation of the AI-Lab, a scaffolded intervention for teaching undergraduate CS and engineering students to use generative AI deliberately. Across two semesters and three CS courses plus one engineering course, the authors collected paired pre/post surveys (N≈830) and six post-intervention focus groups. With Wilcoxon signed-rank tests they find statistically significant increases in several self-reported openness, comfort, and debugging-frequency items, while homework/project use frequency stays flat, and they interpret the corresponding rank-biserial correlations (0.88–0.94) as large effects. The qualitative themes describe students adopting more iterative prompting, greater skepticism of outputs, and clearer boundaries around integrity.","tokens_in":16564,"tokens_out":4365,"duration_ms":43974,"significance":"If the causal claim were supported, this would be a useful contribution to the emerging literature on GenAI pedagogy: it is one of the larger explicitly evaluated interventions, it pairs quantitative and qualitative data, and it explicitly documents a semester-level policy shift that confounds the design (Section 4.2). The honesty about the confound is a strength, and the large paired sample is a real asset. However, the central claim that a short intervention can shift attitudes is not established by the design, and the reported 'large effect sizes' are not supported by the raw distributions. The paper has value as a pilot or exploratory evaluation, but the current framing overstates what the evidence can show.","major_comments":[{"comment":"The pre/post design has no non-intervention comparison group, and the paper itself documents a major confound: between Spring and Fall 2024, Purdue moved from recommending to requiring AI statements in course syllabi, and Figure 5 shows that self-reported baseline homework-use frequency differs sharply by semester (low-use categories: 28.49% in S24 vs 52.30% in F24). The combined-semester Wilcoxon tests therefore cannot separate the intervention effect from concurrent policy change, maturation, or growing familiarity with GenAI over the semester. The causal language in the abstract ('can shift') and in Section 4.5 ('can promote') is not supported by the evidence. To keep the causal claim, the authors should provide a semester-specific analysis with an explicit modeling of the policy change, or reframe the claim as 'students reported shifts' and openly state that the design cannot rule out confounders.","section":"§4.2, Fig. 5"},{"comment":"The claim of 'large effect sizes' is not supported by the raw distributions. For instance, openness to conceptual questions moves from 50.6% to 53.2% at the top category (Table 3), and comfort with conceptual questions moves from 35.3% to 38.0% at the top category (Table 6). With several hundred students, nearly all responses are unchanged, so a rank-biserial correlation of 0.90–0.94 is driven by the few discordant pairs that almost all go in one direction; it is not evidence that most students changed substantially. The authors should report (i) the proportion of students who increased, decreased, and stayed the same for each item, and (ii) a more interpretable effect size such as Cliff's delta with a confidence interval, and should temper the 'substantial and meaningful changes' language in Section 3.2.1.","section":"§2.2.2, §3.2.1, Tables 3–8, Table 13"},{"comment":"Table 13 reports the openness-to-debugging item with p=0.6592 and no effect size, yet the list of statistically significant questions in Section 3.2 includes 'How open are you to using GenAI to get help with debugging?' as item (2). This is a direct contradiction that must be resolved; it appears that the significance list is wrong and the table is correct, but as written the paper's own results are internally inconsistent.","section":"Table 13 and list in §3.2"},{"comment":"The qualitative focus-group themes are retrospective self-reports collected by the same researchers who implemented the intervention, and the statement in §4.4 that 'These reflections were not influenced by pressure from the intervention or researchers' is unsupported. Demand characteristics are a plausible source of the observed reports (e.g., students echoing the intervention's stated goals of 'mindful usage'). The focus groups cannot provide a counterfactual, so they do not rescue the causal interpretation. This limitation should be acknowledged explicitly.","section":"§3.3, §4.4"}],"minor_comments":[{"comment":"RQ2 is phrased as a yes/no question ('Are Students Open to Using GenAI...?'), but the study analyzes shifts in Likert-scale openness/comfort. The research question should be aligned with the outcome measured.","section":"§1.2"},{"comment":"The effect-size interpretation thresholds (0.30–0.50 medium, >0.50 large) are presented without citation; give a reference or justify the cutoffs.","section":"§2.2.2"},{"comment":"The number of focus-group participants and the total number of groups (six) are mentioned, but the number of participants per group is not reported; include this for transparency.","section":"§2.3"},{"comment":"The discussion of peer-usage perceptions is confusing: the text says 'only around 50% of students predicted correctly' but then reports that 84.54% believed ≥50% of peers used GenAI, while 75.42% of students themselves used it on occasion or more. Please clarify the claim and the arithmetic.","section":"§3.4"},{"comment":"The statement that ENGR showed 'a higher intervention impact' is based on the count of significant p-values; this is not a valid comparison of effect size, especially with a small (n=53) sample. Use an effect-size comparison or drop the claim.","section":"§4.1, Table 14"}],"recommendation":"major_revision","confidential_remarks":"The paper has a useful dataset and a transparent discussion of the semester confound, but the causal framing and the 'large effect size' interpretation both need substantial rework. I would encourage the editor to send this back with clear guidance that the authors address the causal identification gap and report more meaningful effect-size metrics. The paper also relies heavily on the authors' own prior work as the framework of record; that is not disqualifying, but the novelty framing should be checked."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: this is the largest evaluation yet of the AI-Lab framework, with paired surveys from 831 students across two semesters and six focus groups. The qualitative material is the real asset. Students describe moving from clumsy, abandoned prompts to more iterative, skeptical engagement, and they articulate their own boundaries around integrity. Those quotes are credible and useful.\n\nThe paper is also transparent in ways that help a reviewer. It publishes full pre/post distributions, breaks down counts by course and semester, and in Section 4.2 openly documents a confound: between Spring and Fall 2024, Purdue moved from recommending to requiring AI statements in syllabi, and baseline homework-use frequency differs sharply by semester (28.5% vs 52.3% in the low-use categories). That honesty is good, but it makes the main analysis hard to defend.\n\nThe soft spots are three. First, the effect sizes are overstated. The rank-biserial values around 0.9 are labeled 'large,' but the raw tables show only a few percentage points shifting. With Likert data and many tied pairs, a Wilcoxon-derived rank-biserial can be high even when the practical change is small. The paper needs a more interpretable effect size, or at least a paragraph reconciling the two. Second, there is an internal contradiction: Section 3.2 lists 'open to using GenAI for debugging' as significant, while Table 13 reports p=0.66 for that item and Section 4.3 says it was not significant. That is a clear error. Third, the central claim that the intervention 'can shift' attitudes is causal, but the design is pre/post with no control, and the documented policy change is a plausible alternative explanation. The focus groups were retrospective and could reflect demand characteristics. The null finding on graded-use frequency is exactly that: a null, not evidence the intervention prevented increases.\n\nWhat holds up: the direction of the shifts is probably real, and the qualitative themes are coherent. The paper is a solid exploratory study, not a rigorous causal evaluation. The authors also get credit for explicitly calling their semester explanation 'pure conjecture.'\n\nBottom line: this deserves a serious referee. A good review should push for reframed causal claims, per-semester analysis or a comparison cohort, a corrected significance list, and a less grandiose effect-size interpretation. I'd bring it to our reading group as a case study in the difference between paired pre/post data and causal inference.","headline":"Useful mixed-methods evaluation of a GenAI literacy lab, but the causal claim outruns the pre/post design and the reported 'large' effects rest on an inflated effect-size metric.","tokens_in":17171,"tokens_out":3081,"would_cite":true,"duration_ms":33032,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A short, structured lab session can shift students' reported comfort and willingness to use generative AI without increasing self-reported use on graded work.","keywords":["Generative AI (GenAI)","core skill development","pedagogical framework","AI-Lab","quantitative qualitative mixed methods","computing education","student perceptions","self-reported usage"],"falsifier":"A reader could settle the causal claim by running the same pre/post surveys in the same courses with a randomly assigned no-intervention control group, or by comparing semesters in which the intervention was offered to those in which it was not while holding the syllabus-policy requirement constant; if control students show the same comfort and openness gains, the AI-Lab is not the active ingredient.","tokens_in":16158,"feed_emoji":"🤖","tokens_out":5600,"duration_ms":49200,"temperature":0.7,"pith_summary":"The paper evaluates the AI-Lab, a short scaffolded intervention in which students learn prompting, critique generated outputs in class, and reflect on a homework assignment using generative AI. Across two semesters in three computer science courses and one first-year engineering course, paired pre/post surveys (831 perception responses, 826 usage responses) and six focus groups tracked how students' reported openness, comfort, frequency of use, and prompting strategies changed. The central claim is that even a short intervention can make students more comfortable and open to using generative AI for conceptual, debugging, and homework help, and can shift how they prompt and evaluate outputs, without increasing self-reported frequency of use on graded homework or projects. The authors argue this matters because it addresses the faculty concern that teaching AI use encourages academic dishonesty, while giving students a more deliberate, learning-oriented way to engage the tools.","feed_headline":"A short lab shifts student AI comfort without raising homework use","feed_subtitle":"Across 778 computer science students, openness and comfort rose while self-reported use on graded work stayed flat.","key_machinery":"The AI-Lab framework is a four-stage pedagogical sequence: instructor preparation (selecting a topic likely to expose AI errors), a prelab with preparatory materials and a baseline survey, an in-class lab in which the instructor demonstrates GenAI output and students critique and correct it collaboratively, and a post-lab homework requiring documented attempts to steer the AI plus a follow-up survey. This staged structure is the mechanism that is claimed to convert naive experimentation into deliberate, critical use. The quantitative evaluation rests on paired Wilcoxon signed-rank tests with rank-biserial effect sizes, and the qualitative analysis on thematic coding of focus-group transcripts.","core_discovery":"The central discovery, stated as the authors would state it, is that the AI-Lab intervention changes the quality of students' engagement with generative AI more than the quantity. Students' self-reported openness to using GenAI for conceptual questions and homework help increased, and comfort increased for conceptual, debugging, and homework scenarios, with rank-biserial correlations in the 'large' range. Frequency of use for homework and programming projects did not change significantly, while frequency of use for debugging increased. In focus groups, students described moving from trial-and-error prompting to iterative, context-rich prompting, becoming skeptical of confidently stated but incorrect outputs, and articulating explicit boundaries about when not to use AI. The paper interprets this as evidence that structured scaffolding can promote mindful, reflective AI use rather than overreliance.","pith_inferences":["An inference the authors leave implicit: because the study has no control group and the university began requiring AI syllabus statements between the two semesters, the reported shifts may partly reflect growing institutional and societal familiarity with GenAI rather than the AI-Lab alone; the paper's own semester comparison documents a large baseline difference.","The qualitative themes suggest a testable extension: behavioral trace data (e.g., actual prompt logs before and after the intervention) could determine whether 'iterative prompting' and 'skepticism' translate into measurable changes in tool interaction rather than only self-report.","A plausible next experiment would separate the intervention's components — the in-class critique, the homework reflection, and the prelab materials — to identify which stage drives the comfort and openness gains.","The semester difference also implies that the intervention's effects may be time-dependent, so replicating it now versus a year ago may yield different baselines; future evaluations should treat institutional AI policy as a covariate."],"forward_implications":["If the central claim holds, computing educators can introduce GenAI tools through a short structured activity without measurable increases in self-reported use on graded assignments, addressing a common faculty worry about cheating.","Students' prompting strategies and skepticism of AI outputs are teachable within a single lab session, suggesting that the 'AI literacy' that matters is less about raw tool exposure and more about critique and reflection.","The finding that self-reported use for debugging increased while homework use stayed flat implies that scaffolding can channel GenAI use toward lower-stakes, skill-building activities.","The larger shift seen in the first-year engineering course, compared with CS courses, suggests the intervention may help students with less prior tool exposure more, motivating adoption outside CS majors.","Reported desire to use GenAI mostly shifted toward more use in CP, DSA-CS, and ENGR but toward less use in DSA-DSAI, hinting that effects depend on students' starting point."],"supporting_citations":[{"why":"Introduces the AI-Lab framework that this paper implements and evaluates.","marker":"[8]"},{"why":"Reports the earlier Spring 2024 pilot of the AI-Lab in two data structures courses, which this study extends.","marker":"[3]"},{"why":"Supplies the scaffolding-and-fading pedagogical approach that frames the intervention.","marker":"[5]"},{"why":"Provides the thematic analysis method used to code the focus-group transcripts.","marker":"[4]"},{"why":"Supplies the coding manual used in the qualitative analysis.","marker":"[20]"},{"why":"Justifies the focus-group methodology for gathering the qualitative data.","marker":"[11]"}],"fun_headline_variants":["AI lab boosts student comfort, not homework use","Scaffolded AI lab changes engagement, not usage","Brief AI training shifts student attitudes, not behavior","AI-Lab: more mindful use, same homework frequency","Short AI intervention improves student comfort, keeps use flat"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the pre-to-post changes in students' self-reported attitudes are caused by the AI-Lab intervention itself, even though the same period included a university policy change on AI syllabus statements, growing student familiarity with GenAI, and no control group.","fun_headline_variants_meta":{"raw":{"variants":["AI lab boosts student comfort, not homework use","Scaffolded AI lab changes engagement, not usage","Brief AI training shifts student attitudes, not behavior","AI-Lab: more mindful use, same homework frequency","Short AI intervention improves student comfort, keeps use flat"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000182,"raw_usage":{"total_tokens":1334,"prompt_tokens":995,"completion_tokens":339,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":611,"completion_tokens_details":{"reasoning_tokens":265}},"tokens_in":611,"tokens_out":339,"duration_ms":3493,"temperature":1.0,"reasoning_tokens":265,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:51:35.926315+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could settle the causal claim by running the same pre/post surveys in the same courses with a randomly assigned no-intervention control group, or by comparing semesters in which the intervention was offered to those in which it was not while holding the syllabus-policy requirement constant; if control students show the same comfort and openness gains, the AI-Lab is not the active ingredient.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the coding manual used in the qualitative analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Justifies the focus-group methodology for gathering the qualitative data."}],"review_version":1}