{"id":"deca8014-68e4-4568-9e26-3d26176be04b","arxiv_id":"2412.15624","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Students found generative AI helpful for incremental learning and initial implementation, but challenging for first-time concept learning and advanced implementation, with causes traced to intrinsic AI faults and gaps.","lead":"This paper interviewed 16 software engineering students to map when and how they use generative AI tools for learning and project work, and where those tools help or hurt. It gives educators a phase-by-phase picture of benefits and challenges, plus their causes and consequences, to guide curriculum decisions.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The four-phase benefit/challenge pattern may be an artifact of post-hoc phase coding of recalled experiences; independent re-coding is needed to confirm.","rationale":"The paper is a carefully conducted qualitative study with transparent methods, and its claims are mostly scoped to participants' perceptions. The reader's ACCEPT is defensible. However, the single most load-bearing element is the strong 'only' claim about the four-phase distribution of benefits and challenges. That claim is the foundation for the teaching recommendations in Section VI. The data are retrospective self-reports, as the paper's Section VII limitation states, and the phase categories were derived post hoc by the same team that reports the pattern. No inter-rater reliability is reported, so we cannot distinguish a true phenomenon from a coding or confirmation artifact. This is not a claim of dishonesty; it is a standard methodological gap. The proposed independent re-coding is a concrete, feasible check that would settle the issue. Because that check is not currently in the paper, I recommend CONDITIONAL acceptance: accept the paper with the condition that the authors either provide the re-coding evidence or soften the 'only' claim to a sample-level tendency, such as 'in our sample, participants predominantly reported benefits in L2 and I1 and challenges in L1 and I2.' This preserves the paper's useful qualitative contribution while aligning its central claim with its evidentiary strength.","tokens_in":20283,"tokens_out":5500,"duration_ms":49748,"concrete_test":"Obtain the interview transcripts and the published codebook from the supplemental material [56]. Have two independent researchers, blind to the paper's four-phase model and hypotheses, code every reported benefit and challenge into L1, L2, I1, I2, or 'unclear.' Compute inter-rater agreement (e.g., Cohen's kappa) and then check whether the diagonal pattern—benefits only in L2/I1, challenges only in L1/I2—replicates in both coders' outputs. If kappa is below 0.6, or if either coder finds a benefit reported in L1/I2 or a challenge in L2/I1, the central claim is not robust and the paper's 'only' phrasing should be softened.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim in Section IV-C is that benefits were perceived 'only' for incremental learning (L2) and initial implementation (I1), while challenges were encountered 'only' for initial learning (L1) and advanced implementation (I2). This clean diagonal pattern depends on reliably classifying each reported experience into L1, L2, I1, or I2. The authors state in Section III-B that two authors open-coded and negotiated, but they report no independent inter-rater reliability metric. They also explicitly acknowledge in Section VII 'the absence of concrete tasks conducted by the students during the study, limiting our results to our participants' recollections.' Because the phase labels were defined post hoc from the same interviews, and participants were not directly observed performing tasks, the phase attribution is an interpretive reconstruction. The 'only' claim is vulnerable to two compounding threats: (1) recall and desirability biases shaping which experiences participants volunteered, and (2) coders unconsciously fitting those recollections into the emerging L1/L2/I1/I2 scheme. If either threat is real, the central pattern could be an artifact rather than a property of students' lived experience. The unusually clean diagonal pattern—benefits on one diagonal, challenges on the other—would be surprising if experiences were heterogeneous, and no inter-rater agreement is provided to show the pattern is not idiosyncratic to the coding team.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a qualitative interview study of 16 software engineering (SE) students and two SE instructors about students' academic use of generative AI (genAI) tools. The authors identify four phases of use—initial learning (L1), incremental learning (L2), initial implementation (I1), and advanced implementation (I2)—and claim that participants perceived benefits only in L2 and I1, while challenges were encountered only in L1 and I2 (Section IV-C). They further analyze the causes of these challenges, attributing them to six intrinsic genAI issues (faults and gaps) that produce five challenge categories (C1–C5), which in turn impact learning, task outcomes, self-perception, and adoption of genAI (Section V). The findings are validated through member checking with students and triangulation with instructors. The authors explicitly acknowledge reliance on retrospective self-report, the absence of concrete tasks during the study, single-university sampling, and other threats to validity (Section VII).","tokens_in":20517,"tokens_out":4818,"duration_ms":44216,"significance":"If the central pattern holds, the paper offers an actionable map for SE educators deciding where genAI can be integrated into curricula and where students are likely to struggle. The study has notable strengths: interviews were anchored in participants' actual conversation histories with genAI tools, saturation was explicitly tested, the analysis used reflexive thematic analysis with consensus-based team meetings, and the authors performed member checking and instructor triangulation. The companion website with the codebook and supplemental material is a further positive. The principal risk is that the strong 'only' claim in Section IV-C rests on post-hoc phase coding of retrospective accounts, without inter-rater reliability evidence or a participant-level breakdown, so the clean diagonal pattern could be an artifact of the coding scheme rather than a robust property of the lived experiences.","major_comments":[{"comment":"The central claim that participants 'perceived the benefits of using genAI only for incremental learning (L2) and initial implementation tasks (I1)' and 'encountered challenges ... for initial learning (L1) and advanced implementations (I2)' is stronger than the evidence currently presented. Section III-B describes open coding with subsequent team negotiation, but no inter-rater reliability metric is reported, and the phase labels (L1, L2, I1, I2) were defined post hoc from the same interview data. Section VII acknowledges the absence of concrete tasks, meaning the phase attribution is entirely retrospective. Given these conditions, the exclusivity of the diagonal pattern could plausibly be an artifact of the coding scheme. Please provide a supplemental matrix showing which participants reported which benefit and challenge codes in which phases, or soften the 'only' formulations to 'clustered in' or 'were reported primarily in.'","section":"Section IV-C, Figure 1"},{"comment":"The acknowledged limitation that the study involved no concrete tasks means that phase attributions rely on participants' recollections. The member checking and instructor triangulation validate the presence of the benefit and challenge categories, but they cannot confirm the exclusivity of the phase mapping, because the mapping itself is an interpretive reconstruction. An independent re-coding of a subset of transcripts by a researcher not involved in the original analysis, with agreement statistics reported, would substantially strengthen the central claim. Without such evidence, the 'only' language in Section IV-C and the teaching recommendations built on it (Section VI) overstate the support.","section":"Section VII"}],"minor_comments":[{"comment":"Figure 1 contains serious rendering artifacts—stray question marks, duplicated text, and irregular line breaks—that obscure the phase-benefit/challenge mapping. Please provide a clean, publication-ready figure.","section":"Figure 1"},{"comment":"The sentence 'we eschewed from discussing frequency or percentage of occurrences of categories' should read 'we eschewed discussing frequency or percentage of occurrences of categories.'","section":"Section VII"},{"comment":"Reference [63] lists 'V . Clark' but the correct author name is 'V. Clarke'; please also check for inconsistent spacing in author initials across other references.","section":"References"},{"comment":"The arrows in Figure 2 from 'genAI's intrinsic issues' to challenges and impacts are based on participants' own causal attributions. The paper should clarify in the text that these are perceived associations, not verified causal links observed by the researchers, to avoid overstating the evidence.","section":"Section V, Figure 2"},{"comment":"The statement 'The first author transcribed the interviews' is useful for transparency, but the paper could also briefly describe the transcription accuracy check (if any) to align with standard reporting in qualitative studies.","section":"Section III-B"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a well-conducted qualitative study with a transparent and applied contribution. The main issue is the mismatch between the strong 'only' claim and the exploratory, retrospective, single-coder-consensus evidence base. This is fixable either by providing a participant-by-phase code distribution or by softening the claim. I do not see grounds for rejection; the concerns are about precise interpretation, not about fundamental methodological invalidity. The self-citations to [19] and [76] are appropriate given the authors' prior work on closely related questions. The paper seems well within scope for a venue interested in computing education or human-AI interaction."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper before the next round of genAI-in-education discussions. It is a solid, honestly-scoped qualitative study that adds a genuinely useful phase map to the literature. The four phases (L1, L2, I1, I2) and the diagonal pattern—benefits in incremental learning and initial implementation, challenges in initial learning and advanced implementation—are new and actionable, even if they aren't proven in a strong causal sense. The reader's ACCEPT verdict is about right, and the stress-test worry is real but shouldn't sink it.\n\nWhat's genuinely new: prior work reported benefits and challenges from students and instructors, but this paper organizes them by learning/implementation phase and chains genAI faults/gaps to five challenge categories to four impacts. That structure is a real extension, and the teaching recommendations follow from it. The method is standard for interview studies: 16 reflective interviews anchored on conversation histories, member checking, instructor triangulation, and a companion site with the interview script and codebook. The paper openly lists threats to validity, including the absence of concrete tasks and reliance on recollections. Good practice, and credit for shipping artifacts.\n\nSoft spots: the central 'only' claim in Section IV-C is the load-bearing one, and it rests on retrospective self-reports plus post-hoc phase coding. No inter-rater reliability metric is reported, and the pattern is suspiciously clean. That does not make the result wrong—novices lacking domain knowledge and advanced implementers needing deep context are both plausible reasons for the diagonal—but it does mean the map should be read as a model of student perceptions, not measured reality. The stress test's demand for independent re-coding is fair; I'd call it a moderate caveat, not a fatal flaw, because the paper frames the claims as perceived and explicitly acknowledges the recall limitation.\n\nThe theory tie-ins (CLT, SCT, SDT, TAM) are used lightly as interpretive lenses, not tests. Fine for this kind of work.\n\nWho it's for: SE educators and computing-education researchers. If you're building policy or teaching materials around genAI, the phase map gives a concrete scaffold: scaffold or restrict in L1 and I2, allow guided use in L2 and I1, and teach verification. This paper deserves a serious referee—the topic is timely, the method is defensible, and the contribution is clear. I'd send it out.","headline":"A transparent, small-scale interview study whose four-phase benefit/challenge map is a real contribution; the 'only' pattern is softer than it looks, but the paper deserves review.","tokens_in":21041,"tokens_out":2333,"would_cite":true,"duration_ms":21454,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"GenAI helps SE students best at mid-steps, not first steps","keywords":["generative AI","software engineering education","student perceptions","benefits and challenges","thematic analysis","human-AI interaction","qualitative interviews","AI literacy"],"falsifier":"Run a controlled task-based study in which students are assigned representative SE tasks from each phase (L1, L2, I1, I2), with and without genAI assistance, and measure learning gains, task completion, time, and frustration. If students show genAI helping in L1 or I2, or hurting in L2 or I1, the claimed four-phase pattern is contradicted.","tokens_in":20118,"feed_emoji":"🤖","tokens_out":5773,"duration_ms":45029,"temperature":0.7,"pith_summary":"This paper tries to establish where generative AI tools genuinely help software engineering students and where they get in the way, and why. Based on reflective interviews with 16 students, validated by member checking and instructor interviews, it argues that benefits concentrate in two specific phases: incremental learning of a concept students already partly know, and the initial steps of a concrete implementation. It argues that challenges concentrate in the other two phases: learning a concept from scratch and pushing an implementation to advanced, integrated work. The paper then traces those challenges back to intrinsic issues in the tools themselves (faults like hallucination and context neglect, gaps like missing scaffolding and weak debugging support) and forward to four kinds of impact on students: learning, task outcomes, self-perception, and adoption of the technology. A sympathetic reader would read this as a map of where genAI should and should not be leaned on in SE education, and why.","feed_headline":"GenAI helps SE students best at mid-steps, not first steps","feed_subtitle":"16 student interviews map where genAI helps and where it fails in SE coursework.","key_machinery":"The central object is a two-by-two phase model that splits SE coursework into initial learning (L1), incremental learning (L2), initial implementation (I1), and advanced implementation (I2). The analysis places perceived benefits and challenges into these quadrants and then builds a cause-consequence network (Figure 2) that connects genAI's intrinsic faults and gaps, through five challenge categories, to impacts on learning, task, self, and adoption. The phase model does the explanatory work: it shows why the same tool can be helpful in one context and harmful in another, and it turns scattered student complaints into a single testable pattern.","core_discovery":"The central claim is that students' lived experience with genAI clusters into a four-phase pattern: benefits appear in incremental learning (L2) and initial implementation (I1), while challenges appear in initial learning (L1) and advanced implementation (I2). The paper further claims that the challenges are not random difficulties but follow a causal chain: genAI's intrinsic faults (reasoning flaws, response-quality issues, deceptive behavior, neglect of student context) and gaps (scaffolding gaps, programming-support gaps) produce five challenge categories (C1-C5: unclear understanding of the tool, difficulty communicating needs, difficulty aligning AI to process and preferences, issues obtaining rationales, and difficulty using responses), which then produce four impacts (on learning, on task completion, on self-perception, and on willingness to adopt genAI). The pattern is meant to guide curriculum design: let students use genAI for clarification and initial scaffolding, but teach novices without it and prepare students for verification, prompt-crafting, and ethical judgment when work becomes advanced.","pith_inferences":["The same four-phase pattern may extend beyond software engineering to other project-based disciplines, where initial concept acquisition and advanced integration are likely the points where AI assistance fails most.","Because the data are retrospective interviews, the paper maps where problems occur but not how often or how strongly; a quantitative survey or log-based study could attach frequencies to the five challenge categories.","The instructor triangulation revealed that instructors did not expect the emotional toll of genAI struggles, which suggests student support should address frustration and self-doubt, not just technical outcomes.","The pattern implies a sharper pedagogical rule than 'use it after mastering basics': genAI is safe for reinforcing known material and for jump-starting concrete tasks, but it is not a reliable tutor for first exposure, so educators should design for that asymmetry."],"forward_implications":["Curriculum designers can use the phase map to decide where genAI use should be encouraged, scaffolded, or restricted, rather than banning it outright.","Novice students need explicit instruction in prompt crafting, output verification, and adapting AI responses, because those are the skills that fail hardest in the challenging phases.","Assignments that require students to explain and justify AI-suggested solutions would directly target the reported lack of rationales (C4).","Improving the tools themselves, such as reducing deceptive behavior and adding debugging support, would shrink the challenge categories at L1 and I2.","Universities need clear authorship and ethical-use policies, since ethical uncertainty (C1) already steers some students away from using genAI in advanced work."],"supporting_citations":[{"why":"It supplies the closest prior study of ChatGPT in SE tasks, which this work extends beyond controlled tasks.","marker":"[19]"},{"why":"It contributes the instructor-interview framing and the technique of anchoring interviews in artifacts, which this study adopts.","marker":"[21]"},{"why":"It provides prior evidence that code-generation tools can frustrate programmers, grounding the challenge categories.","marker":"[36]"},{"why":"It underlies the claim that novices over-rely on generated code because they cannot verify it.","marker":"[12]"},{"why":"It frames over-reliance on Copilot as an educational risk, motivating the study.","marker":"[13]"},{"why":"It supplies the broader debate on generative AI in computing education that this paper situates itself within.","marker":"[9]"},{"why":"It documents aligned student and instructor concerns about over-reliance and trustworthiness, which this paper refines into a phase map.","marker":"[31]"},{"why":"It provides the reflexive thematic analysis method used to derive the codes and themes.","marker":"[63]"}],"fun_headline_variants":["GenAI aids SE students mid-task, not at first or last steps","Four-phase pattern: genAI helps only in middle steps of SE work","SE student interviews map genAI's help to mid-stages, not ends","GenAI's benefit to SE students peaks mid-project, fades at edges","Study maps genAI aid to mid-steps for SE students, not first or last"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire four-phase pattern rests on students' retrospective self-reports, gathered in interviews and anchored on past conversation histories, about where genAI helped or hurt their learning and implementation; if those recollections are distorted, the benefit-challenge map and its cause-consequence claims are not established.","fun_headline_variants_meta":{"raw":{"variants":["GenAI aids SE students mid-task, not at first or last steps","Four-phase pattern: genAI helps only in middle steps of SE work","SE student interviews map genAI's help to mid-stages, not ends","GenAI's benefit to SE students peaks mid-project, fades at edges","Study maps genAI aid to mid-steps for SE students, not first or last"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000987,"raw_usage":{"total_tokens":4150,"prompt_tokens":873,"completion_tokens":3277,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":489,"completion_tokens_details":{"reasoning_tokens":3176}},"tokens_in":489,"tokens_out":3277,"duration_ms":20883,"temperature":1.0,"reasoning_tokens":3176,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T11:14:28.897199+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a controlled task-based study in which students are assigned representative SE tasks from each phase (L1, L2, I1, I2), with and without genAI assistance, and measure learning gains, task completion, time, and frustration. If students show genAI helping in L1 or I2, or hurting in L2 or I1, the claimed four-phase pattern is contradicted.","supporting_citations":[{"cited_title":"How Far Are We? The Triumphs and Trials of Generative AI in Learning Software Engineering,","cited_arxiv_id":null,"evidence_quote":"It supplies the closest prior study of ChatGPT in SE tasks, which this work extends beyond controlled tasks."},{"cited_title":"From” ban it till we understand it","cited_arxiv_id":null,"evidence_quote":"It contributes the instructor-interview framing and the technique of anchoring interviews in artifacts, which this study adopts."},{"cited_title":"Computing education in the era of generative AI,","cited_arxiv_id":null,"evidence_quote":"It supplies the broader debate on generative AI in computing education that this paper situates itself within."},{"cited_title":"Generative ai in computing education: Perspectives of students and instructors,","cited_arxiv_id":null,"evidence_quote":"It documents aligned student and instructor concerns about over-reliance and trustworthiness, which this paper refines into a phase map."},{"cited_title":"Using thematic analysis in psychology,","cited_arxiv_id":null,"evidence_quote":"It provides the reflexive thematic analysis method used to derive the codes and themes."}],"review_version":1}