Pith. sign in

REVIEW 2 major objections 5 minor 80 references

Insights from the Frontline: GenAI Utilization Among Software Engineering Students

T0 review · 2 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read GenAI helps SE students best at mid-steps, not first steps

desk verdict A transparent, small-scale interview study whose four-phase benefit/challenge map is a real contribution; the 'only' pattern is softer than it looks, but the paper deserves review. read the letter →

arxiv 2412.15624 v1 pith:WYF4YTE4 submitted 2024-12-20 cs.HC cs.SE

classification cs.HCcs.SE
keywords generativeAIsoftwareengineeringeducationstudentperceptionsbenefitsandchallengesthematicanalysishuman-AIinteractionqualitativeinterviewsliteracy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish where generative AI tools genuinely help software engineering students and where they get in the way, and why. Based on reflective interviews with 16 students, validated by member checking and instructor interviews, it argues that benefits concentrate in two specific phases: incremental learning of a concept students already partly know, and the initial steps of a concrete implementation. It argues that challenges concentrate in the other two phases: learning a concept from scratch and pushing an implementation to advanced, integrated work. The paper then traces those challenges back to intrinsic issues in the tools themselves (faults like hallucination and context neglect, gaps like missing scaffolding and weak debugging support) and forward to four kinds of impact on students: learning, task outcomes, self-perception, and adoption of the technology. A sympathetic reader would read this as a map of where genAI should and should not be leaned on in SE education, and why.

What carries the argument

The central object is a two-by-two phase model that splits SE coursework into initial learning (L1), incremental learning (L2), initial implementation (I1), and advanced implementation (I2). The analysis places perceived benefits and challenges into these quadrants and then builds a cause-consequence network (Figure 2) that connects genAI's intrinsic faults and gaps, through five challenge categories, to impacts on learning, task, self, and adoption. The phase model does the explanatory work: it shows why the same tool can be helpful in one context and harmful in another, and it turns scattered student complaints into a single testable pattern.

What would settle it

Run a controlled task-based study in which students are assigned representative SE tasks from each phase (L1, L2, I1, I2), with and without genAI assistance, and measure learning gains, task completion, time, and frustration. If students show genAI helping in L1 or I2, or hurting in L2 or I1, the claimed four-phase pattern is contradicted.

Watch

Extended reading notes

Core claim

The central claim is that students' lived experience with genAI clusters into a four-phase pattern: benefits appear in incremental learning (L2) and initial implementation (I1), while challenges appear in initial learning (L1) and advanced implementation (I2). The paper further claims that the challenges are not random difficulties but follow a causal chain: genAI's intrinsic faults (reasoning flaws, response-quality issues, deceptive behavior, neglect of student context) and gaps (scaffolding gaps, programming-support gaps) produce five challenge categories (C1-C5: unclear understanding of the tool, difficulty communicating needs, difficulty aligning AI to process and preferences, issues obtaining rationales, and difficulty using responses), which then produce four impacts (on learning, on task completion, on self-perception, and on willingness to adopt genAI). The pattern is meant to guide curriculum design: let students use genAI for clarification and initial scaffolding, but teach novices without it and prepare students for verification, prompt-crafting, and ethical judgment when work becomes advanced.

Load-bearing premise

The entire four-phase pattern rests on students' retrospective self-reports, gathered in interviews and anchored on past conversation histories, about where genAI helped or hurt their learning and implementation; if those recollections are distorted, the benefit-challenge map and its cause-consequence claims are not established.

Editorial extensions

If this is right

  • Curriculum designers can use the phase map to decide where genAI use should be encouraged, scaffolded, or restricted, rather than banning it outright.
  • Novice students need explicit instruction in prompt crafting, output verification, and adapting AI responses, because those are the skills that fail hardest in the challenging phases.
  • Assignments that require students to explain and justify AI-suggested solutions would directly target the reported lack of rationales (C4).
  • Improving the tools themselves, such as reducing deceptive behavior and adding debugging support, would shrink the challenge categories at L1 and I2.
  • Universities need clear authorship and ethical-use policies, since ethical uncertainty (C1) already steers some students away from using genAI in advanced work.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same four-phase pattern may extend beyond software engineering to other project-based disciplines, where initial concept acquisition and advanced integration are likely the points where AI assistance fails most.
  • Because the data are retrospective interviews, the paper maps where problems occur but not how often or how strongly; a quantitative survey or log-based study could attach frequencies to the five challenge categories.
  • The instructor triangulation revealed that instructors did not expect the emotional toll of genAI struggles, which suggests student support should address frustration and self-doubt, not just technical outcomes.
  • The pattern implies a sharper pedagogical rule than 'use it after mastering basics': genAI is safe for reinforcing known material and for jump-starting concrete tasks, but it is not a reliable tutor for first exposure, so educators should design for that asymmetry.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This paper reports a qualitative interview study of 16 software engineering (SE) students and two SE instructors about students' academic use of generative AI (genAI) tools. The authors identify four phases of use—initial learning (L1), incremental learning (L2), initial implementation (I1), and advanced implementation (I2)—and claim that participants perceived benefits only in L2 and I1, while challenges were encountered only in L1 and I2 (Section IV-C). They further analyze the causes of these challenges, attributing them to six intrinsic genAI issues (faults and gaps) that produce five challenge categories (C1–C5), which in turn impact learning, task outcomes, self-perception, and adoption of genAI (Section V). The findings are validated through member checking with students and triangulation with instructors. The authors explicitly acknowledge reliance on retrospective self-report, the absence of concrete tasks during the study, single-university sampling, and other threats to validity (Section VII).

Significance. If the central pattern holds, the paper offers an actionable map for SE educators deciding where genAI can be integrated into curricula and where students are likely to struggle. The study has notable strengths: interviews were anchored in participants' actual conversation histories with genAI tools, saturation was explicitly tested, the analysis used reflexive thematic analysis with consensus-based team meetings, and the authors performed member checking and instructor triangulation. The companion website with the codebook and supplemental material is a further positive. The principal risk is that the strong 'only' claim in Section IV-C rests on post-hoc phase coding of retrospective accounts, without inter-rater reliability evidence or a participant-level breakdown, so the clean diagonal pattern could be an artifact of the coding scheme rather than a robust property of the lived experiences.

major comments (2)
  1. [Section IV-C, Figure 1] The central claim that participants 'perceived the benefits of using genAI only for incremental learning (L2) and initial implementation tasks (I1)' and 'encountered challenges ... for initial learning (L1) and advanced implementations (I2)' is stronger than the evidence currently presented. Section III-B describes open coding with subsequent team negotiation, but no inter-rater reliability metric is reported, and the phase labels (L1, L2, I1, I2) were defined post hoc from the same interview data. Section VII acknowledges the absence of concrete tasks, meaning the phase attribution is entirely retrospective. Given these conditions, the exclusivity of the diagonal pattern could plausibly be an artifact of the coding scheme. Please provide a supplemental matrix showing which participants reported which benefit and challenge codes in which phases, or soften the 'only' formulations to 'clustered in' or 'were reported primarily in.'
  2. [Section VII] The acknowledged limitation that the study involved no concrete tasks means that phase attributions rely on participants' recollections. The member checking and instructor triangulation validate the presence of the benefit and challenge categories, but they cannot confirm the exclusivity of the phase mapping, because the mapping itself is an interpretive reconstruction. An independent re-coding of a subset of transcripts by a researcher not involved in the original analysis, with agreement statistics reported, would substantially strengthen the central claim. Without such evidence, the 'only' language in Section IV-C and the teaching recommendations built on it (Section VI) overstate the support.
minor comments (5)
  1. [Figure 1] Figure 1 contains serious rendering artifacts—stray question marks, duplicated text, and irregular line breaks—that obscure the phase-benefit/challenge mapping. Please provide a clean, publication-ready figure.
  2. [Section VII] The sentence 'we eschewed from discussing frequency or percentage of occurrences of categories' should read 'we eschewed discussing frequency or percentage of occurrences of categories.'
  3. [References] Reference [63] lists 'V . Clark' but the correct author name is 'V. Clarke'; please also check for inconsistent spacing in author initials across other references.
  4. [Section V, Figure 2] The arrows in Figure 2 from 'genAI's intrinsic issues' to challenges and impacts are based on participants' own causal attributions. The paper should clarify in the text that these are perceived associations, not verified causal links observed by the researchers, to avoid overstating the evidence.
  5. [Section III-B] The statement 'The first author transcribed the interviews' is useful for transparency, but the paper could also briefly describe the transcription accuracy check (if any) to align with standard reporting in qualitative studies.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the benefit/challenge pattern and cause-consequence taxonomy are emergent from fresh interview data, not equivalent to the paper's inputs or self-citations.

full rationale

The paper's central claims are descriptive qualitative results, not derived from any equation, fitted parameter, or prediction that is defined in terms of its own output. The L1/L2/I1/I2 phases are analytic context labels built from participants' reported uses (e.g., "Learning (Initial - L1) corresponded to situations where participants used genAI to learn SE concepts from scratch"), and the diagonal benefit/challenge map in Section IV-C is a coding outcome supported by direct participant quotes, not an identity. RQ2's causes/consequences taxonomy likewise emerges from open coding of the 16 interviews, with categories such as reasoning flaws and scaffolding gaps presented as classifications of participants' descriptions rather than as quantities derived from those categories. Self-citations appear ([19], [16], [76]), but they are contextual: [19] is described as "Closest to our work" and contrasted with the present study's unrestricted scope, and [76] is cited only when recommending that educators help students "scope their trust in AI." None of these citations carries the burden of establishing the empirical pattern. The paper itself flags in Section VII "the absence of concrete tasks conducted by the students during the study, limiting our results to our participants' recollections," and the analysis uses negotiated agreement rather than an inter-rater reliability metric; these are methodological and evidentiary limitations, not cases where a result reduces to its input by construction. No uniqueness theorem, no imported ansatz, and no renaming of a known result is load-bearing. The verdict is no significant circularity; score 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No free parameters were fitted. The central claims depend on four qualitative assumptions: accuracy of self-report, sufficiency of saturation, validity of member checking as confirmation, and appropriateness of external theory lenses. These are reasonable for exploratory qualitative work but limit the strength of the conclusions.

assumptions (4)
  • domain assumption Students' retrospective recollections, anchored on past conversation logs, accurately capture where genAI helped or hurt their learning and implementation.
    The RQ1 phase-benefit/challenge map is built on self-report; the paper acknowledges in Section VII that there was no concrete task observation and results are limited to recollections.
  • domain assumption Thematic saturation was reached at 10 interviews and the additional 6 confirm it, making 16 interviews sufficient for the emergent categories.
    Section III-B: saturation is asserted on the basis of no new insights, which is standard but subjective in qualitative research.
  • domain assumption Member checking with 14 of 16 participants and triangulation with instructors provide sufficient validation of the interpreted findings.
    Section III-B: these checks confirm plausibility and did not produce disagreement, but they cannot confirm objective truth about learning outcomes.
  • domain assumption Post-hoc application of Cognitive Load Theory, Social Cognitive Theory, Self-Determination Theory, and the Technology Acceptance Model is an appropriate way to explain the consequences.
    Section V-B: these external theories are used interpretively, not derived from the data.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Insights from the Frontline: GenAI Utilization Among Software Engineering Students." pith.science (2026). https://pith.science/paper/WYF4YTE4

@misc{pith2026241215624,
  author       = {Pith},
  title        = {Pith review of: Insights from the Frontline: GenAI Utilization Among Software Engineering Students},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WYF4YTE4}},
  note         = {Machine review of arXiv:2412.15624}
}
read the original abstract

Generative AI (genAI) tools (e.g., ChatGPT, Copilot) have become ubiquitous in software engineering (SE). As SE educators, it behooves us to understand the consequences of genAI usage among SE students and to create a holistic view of where these tools can be successfully used. Through 16 reflective interviews with SE students, we explored their academic experiences of using genAI tools to complement SE learning and implementations. We uncover the contexts where these tools are helpful and where they pose challenges, along with examining why these challenges arise and how they impact students. We validated our findings through member checking and triangulation with instructors. Our findings provide practical considerations of where and why genAI should (not) be used in the context of supporting SE students.

Figures

Figures reproduced from arXiv: 2412.15624 by the authors.

Figure 1
Figure 1. Perceived benefits and challenges of genAI usage among students: Benefits were in incremental learning (L2) & initial implementations (I1), with challenges in initial learning (L1) & advanced implementations (I2). Stars indicate perceptions unique to the quadrant. 4) Using course materials to query AI enabled participants to ask focused questions to refine or expand on their knowl￾edge. For instance, P4 described, “… view at source ↗
Figure 2
Figure 2. Associations between genAI’s intrinsic issues (faults and gaps), challenges (C1-C5), and the resulting impacts. Issues in genAI contributed to various challenges encountered by participants, which subsequently had impacts on learning, task, self, and genAI adoption. wrong information. P1 reflected on this, stating, “Sometimes, genAI gives wrong information due to mistakes in how I’ve inquired as I didn’t know enough… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

80 extracted references · 63 canonical work pages

  1. [1]

    OpenAI, “Gpt-4,” https://openai.com/product/gpt-4, 2024

  2. [2]

    Google, “Gemini,” https://gemini.google.com, 2024

  3. [3]

    Copilot,

    Microsoft, “Copilot,” https://copilot.microsoft.com, 2024

  4. [4]

    Large language models for software engineering: Sur- vey and open problems,

    A. Fan, B. Gokkaya, M. Harman, M. Lyubarskiy, S. Sengupta, S. Yoo, and J. M. Zhang, “Large language models for software engineering: Sur- vey and open problems,” in 2023 IEEE/ACM International Conference on Software Engineering: Future of Software Engineering (ICSE-FoSE). IEEE, 2023, pp. 31–53

  5. [5]

    Generative AI for software practitioners,

    C. Ebert and P. Louridas, “Generative AI for software practitioners,” IEEE Software, vol. 40, no. 4, pp. 30–38, 2023

  6. [6]

    Navigating the complexity of generative AI adoption in software engineering,

    D. Russo, “Navigating the complexity of generative AI adoption in software engineering,” ACM Transactions on Software Engineering and Methodology, 2024

  7. [7]

    Generative AI assistants in software develop- ment education: A vision for integrating generative AI into educational practice, not instinctively defending against it

    C. Bull and A. Kharrufa, “Generative AI assistants in software develop- ment education: A vision for integrating generative AI into educational practice, not instinctively defending against it.” IEEE Software, 2023

  8. [8]

    Generative artificial intelligence for software engineering–a research agenda,

    A. Nguyen-Duc, B. Cabrero-Daniel, A. Przybylek, C. Arora, D. Khanna, T. Herda, U. Rafiq, J. Melegati, E. Guerra, K.-K. Kemell et al. , “Generative artificial intelligence for software engineering–a research agenda,” arXiv preprint arXiv:2310.18648 , 2023

Show all 80 references
  1. [9]

    Computing education in the era of generative AI,

    P. Denny, J. Prather, B. A. Becker, J. Finnie-Ansley, A. Hellas, J. Leinonen, A. Luxton-Reilly, B. N. Reeves, E. A. Santos, and S. Sarsa, “Computing education in the era of generative AI,” Communications of the ACM, vol. 67, no. 2, pp. 56–67, 2024

  2. [10]

    The end of programming,

    M. Welsh, “The end of programming,” Communications of the ACM , vol. 66, no. 1, pp. 34–35, 2022

  3. [11]

    The premature obituary of programming,

    D. M. Yellin, “The premature obituary of programming,” Communica- tions of the ACM , vol. 66, no. 2, pp. 41–44, 2023

  4. [12]

    Evaluating large language models trained on code,

    M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y . Burda, N. Joseph, G. Brockman et al., “Evaluating large language models trained on code,” arXiv preprint arXiv:2107.03374 , 2021

  5. [13]

    On the opportunities and risks of foundation models,

    R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill et al. , “On the opportunities and risks of foundation models,” arXiv preprint arXiv:2108.07258, 2021

  6. [14]

    Towards characterizing trust in gen- erative artificial intelligence among students,

    M. Amoozadeh, D. Daniels, S. Chen, D. Nam, A. Kumar, M. Hilton, M. A. Alipour, and S. S. Ragavan, “Towards characterizing trust in gen- erative artificial intelligence among students,” in 2023 ACM Conference on International Computing Education Research-Volume 2 , 2023, pp. 3–4

  7. [15]

    Programming is hard-or at least it used to be: Educational opportunities and challenges of AI code generation,

    B. A. Becker, P. Denny, J. Finnie-Ansley, A. Luxton-Reilly, J. Prather, and E. A. Santos, “Programming is hard-or at least it used to be: Educational opportunities and challenges of AI code generation,” in Proceedings of the 54th ACM Technical Symposium on Computer Science Edu...

  8. [16]

    An- ticipating user needs: Insights from design fiction on conversational agents for computational thinking,

    J. Penney, J. F. Pimentel, I. Steinmacher, and M. A. Gerosa, “An- ticipating user needs: Insights from design fiction on conversational agents for computational thinking,” in Chatbot Research and Design , A. Følstad, T. Araujo, S. Papadopoulos, E. L.-C. Law, E. Luger, M. Goodw...

  9. [17]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017

  10. [18]

    Gpt-4 technical report,

    J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat et al. , “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774 , 2023

  11. [19]

    How Far Are We? The Triumphs and Trials of Generative AI in Learning Software Engineering,

    R. Choudhuri, D. Liu, I. Steinmacher, M. Gerosa, and A. Sarma, “How Far Are We? The Triumphs and Trials of Generative AI in Learning Software Engineering,” in Proceedings of the IEEE/ACM 46th International Conference on Software Engineering , 2024, pp. 1–13

  12. [20]

    More than calculators: Why large language models threaten public education,

    A. J. Ko, “More than calculators: Why large language models threaten public education,” Jan 2024

  13. [21]

    From” ban it till we understand it

    S. Lau and P. Guo, “From” ban it till we understand it” to” resistance is futile”: How university programming instructors plan to adapt as more students use AI code generation and explanation tools such as chatgpt and github copilot,” in 2023 ACM Conference on International Co...

  14. [22]

    A large-scale survey on the usability of AI programming assistants: Successes and challenges,

    J. T. Liang, C. Yang, and B. A. Myers, “A large-scale survey on the usability of AI programming assistants: Successes and challenges,” in Proceedings of the 46th IEEE/ACM International Conference on Software Engineering, 2024, pp. 1–13

  15. [23]

    How AI is helping to identify skills gaps and future jobs,

    K. Whiting, “How AI is helping to identify skills gaps and future jobs,” May 2023. [Online]. Available: https://www.weforum.org/agenda/2023/ 05/ai-skills-gaps-future-jobs/

  16. [24]

    How knowledge workers think generative AI will (not) transform their industries,

    A. Woodruff, R. Shelby, P. G. Kelley, S. Rousso-Schindler, J. Smith- Loud, and L. Wilcox, “How knowledge workers think generative AI will (not) transform their industries,” in Proceedings of the CHI Conference on Human Factors in Computing Systems , 2024, pp. 1–26

  17. [25]

    The robots are here: Navigating the generative AI revolution in computing education,

    J. Prather, P. Denny, J. Leinonen, B. A. Becker, I. Albluwi, M. Craig, H. Keuning, N. Kiesler, T. Kohn, A. Luxton-Reillyet al., “The robots are here: Navigating the generative AI revolution in computing education,” in Proceedings of the 2023 Working Group Reports on Innovation...

  18. [26]

    “so what if chatgpt wrote it?

    T. Malik, Y . Dwivedi, N. Kshetri, L. Hughes, E. L. Slade, A. Jeyaraj, A. K. Kar, A. M. Baabdullah, A. Koohang, V . Raghavanet al., ““so what if chatgpt wrote it?” multidisciplinary perspectives on opportunities, challenges and implications of generative conversational ai for ...

  19. [27]

    Exploring chatgpt’s impact on post-secondary education: A qualitative study,

    P. Rajabi, P. Taghipour, D. Cukierman, and T. Doleck, “Exploring chatgpt’s impact on post-secondary education: A qualitative study,” in Proceedings of the 25th Western Canadian Conference on Computing Education, 2023, pp. 1–6

  20. [28]

    From” let’s google

    I. Joshi, R. Budhiraja, P. D. Tanna, L. Jain, M. Deshpande, A. Srivastava, S. Rallapalli, H. D. Akolekar, J. S. Challa, and D. Kumar, “From” let’s google” to” let’s chatgpt”: Student and instructor perspectives on the influence of llms on undergraduate engineering education,” ...

  21. [29]

    University students as early adopters of chatgpt: Innovation diffusion study,

    R. Raman, S. Mandal, P. Das, T. Kaur, J. Sanjanasri, and P. Nedungadi, “University students as early adopters of chatgpt: Innovation diffusion study,” 2023

  22. [30]

    Towards adapting computer science courses to AI assistants’ capabilities,

    T. Wang, D. Vargas-Diaz, C. Brown, and Y . Chen, “Towards adapting computer science courses to AI assistants’ capabilities,” arXiv preprint arXiv:2306.03289, 2023

  23. [31]

    Generative ai in computing education: Perspectives of students and instructors,

    C. Zastudil, M. Rogalska, C. Kapp, J. Vaughn, and S. MacNeil, “Generative ai in computing education: Perspectives of students and instructors,” in 2023 IEEE Frontiers in Education Conference (FIE) . IEEE, 2023, pp. 1–9

  24. [32]

    Grounded copilot: How programmers interact with code-generating models,

    S. Barke, M. B. James, and N. Polikarpova, “Grounded copilot: How programmers interact with code-generating models,” Proceedings of the ACM on Programming Languages , vol. 7, no. OOPSLA1, pp. 85–111, 2023

  25. [33]

    Studying the effect of AI code generators on supporting novice learners in introductory programming,

    M. Kazemitabaar, J. Chow, C. K. T. Ma, B. J. Ericson, D. Weintrop, and T. Grossman, “Studying the effect of AI code generators on supporting novice learners in introductory programming,” in Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems , 2023, pp. 1–23

  26. [34]

    Reading between the lines: Modeling user behavior and costs in ai-assisted programming,

    H. Mozannar, G. Bansal, A. Fourney, and E. Horvitz, “Reading between the lines: Modeling user behavior and costs in ai-assisted programming,” arXiv preprint arXiv:2210.14306 , 2022

  27. [35]

    The programmer’s assistant: Conversational interaction with a large language model for software development,

    S. I. Ross, F. Martinez, S. Houde, M. Muller, and J. D. Weisz, “The programmer’s assistant: Conversational interaction with a large language model for software development,” in Proceedings of the 28th International Conference on Intelligent User Interfaces , 2023, pp. 491– 514

  28. [36]

    Expectation vs. experi- ence: Evaluating the usability of code generation tools powered by large language models,

    P. Vaithilingam, T. Zhang, and E. L. Glassman, “Expectation vs. experi- ence: Evaluating the usability of code generation tools powered by large language models,” in Chi conference on human factors in computing systems extended abstracts , 2022, pp. 1–7

  29. [37]

    In-ide code generation from natural language: Promise and challenges,

    F. F. Xu, B. Vasilescu, and G. Neubig, “In-ide code generation from natural language: Promise and challenges,” ACM Transactions on Soft- ware Engineering and Methodology (TOSEM) , vol. 31, no. 2, pp. 1–47, 2022

  30. [38]

    Productivity assessment of neural code completion,

    A. Ziegler, E. Kalliamvakou, X. A. Li, A. Rice, D. Rifkin, S. Simister, G. Sittampalam, and E. Aftandilian, “Productivity assessment of neural code completion,” in Proceedings of the 6th ACM SIGPLAN Interna- tional Symposium on Machine Programming , 2022, pp. 21–29

  31. [39]

    “it’s weird that it knows what i want

    J. Prather, B. N. Reeves, P. Denny, B. A. Becker, J. Leinonen, A. Luxton- Reilly, G. Powell, J. Finnie-Ansley, and E. A. Santos, ““it’s weird that it knows what i want”: Usability and interactions with copilot for novice programmers,” ACM Transactions on Computer-Human Interac...

  32. [40]

    Taking flight with copilot,

    C. Bird, D. Ford, T. Zimmermann, N. Forsgren, E. Kalliamvakou, T. Lowdermilk, and I. Gazit, “Taking flight with copilot,” Communi- cations of the ACM , vol. 66, no. 6, pp. 56–62, 2023

  33. [41]

    Guidelines for human- AI interaction,

    S. Amershi, D. Weld, M. V orvoreanu, A. Fourney, B. Nushi, P. Collisson, J. Suh, S. Iqbal, P. N. Bennett, K. Inkpen et al., “Guidelines for human- AI interaction,” in Proceedings of the 2019 chi conference on human factors in computing systems , 2019, pp. 1–13

  34. [42]

    The widening gap: The benefits and harms of generative AI for novice programmers,

    J. Prather, B. N. Reeves, J. Leinonen, S. MacNeil, A. S. Randrianasolo, B. A. Becker, B. Kimmel, J. Wright, and B. Briggs, “The widening gap: The benefits and harms of generative AI for novice programmers,” in Proceedings of the 2024 ACM Conference on International Computing E...

  35. [43]

    An investigation of the drivers of novice programmers’ intentions to use web search and GenAI,

    J. Skripchuk, J. Bacher, and T. Price, “An investigation of the drivers of novice programmers’ intentions to use web search and GenAI,” in Proceedings of the 2024 ACM Conference on International Computing Education Research-Volume 1, 2024, pp. 487–501

  36. [44]

    AI-driven development is here: Should you worry?

    N. A. Ernst and G. Bavota, “AI-driven development is here: Should you worry?” IEEE Software, vol. 39, no. 2, pp. 106–110, 2022

  37. [45]

    Com- piler error messages considered unhelpful: The landscape of text-based programming error message research,

    B. A. Becker, P. Denny, R. Pettit, D. Bouchard, D. J. Bouvier, B. Har- rington, A. Kamil, A. Karkare, C. McDonald, P.-M. Osera et al., “Com- piler error messages considered unhelpful: The landscape of text-based programming error message research,” Proceedings of the working g...

  38. [46]

    The robots are coming: Exploring the implications of openai codex on introductory programming,

    J. Finnie-Ansley, P. Denny, B. A. Becker, A. Luxton-Reilly, and J. Prather, “The robots are coming: Exploring the implications of openai codex on introductory programming,” in Proceedings of the 24th Australasian Computing Education Conference , 2022, pp. 10–19

  39. [47]

    Automatic generation of programming exercises and code explanations using large language models,

    S. Sarsa, P. Denny, A. Hellas, and J. Leinonen, “Automatic generation of programming exercises and code explanations using large language models,” in Proceedings of the 2022 ACM Conference on International Computing Education Research-Volume 1 , 2022, pp. 27–43

  40. [48]

    Can genera- tive pre-trained transformers (gpt) pass assessments in higher education programming courses?

    J. Savelka, A. Agarwal, C. Bogart, Y . Song, and M. Sakr, “Can genera- tive pre-trained transformers (gpt) pass assessments in higher education programming courses?” in Proceedings of the 2023 Conference on Innovation and Technology in Computer Science Education V . 1 , 2023, ...

  41. [49]

    Comparing code explanations created by students and large language models,

    J. Leinonen, P. Denny, S. MacNeil, S. Sarsa, S. Bernstein, J. Kim, A. Tran, and A. Hellas, “Comparing code explanations created by students and large language models,” in Proceedings of the 2023 Con- ference on Innovation and Technology in Computer Science Education V . 1, 202...

  42. [50]

    Experiences from using code explanations generated by large language models in a web software development e-book,

    S. MacNeil, A. Tran, A. Hellas, J. Kim, S. Sarsa, P. Denny, S. Bernstein, and J. Leinonen, “Experiences from using code explanations generated by large language models in a web software development e-book,” in Proceedings of the 54th ACM Technical Symposium on Computer Science...

  43. [51]

    Using GitHub copilot to solve simple programming problems,

    M. Wermelinger, “Using GitHub copilot to solve simple programming problems,” in Proceedings of the 54th ACM Technical Symposium on Computer Science Education V . 1, 2023, pp. 172–178

  44. [52]

    Using large language models to enhance programming error messages,

    J. Leinonen, A. Hellas, S. Sarsa, B. Reeves, P. Denny, J. Prather, and B. A. Becker, “Using large language models to enhance programming error messages,” in Proceedings of the 54th ACM Technical Symposium on Computer Science Education V . 1, 2023, pp. 563–569

  45. [53]

    The AI generation gap: Are Gen Z students more interested in adopting generative AI such as chatgpt in teaching and learning than their Gen X and millennial generation teachers?

    C. K. Y . Chan and K. K. Lee, “The AI generation gap: Are Gen Z students more interested in adopting generative AI such as chatgpt in teaching and learning than their Gen X and millennial generation teachers?” Smart Learning Environments, vol. 10, no. 1, p. 60, 2023

  46. [54]

    Insights from social shaping theory: The appropriation of large language models in an undergraduate programming course,

    A. Padiyath, X. Hou, A. Pang, D. Viramontes Vargas, X. Gu, T. Nelson- Fromm, Z. Wu, M. Guzdial, and B. Ericson, “Insights from social shaping theory: The appropriation of large language models in an undergraduate programming course,” in Proceedings of the 2024 ACM Conference o...

  47. [55]

    Chatgpt in data visualiza- tion education: A student perspective,

    N. W. Kim, H.-K. Ko, G. Myers, and B. Bach, “Chatgpt in data visualiza- tion education: A student perspective,”arXiv preprint arXiv:2405.00748, 2024

  48. [56]

    Supplemental Material for GenAI Utilization Among SE Students,

    Anonymous, “Supplemental Material for GenAI Utilization Among SE Students,” Anonymous, October 2024. [Online]. Available: https://doi.org/10.5281/zenodo.13882829

  49. [57]

    A revision of bloom’s taxonomy: An overview,

    D. R. Krathwohl, “A revision of bloom’s taxonomy: An overview,” Theory into practice , vol. 41, no. 4, pp. 212–218, 2002

  50. [58]

    Bloom’s taxonomy: Original and revised,

    M. Forehand et al., “Bloom’s taxonomy: Original and revised,” Emerg- ing perspectives on learning, teaching, and technology, vol. 8, pp. 41–44, 2005

  51. [59]

    Using software engineering design principles as tools for freshman students learning,

    I. Cabezas, R. Segovia, P. Caratozzolo, and E. Webb, “Using software engineering design principles as tools for freshman students learning,” in 2020 IEEE Frontiers in Education Conference (FIE) . IEEE, 2020, pp. 1–5

  52. [60]

    The mind in the middle,

    J. A. Bargh and T. L. Chartrand, “The mind in the middle,” Handbook of research methods in social and personality psychology , vol. 2, pp. 253–285, 2000

  53. [61]

    A literature review of the anchoring effect,

    A. Furnham and H. C. Boo, “A literature review of the anchoring effect,” The journal of socio-economics , vol. 40, no. 1, pp. 35–42, 2011

  54. [62]

    J. W. Creswell and C. N. Poth, Qualitative inquiry and research design: Choosing among five approaches . Sage publications, 2016

  55. [63]

    Using thematic analysis in psychology,

    V . Braun and V . Clark, “Using thematic analysis in psychology,” Qualitative research in psychology , vol. 3, no. 2, pp. 77–101, 2006

  56. [64]

    Conceptual and design thinking for thematic analysis

    V . Braun and V . Clarke, “Conceptual and design thinking for thematic analysis.” Qualitative Psychology, vol. 9, no. 1, p. 3, 2022

  57. [65]

    Atlas.ti (version 3) [computer software],

    “Atlas.ti (version 3) [computer software],” 2024, accessed: 2024-09-07. [Online]. Available: https://atlasti.com/

  58. [66]

    Reliability and inter-rater reliability in qualitative research: Norms and guidelines for CSCW and HCI practice,

    N. McDonald, S. Schoenebeck, and A. Forte, “Reliability and inter-rater reliability in qualitative research: Norms and guidelines for CSCW and HCI practice,” Proceedings of the ACM on human-computer interaction, vol. 3, no. CSCW, pp. 1–23, 2019

  59. [67]

    What is an adequate sample size? operationalising data saturation for theory-based interview studies,

    J. J. Francis, M. Johnston, C. Robertson, L. Glidewell, V . Entwistle, M. P. Eccles, and J. M. Grimshaw, “What is an adequate sample size? operationalising data saturation for theory-based interview studies,” Psychology and health , vol. 25, no. 10, pp. 1229–1245, 2010

  60. [68]

    Evaluating errors and improving performance of chatgpt,

    S. Biswas, “Evaluating errors and improving performance of chatgpt,” International Journal of Clinical and Medical Education Research , vol. 2, no. 6, pp. 182–188, 2023

  61. [69]

    Cognitive load theory,

    J. L. Plass, R. Moreno, and R. Br ¨unken, “Cognitive load theory,” 2010

  62. [70]

    Social foundations of thought and action,

    A. Bandura et al. , “Social foundations of thought and action,” Engle- wood Cliffs, NJ , vol. 1986, no. 23-28, p. 2, 1986

  63. [71]

    Self-determination theory,

    E. L. Deci and R. M. Ryan, “Self-determination theory,” Handbook of theories of social psychology , vol. 1, no. 20, pp. 416–436, 2012

  64. [72]

    Technology acceptance model,

    F. D. Davis, R. Bagozzi, and P. Warshaw, “Technology acceptance model,” J Manag Sci , vol. 35, no. 8, pp. 982–1003, 1989

  65. [73]

    AI Guidance & FAQs,

    Harvard Office of Undergraduate Education, “AI Guidance & FAQs,” https://oue.fas.harvard.edu/ai-guidance-faqs, 2023

  66. [74]

    Guidance for the use of Generative AI,

    UCLA Center for the Advancement of Teaching, “Guidance for the use of Generative AI,” https://teaching.ucla.edu/resources/ guidance-for-the-use-of-generative-ai/, 2023

  67. [75]

    Trust in Generative AI among students: An exploratory study,

    M. Amoozadeh, D. Daniels, D. Nam, A. Kumar, S. Chen, M. Hilton, S. Srinivasa Ragavan, and M. A. Alipour, “Trust in Generative AI among students: An exploratory study,” in 55th ACM Technical Symposium on Computer Science Education V . 1, 2024, pp. 67–73

  68. [76]

    What Guides Our Choices? Modeling Developers’ Trust and Behavioral Intentions To- wards GenAI,

    R. Choudhuri, B. Trinkenreich, R. Pandita, E. Kalliamvakou, I. Stein- macher, M. Gerosa, C. Sanchez, and A. Sarma, “What Guides Our Choices? Modeling Developers’ Trust and Behavioral Intentions To- wards GenAI,” arXiv preprint arXiv:2409.04099 , 2024

  69. [77]

    Taking flight with copilot: Early in- sights and opportunities of ai-powered pair-programming tools,

    C. Bird, D. Ford, T. Zimmermann, N. Forsgren, E. Kalliamvakou, T. Lowdermilk, and I. Gazit, “Taking flight with copilot: Early in- sights and opportunities of ai-powered pair-programming tools,” Queue, vol. 20, no. 6, pp. 35–57, 2022

  70. [78]

    A prompt pattern catalog to enhance prompt engineering with chatgpt,

    J. White, Q. Fu, S. Hays, M. Sandborn, C. Olea, H. Gilbert, A. El- nashar, J. Spencer-Smith, and D. C. Schmidt, “A prompt pattern catalog to enhance prompt engineering with chatgpt,” arXiv preprint arXiv:2302.11382, 2023

  71. [79]

    S. B. Merriam and E. J. Tisdell, Qualitative research: A guide to design and implementation. John Wiley & Sons, 2015

  72. [80]

    N. K. Denzin and Y . S. Lincoln, The Sage handbook of qualitative research. sage, 2011

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.