REVIEW 3 major objections 9 minor 85 references
OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning
T0 review · 3 major / 9 minor · reviewed 2026-07-05 · glm-5.2
Pith's one-line read AI pair programmer boosts speed but flattens emotion and learning
desk verdict Well-designed study of human vs. AI pair programming finds AI boosts performance but humans boost emotion; sample size limits retention claims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The experiment uses a within-subjects counterbalanced design where each participant programs in both conditions (human teammate and Copilot) across two task pools drawn from HumanEval, with a one-week retention interval followed by a solo retest. The analytical machinery is Ordinary Least Squares regression with Wild Cluster Bootstrapping to handle the small sample and nested data structure (participants within pairs, repeated measures). The theoretical interpretive frame is Pekrun's Control-Value Theory, which predicts that emotions follow from appraisals of both control over an activity and the value of its outcome, explaining why Copilot's high competence raises outcome value while lowing
What would settle it
A larger replication with more participants and independently validated task difficulty equivalence that finds no retention difference between conditions, or that finds the performance gap disappears when task pools are more rigorously matched.
Extended reading notes
Core claim
The paper's central discovery is that AI-assisted programming and human pair programming produce opposite profiles across four measured dimensions. Copilot yielded higher task performance (approximately 14 points on a 100-point scale, a large effect) and lower mental demand, temporal demand, and effort (all large effects), but the human condition produced significantly more positive emotional valence and higher arousal (both large effects). Retention performance one week later was statistically comparable in absolute terms between conditions, but the drop from initial team performance to solo retest was significantly larger for tasks learned with Copilot, suggesting that the higher initialAI
Load-bearing premise
The paper assumes that the two sets of programming tasks (Pool X and Pool Y) are equivalent in difficulty, and with only 22 participants and 11 pairs, the study has limited statistical power to detect a task-pool-by-condition interaction that could confound the results.
Editorial extensions
If this is right
- Educators who replace pair programming with AI-assisted programming may gain short-term productivity but risk eroding the emotional engagement and social learning dynamics that support deeper processing.
- AI tool designers could explore rate-limiting or scaffolding features that withhold full solutions, forcing the learner to maintain cognitive control and potentially preserving both engagement and retention.
- The finding that stronger teammates learned less when their first exposure was through Copilot suggests that the act of teaching a peer may itself be a learning mechanism that AI assistance bypasses.
- The emotional flatness associated with Copilot, if generalizable, could have cumulative effects on motivation and persistence in programming courses where students increasingly default to AI tools.
Reading between the lines
- If the emotional and retention differences hold in larger samples, the result would imply an inverse relationship between immediate task performance and durable learning in AI-assisted contexts, a pattern reminiscent of 'desirable difficulties' in cognitive psychology.
- The marginally significant retention effects, combined with the small sample, suggest that a larger replication could either confirm a genuine learning cost of AI assistance or reveal that the effect is unstable; either outcome would be informative.
- The finding that participants did not credit themselves for Copilot-enhanced performance raises the possibility that AI tools may undermine self-efficacy formation in ways not captured by standard self-report measures.
- If task difficulty equivalence between the two pools is not perfect, some portion of the performance gap could reflect task-Copilot interaction rather than the teammate condition alone, though the paper's Wald tests found no significant confound.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript reports a within-subjects controlled experiment (n=22, 11 pairs) comparing human-human pair programming to human-AI pair programming (GitHub Copilot) among novice and intermediate programmers. The study measures four outcomes: programming performance (RQ1), learning retention after one week (RQ2), subjective workload via NASA-TLX (RQ3), and emotional impact via valence/arousal change scores (RQ4). The central findings are that AI assistance significantly improves performance and reduces workload but diminishes emotional engagement compared to human pair programming. Learning retention differences were mostly non-significant, with only a marginally significant relative performance decrement in the AI condition. The study is grounded in Control-Value Theory of emotion and Cognitive Load Theory. The experimental design is careful: counterbalancing across eight permutations, deception with debriefing, pre-registered-style task pools, and Wild Cluster Bootstrapping for inference given the small clustered sample. The paper also provides qualitative triangulation from exit questionnaires. Note: the abstract and arXiv metadata describe an RL/LLM paper (OGER), but the full text is an HCI/computing education paper; this review addresses the actual full-text manuscript.
Significance. The paper makes a timely contribution to computing education research by directly comparing human-human and human-AI pair programming on a multi-dimensional outcome set (performance, learning, workload, emotion) within the same participants. The emotion finding (RQ4) is the most distinctive contribution, as prior work has focused primarily on performance and workload. The within-subjects design with counterbalancing, the use of WCB inference appropriate for few-cluster data, and the BH correction within RQ families are methodological strengths. The authors provide open appendices and task materials via Zenodo, supporting reproducibility. The theoretical framing via CVT is well-motivated and the qualitative data triangulates the quantitative findings effectively.
major comments (3)
- §4.4, Table 4: The emotion findings (valence change B=-1.18, g=-1.61; arousal change B=-2.09, g=-1.74) are the paper's most novel contribution, but they rely on sequential within-subjects change scores where differential carryover could inflate the apparent AI-vs-human gap. If the human condition produces a strong positive emotional response (which is the finding), participants who experience it first carry an elevated baseline into the subsequent AI trial, making the AI condition's change score artificially negative. This is an asymmetric carryover effect that counterbalancing cannot fully address. The paper acknowledges this in §6 as a caveat, but the Wald test for ordering effects (p>.1, §3.7) has very low power at 11 clusters to detect such an interaction. The paper should either (a) report the emotion results stratified by ordering (human-first vs. AI-first) so readers can assess to
- §3.3: The assumption that the two task pools (Pool X: tasks 14, 52, 72, 0; Pool Y: tasks 35, 9, 24, 18) are equivalent in difficulty is load-bearing for all four RQs. The paper states pools were 'as similar as possible in difficulty' and reports Wald tests (p>.1), but with only 11 pairs, the power to detect a task-pool-by-condition interaction is very low. If one pool happened to be more amenable to Copilot's strengths (e.g., tasks requiring more boilerplate vs. algorithmic reasoning), the performance difference could be partially attributable to task selection. The paper should report descriptive statistics (mean scores) for each pool separately by condition, or at minimum acknowledge this power limitation more explicitly in §6.
- §4.2, Table 2: The learning retention analysis (RQ2) reports a marginally significant relative retest decrement (p_adj=.054, g=-1.13) but non-significant absolute retest performance (p_adj=.529, g=-0.27). The large discrepancy between the relative and absolute effect sizes (g=-1.13 vs. g=-0.27) is driven entirely by the higher Session 1 performance with AI. This means the 'large' effect for relative retest change is a mathematical consequence of the RQ1 performance finding, not independent evidence of a learning difference. The paper should clarify that the relative retest result does not provide independent evidence of differential learning beyond what RQ1 already establishes.
minor comments (9)
- The arXiv abstract and metadata describe a completely different paper (OGER: Offline-Guided Exploration Reward for RL). The title, abstract, and metadata should be corrected to match the actual manuscript content.
- §3.7: 'Benjimini-Hochberg' is misspelled (should be 'Benjamini-Hochberg') in multiple places, including Tables 1-4 captions and the main text.
- §3.6.1: The emotion measure uses 7-point Likert scales for valence and arousal, but the paper does not specify the exact anchor labels. Including these would help readers interpret the magnitude of the changes.
- Fig. 2a/2b: The figure captions reference subfigures (a) and (b) but the timeline details are small and difficult to read. Consider enlarging or providing a text-based timeline in addition to the figure.
- §4.1: The claim that the performance advantage corresponds to '7/10 of a task or the same number of tasks but faster by a margin exceeding half of the total time limit' is somewhat informal. A more precise statement of the effect in terms of the scoring formula (Eq. 2) would be clearer.
- §5.2: The discussion of 'near-peer teaching' as an explanation for stronger teammates' lower AI-condition retest performance is speculative. The paper should either flag this more clearly as post-hoc or provide additional supporting evidence.
- §3.1: The demographic skew (16 male, 15 Asian out of 22) is noted but not discussed as a potential limitation on generalizability. A brief mention in §6 would be appropriate.
- Table 3 caption states 'Correcting p-values for multiple testing would not change the conclusions here, even with the very conservative Bonferroni adjustment.' This is an editorial claim that should be verified and, if accurate, stated more precisely.
- §4.4: The paper notes that 'valence and arousal simply tended to move together.' A correlation between valence and arousal change scores could be reported to quantify this.
Simulated Author's Rebuttal
We thank the referee for a careful and constructive review. The referee raises three substantive points concerning (1) potential asymmetric carryover in the emotion change scores, (2) the equivalence of the two task pools, and (3) the interpretation of the relative retest decrement. We agree with all three points and will revise the manuscript accordingly. Below we address each in turn.
read point-by-point responses
-
Referee: §4.4, Table 4: The emotion findings rely on sequential within-subjects change scores where differential carryover could inflate the apparent AI-vs-human gap. The paper should report emotion results stratified by ordering (human-first vs. AI-first) or otherwise address this concern.
Authors: The referee is correct that asymmetric carryover is a genuine threat to the emotion change-score analysis, and we agree that the Wald test for ordering effects is underpowered at 11 clusters to detect such an interaction. We will address this in revision by reporting the emotion results (valence change and arousal change) stratified by ordering group (human-first vs. AI-first) in a supplementary table, so readers can directly assess whether the pattern is consistent across orderings. We will also expand the caveat in §6 to explicitly state that the low power of the ordering test means it cannot rule out asymmetric carryover, and that the stratified results should be interpreted accordingly. We want to be transparent: if the stratified analysis reveals that the effect is concentrated in one ordering group, this would weaken the causal interpretation of the emotion finding, and we will say so plainly. revision: partial
-
Referee: §3.3: The assumption that the two task pools are equivalent in difficulty is load-bearing for all four RQs. The paper should report descriptive statistics for each pool separately by condition, or at minimum acknowledge this power limitation more explicitly in §6.
Authors: We agree that task-pool equivalence is a critical assumption and that the Wald test is underpowered to detect a task-pool-by-condition interaction at this sample size. We will add a supplementary table reporting mean scores (performance, retest, workload subscales, and emotion change) for each pool separately by condition. This will allow readers to assess whether any pool appears systematically more or less amenable to AI assistance. We will also strengthen the discussion in §6 to explicitly acknowledge that the non-significant Wald test does not constitute strong evidence of pool equivalence given the limited power, and that the task-pool confound remains a threat to internal validity that cannot be fully resolved with our sample size. revision: partial
-
Referee: §4.2, Table 2: The large discrepancy between the relative and absolute retest effect sizes (g=-1.13 vs. g=-0.27) is driven entirely by the higher Session 1 performance with AI. The paper should clarify that the relative retest result does not provide independent evidence of differential learning beyond what RQ1 already establishes.
Authors: The referee is entirely correct. The large effect size for the relative retest change (g=-1.13) is a mathematical consequence of the higher Session 1 performance in the AI condition (established in RQ1), combined with comparable absolute retest performance across conditions (g=-0.27, non-significant). The relative retest result therefore does not constitute independent evidence of differential learning; it is a restatement of the performance difference already captured by RQ1. We will revise §4.2 and the corresponding discussion in §5.2 to make this explicit, removing any framing that could be read as the relative retest decrement providing additional evidence of a learning deficit. We will state clearly that the only learning-related finding is the non-significant trend toward lower absolute retest performance in the AI condition, and that even this should be interpreted cautiously given the sample size. revision: yes
Circularity Check
No circularity: empirical study with external theoretical framework and independently collected data
full rationale
This is an empirical HCI study comparing human-human pair programming to human-AI pair programming. The theoretical frameworks (Control-Value Theory by Pekrun, Cognitive Load Theory by Sweller, Russell's circumplex model of affect) are all external to the authors and not self-cited. The hypotheses (AI improves performance, human provides better emotional experience) are tested against data collected through a controlled within-subjects experiment with counterbalancing. No prediction or result is defined in terms of the data it claims to explain. The performance metric (Eq. 2) is defined by task completion and time, not fitted to outcomes. The emotion change scores (after - before) are computed from independent Likert ratings, not from a model fitted to the same data. Self-citation is minimal (reference [29] is a prior study by some of the same authors, used for task selection and comparison, not as a load-bearing premise). The Wald tests for confounds and the Wild Cluster Bootstrapping are standard statistical methods. The derivation chain from data collection to analysis to conclusions contains no step that reduces to its own inputs by construction.
Assumptions & free parameters
free parameters (3)
- Task pool assignment =
Pool X: tasks 14,52,72,0; Pool Y: tasks 35,9,24,18
- 20-minute time limit =
20 minutes per trial
- Scoring formula weights =
20 pts/task + 5 pts/task * (%time to spare)
assumptions (4)
- domain assumption Control-Value Theory of Achievement Emotions (Pekrun 2006, 2024) accurately describes emotional responses in programming contexts
- domain assumption The two HumanEval task pools are equivalent in difficulty
- domain assumption Self-reported emotion on 7-point Likert scales validly captures valence and arousal changes
- domain assumption One-week retention test measures learning from the team condition
Cite this review
Pith. "Pith review of OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning." pith.science (2026). https://pith.science/paper/HKCEZK66
@misc{pith2026260418530,
author = {Pith},
title = {Pith review of: OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/HKCEZK66}},
note = {Machine review of arXiv:2604.18530}
}
read the original abstract
Recent advancements in Reinforcement Learning with Verifiable Rewards (RLVR) have significantly improved Large Language Model (LLM) reasoning, yet models often struggle to explore novel trajectories beyond their initial policy distribution. While offline teacher guidance and entropy-driven strategies have been proposed to address this, they often lack deep integration or are constrained by the model's inherent capacity. In this paper, we propose OGER (Offline-Guided Exploration Reward), a novel framework that unifies offline teacher guidance and online reinforcement learning through a specialized reward modeling lens. OGER employs multi-teacher collaborative training and constructs an auxiliary exploration reward that leverages both offline trajectories and the model's own entropy to incentivize autonomous exploration. Extensive experiments across mathematical and general reasoning benchmarks demonstrate that OGER consistently outperforms competitive baselines, achieving substantial gains in mathematical reasoning while maintaining robust generalization to out-of-domain tasks. We provide a comprehensive analysis of training dynamics and conduct detailed ablation studies to validate the effectiveness of our entropy-aware reward modulation. Our code is available at https://github.com/ecoli-hit/OGER.git.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Kazunori Akizuki and Yukari Ohashi. 2015. Measurement of functional task difficulty during motor learning: What level of difficulty corresponds to the optimal challenge point?Human Movement Science43 (Oct. 2015), 107–117. doi:10.1016/j.humov.2015.07.007
-
[2]
Matin Amoozadeh, Daye Nam, Daniel Prol, Ali Alfageeh, James Prather, Michael Hilton, Sruti Srinivasa Ragavan, and Amin Alipour. 2024. Student-AI Interaction: A Case Study of CS1 students. InProceedings of the 24th Koli Calling International Conference on Computing Education Research (Koli Calling ’24). Association for Computing Machinery, New York, NY, US...
-
[3]
Lisanne Bainbridge. 1983. Ironies of automation. InAnalysis, design and evaluation of man–machine systems. Elsevier, 129–135
work page 1983
-
[4]
André Barcaui. 2025. ChatGPT as a cognitive crutch: Evidence from a randomized controlled trial on knowledge retention.Social Sciences & Humanities Open12 (2025), 102287. doi:10.1016/j.ssaho.2025.102287
-
[5]
Shraddha Barke, Michael B. James, and Nadia Polikarpova. 2023. Grounded copilot: How programmers interact with code-generating models. Proceedings of the ACM on Programming Languages7, OOPSLA1 (2023), 85–111. Number: OOPSLA1
work page 2023
-
[6]
Basawapatna, Alexander Repenning, Kyu Han Koh, and Hilarie Nickerson
Ashok R. Basawapatna, Alexander Repenning, Kyu Han Koh, and Hilarie Nickerson. 2013. The zones of proximal flow: guiding students through a space of computational thinking skills and challenges. InProceedings of the ninth annual international ACM conference on International computing education research. ACM, San Diego San California USA, 67–74. doi:10.114...
-
[7]
2004.Extreme Programming Explained: Embrace Change
Kent Beck, Cynthia Andres, and O’Reilly Online Learning: Academic/Public Library Edition. 2004.Extreme Programming Explained: Embrace Change. Addison Wesley Professional, Boston. http://RE5QY4SB7X.search.serialssolutions.com/?V=1.0&L=RE5QY4SB7X&S=JCs&C=TC0000074886&T=marc
work page 2004
-
[8]
Brett A. Becker, Paul Denny, James Finnie-Ansley, Andrew Luxton-Reilly, James Prather, and Eddie Antonio Santos. 2023. Programming Is Hard - Or at Least It Used to Be: Educational Opportunities and Challenges of AI Code Generation. InProceedings of the 54th ACM Technical Symposium on Computer Science Education V. 1 (SIGCSE 2023). Association for Computing...
Show all 85 references
-
[9]
Ben-Shachar, Daniel Lüdecke, and Dominique Makowski
Mattan S. Ben-Shachar, Daniel Lüdecke, and Dominique Makowski. 2020. effectsize: Estimation of Effect Size Indices and Standardized Parameters. Journal of Open Source Software5, 56 (2020), 2815. doi:10.21105/joss.02815
2020 doi
-
[10]
Yoav Benjamini and Yosef Hochberg. 1995. Controlling the false discovery rate: a practical and powerful approach to multiple testing.Journal of the Royal statistical society: series B (Methodological)57, 1 (1995), 289–300. Number: 1
1995
-
[11]
Simon Berger and Julian Voigt. 2025. My Way or the AI Way? How Autonomy Frustration Undermines Creative Performance Through Editing AI Advice.How Autonomy Frustration Undermines Creative Performance Through Editing AI Advice (April 28, 2025)(2025)
2025
-
[12]
Christian Bird, Denae Ford, Thomas Zimmermann, Nicole Forsgren, Eirini Kalliamvakou, Travis Lowdermilk, and Idan Gazit. 2023. Taking flight with copilot.Commun. ACM66, 6 (2023), 56–62
2023
-
[13]
Robert A. Bjork. 1994. Memory and Metamemory Considerations in the Training of Human Beings. InMetacognition: Knowing About Knowing, Janet Metcalfe and Arthur P. Shimamura (Eds.). MIT Press, Cambridge, MA, 185–205
1994
-
[14]
Colin Cameron, Jonah B
A. Colin Cameron, Jonah B. Gelbach, and Douglas L. Miller. 2008. Bootstrap-based improvements for inference with clustered errors.The review of economics and statistics90, 3 (2008), 414–427. https://direct.mit.edu/rest/article-abstract/90/3/414/57731
2008
-
[15]
Colin Cameron and Douglas L
A. Colin Cameron and Douglas L. Miller. 2015. A practitioner’s guide to cluster-robust inference.Journal of human resources50, 2 (2015), 317–372. https://jhr.uwpress.org/content/50/2/317.short
2015
-
[16]
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, and Greg Brockman. 2021. Evaluating large language models trained on code.arXiv preprint arXiv:2107.03374(2021)
2021 arXiv
-
[17]
Myra Cheng, Cinoo Lee, Pranav Khadpe, Sunny Yu, Dyllan Han, and Dan Jurafsky. 2025. Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence. doi:10.48550/arXiv.2510.01395 arXiv:2510.01395 [cs]
2025 doi
-
[18]
2013.Statistical power analysis for the behavioral sciences
Jacob Cohen. 2013.Statistical power analysis for the behavioral sciences. Academic press
2013
-
[19]
Livesey, and Justin A
Ben Colagiuri, Evan J. Livesey, and Justin A. Harris. 2011. Can expectancies produce placebo effects for implicit learning?Psychonomic Bulletin & Review18, 2 (April 2011), 399–405. doi:10.3758/s13423-010-0041-1
2011 doi
-
[20]
2013.Understanding the new statistics: Effect sizes, confidence intervals, and meta-analysis
Geoff Cumming. 2013.Understanding the new statistics: Effect sizes, confidence intervals, and meta-analysis. Routledge. https://api.taylorfrancis.com/ content/books/mono/download?identifierName=doi&identifierValue=10.4324/9780203807002&type=googlepdf Manuscript submitted to AC...
2013 doi
-
[21]
Elise Deitrick, R Benjamin Shapiro, and Brian Gravel. 2016. How do we assess equity in programming pairs? Singapore: International Society of the Learning Sciences
2016
-
[22]
Becker, and Brent N
Paul Denny, Juho Leinonen, James Prather, Andrew Luxton-Reilly, Thezyrie Amarouche, Brett A. Becker, and Brent N. Reeves. 2024. Prompt Problems: A New Programming Exercise for the Generative AI Era. InProceedings of the 55th ACM Technical Symposium on Computer Science Educatio...
2024 doi
-
[23]
Becker, James Finnie-Ansley, Arto Hellas, Juho Leinonen, Andrew Luxton-Reilly, Brent N
Paul Denny, James Prather, Brett A. Becker, James Finnie-Ansley, Arto Hellas, Juho Leinonen, Andrew Luxton-Reilly, Brent N. Reeves, Eddie Antonio Santos, and Sami Sarsa. 2024. Computing Education in the Era of Generative AI.Commun. ACM67, 2 (Jan. 2024), 56–67. doi:10.1145/3624720
2024 doi
-
[24]
Maria TM Dijkstra and Astrid C. Homan. 2016. Engaging in rather than disengaging from stress: Effective coping and perceived control.Frontiers in psychology7 (2016), 1415. https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2016.01415/full
2016 doi
-
[25]
Guangrui Fan, Dandan Liu, Rui Zhang, and Lihu Pan. 2025. The impact of AI-assisted pair programming on student motivation, programming anxiety, collaborative learning, and programming performance: a comparative study with traditional pair programming and individual approaches....
2025 doi
-
[26]
James Finnie-Ansley, Paul Denny, Andrew Luxton-Reilly, Eddie Antonio Santos, James Prather, and Brett A. Becker. 2023. My AI Wants to Know if This Will Be on the Exam: Testing OpenAI’s Codex on CS2 Programming Exercises. InProceedings of the 25th Australasian Computing Educati...
2023 doi
-
[27]
2004.The Midnight Disease: The Drive to Write, Writer’s Block, and the Creative Brain
Alice Weaver Flaherty. 2004.The Midnight Disease: The Drive to Write, Writer’s Block, and the Creative Brain. Houghton Mifflin, Boston
2004
-
[28]
Nat Friedman. 2021. Introducing GitHub Copilot: your AI pair programmer. https://github.blog/2021-06-29-introducing-github-copilot-ai-pair- programmer/
2021
-
[29]
Nicholas Gardella, Raymond Pettit, and Sara L. Riggs. 2024. Performance, Workload, Emotion, and Self-Efficacy of Novice Programmers Using AI Code Generation. InProceedings of the 2024 on Innovation and Technology in Computer Science Education V. 1. ACM, Milan Italy, 290–296. d...
2024 doi
-
[30]
Nicholas Gardella, Joseph Shelton, Isabella Graßl, and Sara Riggs. 2025. HBCU Student Perspectives on Identity, Persistence, and Code-Generating AI in CS Education: A Case Study. InProceedings of the 25th Koli Calling International Conference on Computing Education Research. A...
2025 doi
-
[31]
Gian Luca Scoccia. 2023. Exploring Early Adopters’ Perceptions of ChatGPT as a Code Generation Tool. (Sept. 2023). doi:10.1109/asew60602.2023.00016 MAG ID: 4388214696
2023 doi
-
[32]
Brian Hanks, Sue Fitzgerald, Renée McCauley, Laurie Murphy, and Carol Zander. 2011. Pair programming in education: A literature review.Computer Science Education21, 2 (2011), 135–173
2011
-
[33]
Hannay, Tore Dybå, Erik Arisholm, and Dag I
Jo E. Hannay, Tore Dybå, Erik Arisholm, and Dag I. K. Sjøberg. 2009. The effectiveness of pair programming: A meta-analysis.Information and Software Technology51, 7 (July 2009), 1110–1122. doi:10.1016/j.infsof.2009.02.001
2009 doi
-
[34]
Hart and Lowell E
Sandra G. Hart and Lowell E. Staveland. 1988. Development of NASA-TLX (Task Load Index): Results of empirical and theoretical research. In Advances in psychology. Vol. 52. Elsevier, 139–183
1988
-
[35]
Anja Hawlitschek, Sarah Berndt, and Sandra Schulz. 2023. Empirical research on pair programming in higher education: a literature review. Computer Science Education33, 3 (July 2023), 400–428. doi:10.1080/08993408.2022.2039504
2023 doi
-
[36]
Toni Honicke and Jaclyn Broadbent. 2016. The influence of academic self-efficacy on academic performance: A systematic review.Educational research review17 (2016), 63–84. https://www.sciencedirect.com/science/article/pii/S1747938X15000639
2016
-
[37]
Irene Hou, Owen Man, Kate Hamilton, Srishty Muthusekaran, Jeffin Johnykutty, Leili Zadeh, and Stephen MacNeil. 2025. ’All Roads Lead to ChatGPT’: How Generative AI is Eroding Social Interactions and Student Learning Communities. InProceedings of the 30th ACM Conference on Inno...
2025 doi
-
[38]
Irene Hou, Sophia Mettille, Owen Man, Zhuo Li, Cynthia Zastudil, and Stephen MacNeil. 2024. The Effects of Generative AI on Computing Students’ Help-Seeking Preferences. InProceedings of the 26th Australasian Computing Education Conference (ACE ’24). Association for Computing ...
2024 doi
-
[39]
Ross Ihaka and Robert Gentleman. 1996. R: A Language for Data Analysis and Graphics.Journal of Computational and Graphical Statistics5, 3 (Sept. 1996), 299–314. doi:10.1080/10618600.1996.10474713
1996 doi
-
[40]
Sumeet Jeswani, Akshay Mittal, and Soni Sewlani. 2025. Premature Trust: How Student Overreliance on GenAI is Seeding Tomorrow’s Security Gaps. InProceedings of the 26th ACM Annual Conference on Cybersecurity & Information Technology Education (SIGCITE ’25). Association for Com...
2025 doi
-
[41]
Martin Jonsson and Jakob Tholander. 2022. Cracking the code: Co-coding with AI in creative programming education. InProceedings of the 14th Conference on Creativity and Cognition (C&C ’22). Association for Computing Machinery, New York, NY, USA, 5–14. doi:10.1145/3527927.3532801
2022 doi
-
[42]
Gregor Jošt, Viktor Taneski, and Sašo Karakatič. 2024. The Impact of Large Language Models on Programming Education and Student Learning Outcomes.Applied Sciences14, 10 (Jan. 2024), 4115. doi:10.3390/app14104115 Number: 10
2024 doi
-
[43]
Maria Kallia. 2025. To Be, or to Be Otherwise? Silicon Souls in Search of Dasein and Authentic Human Engagement in the Age of Generative AI. In Proceedings of the 2025 ACM Conference on International Computing Education Research V.1 (ICER ’25). Association for Computing Machin...
2025 doi
-
[44]
Ericson, David Weintrop, and Tovi Grossman
Majeed Kazemitabaar, Justin Chow, Carl Ka To Ma, Barbara J. Ericson, David Weintrop, and Tovi Grossman. 2023. Studying the effect of AI Code Generators on Supporting Novice Learners in Introductory Programming. InProceedings of the 2023 CHI Conference on Human Factors in Compu...
2023 doi
-
[45]
Majeed Kazemitabaar, Xinying Hou, Austin Henley, Barbara Jane Ericson, David Weintrop, and Tovi Grossman. 2024. How Novices Use LLM-based Code Generators to Solve CS1 Coding Tasks in a Self-Paced Learning Environment. InProceedings of the 23rd Koli Calling International Confer...
2024 doi
-
[46]
Majeed Kazemitabaar, Runlong Ye, Xiaoning Wang, Austin Zachary Henley, Paul Denny, Michelle Craig, and Tovi Grossman. 2024. CodeAid: Evaluating a Classroom Deployment of an LLM-based Programming Assistant that Balances Student and Educator Needs. InProceedings of the 2024 CHI ...
2024 doi
-
[47]
András Komócsi, Gergely Csaba, Eszter Veronika Csöngei, Krisztina Fischer, Kristóf Filipánits, László Czopf, and Andrea Tamás. 2026. Academic performance and progression among near-peer tutors: A comparative analysis in undergraduate medical education.Medical Teacher48, 3 (Mar...
2026 doi
-
[48]
Sandeep Kaur Kuttal, Bali Ong, Kate Kwasny, and Peter Robe. 2021. Trade-offs for Substituting a Human with an Agent in a Pair Programming Context: The Good, the Bad, and the Ugly. InProceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI ’21). Associa...
2021 doi
-
[49]
Liang, Chenyang Yang, and Brad A
Jenny T. Liang, Chenyang Yang, and Brad A. Myers. 2024. A Large-Scale Survey on the Usability of AI Programming Assistants: Successes and Challenges. InProceedings of the IEEE/ACM 46th International Conference on Software Engineering (ICSE ’24). Association for Computing Machi...
2024 doi
-
[50]
Mark Liffiton, Brad E Sheese, Jaromir Savelka, and Paul Denny. 2023. CodeHelp: Using Large Language Models with Guardrails for Scalable Support in Programming Classes. InProceedings of the 23rd Koli Calling International Conference on Computing Education Research. ACM, Koli Fi...
2023 doi
-
[51]
Wenhan Lyu, Yimeng Wang, Yifan Sun, and Yixuan Zhang. 2025. Will Your Next Pair Programming Partner Be Human? An Empirical Evaluation of Generative AI as a Collaborative Teammate in a Semester-Long Classroom Setting. InProceedings of the Twelfth ACM Conference on Learning @ Sc...
2025 doi
-
[52]
MacKinnon, Morten Ørregaard Nielsen, and Matthew D
James G. MacKinnon, Morten Ørregaard Nielsen, and Matthew D. Webb. 2023. Fast and reliable jackknife and bootstrap methods for cluster-robust inference.Journal of Applied Econometrics38, 5 (Aug. 2023), 671–694. doi:10.1002/jae.2969
2023 doi
-
[53]
Margulieux, James Prather, Brent N
Lauren E. Margulieux, James Prather, Brent N. Reeves, Brett A. Becker, Gozde Cetin Uzun, Dastyni Loksa, Juho Leinonen, and Paul Denny. 2024. Self-Regulation, Self-Efficacy, and Fear of Failure Interactions with How Novices Use LLMs to Solve Programming Problems. InProceedings ...
2024 doi
-
[54]
Daye Nam, Andrew Macvean, Vincent Hellendoorn, Bogdan Vasilescu, and Brad Myers. 2024. Using an LLM to Help With Code Understanding. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering. ACM, Lisbon Portugal, 1–13. doi:10.1145/3597503.3639187
2024 doi
-
[55]
Akilan, P
Palanichamy Naveen, T. Akilan, P. Manikandan, C. Swedheetha, and M. Saravanan. 2025. Comparative Review of Large Language Models. In2025 5th International Conference on Soft Computing for Security Applications (ICSCSA). IEEE, 2120–2124. https://ieeexplore.ieee.org/abstract/doc...
2025
-
[56]
2021.The Extended Mind: The Power of Thinking Outside the Brain
Annie Murphy Paul. 2021.The Extended Mind: The Power of Thinking Outside the Brain. Houghton Mifflin Harcourt, Boston, MA
2021
-
[57]
Reinhard Pekrun. 2006. The Control-Value Theory of Achievement Emotions: Assumptions, Corollaries, and Implications for Educational Research and Practice.Educational Psychology Review18, 4 (Nov. 2006), 315–341. doi:10.1007/s10648-006-9029-9
2006 doi
-
[58]
Reinhard Pekrun. 2024. Control-Value Theory: From Achievement Emotion to a General Theory of Human Emotions.Educational Psychology Review36, 3 (Sept. 2024), 83. doi:10.1007/s10648-024-09909-7
2024 doi
-
[59]
Jacob Penney, Pawan Acharya, Peter Hilbert, Priyanka Parekh, Anita Sarma, Igor Steinmacher, and Marco Gerosa. 2025. Understanding Programming Students’ Help-Seeking Preferences in the Era of Generative AI. InProceedings of the ACM Global Computing Education Conference 2025 - V...
2025 doi
-
[60]
Reeves, Jaromir Savelka, IV Smith, David H., Sven Strickroth, and Daniel Zingaro
James Prather, Juho Leinonen, Natalie Kiesler, Jamie Gorson Benario, Sam Lau, Stephen MacNeil, Narges Norouzi, Simone Opel, Vee Pettit, Leo Porter, Brent N. Reeves, Jaromir Savelka, IV Smith, David H., Sven Strickroth, and Daniel Zingaro. 2025. Beyond the Hype: A Comprehensive...
2025 doi
-
[61]
It’s Weird That it Knows What I Want
James Prather, Brent N. Reeves, Paul Denny, Brett A. Becker, Juho Leinonen, Andrew Luxton-Reilly, Garrett Powell, James Finnie-Ansley, and Eddie Antonio Santos. 2024. “It’s Weird That it Knows What I Want”: Usability and Interactions with Copilot for Novice Programmers.ACM Tra...
2024 doi
-
[62]
Becker, Bailey Kimmel, Jared Wright, and Ben Briggs
James Prather, Brent N Reeves, Juho Leinonen, Stephen MacNeil, Arisoa S Randrianasolo, Brett A. Becker, Bailey Kimmel, Jared Wright, and Ben Briggs. 2024. The Widening Gap: The Benefits and Harms of Generative AI for Novice Programmers. InProceedings of the 2024 ACM Conference...
2024 doi
-
[63]
Ira J. Roseman. 1996. Appraisal Determinants of Emotions: Constructing a More Accurate and Comprehensive Theory.Cognition & Emotion10, 3 (May 1996), 241–278. doi:10.1080/026999396380240 Manuscript submitted to ACM 22 Gardella et al
1996 doi
-
[64]
Ross, Fernando Martinez, Stephanie Houde, Michael Muller, and Justin D
Steven I. Ross, Fernando Martinez, Stephanie Houde, Michael Muller, and Justin D. Weisz. 2023. The Programmer’s Assistant: Conversational Interaction with a Large Language Model for Software Development. InProceedings of the 28th International Conference on Intelligent User In...
2023 doi
-
[65]
James A. Russell. 1980. A circumplex model of affect.Journal of personality and social psychology39, 6 (1980), 1161. https://psycnet.apa.org/journals/ psp/39/6/1161/
1980
-
[66]
Norsaremah Salleh, Emilia Mendes, and John Grundy. 2011. Empirical Studies of Pair Programming for CS/SE Teaching in Higher Education: A Systematic Literature Review.IEEE Transactions on Software Engineering37, 4 (July 2011), 509–525. doi:10.1109/TSE.2010.59
2011 doi
-
[67]
Sharpe and Ian Tyndall
Benjamin T. Sharpe and Ian Tyndall. 2025. The Sustained Attention Paradox: A Critical Commentary on the Theoretical Impossibility of Perfect Vigilance.Cognitive Science49, 4 (2025), e70061. doi:10.1111/cogs.70061 _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/cogs.70061
2025 doi
-
[68]
Abdulhadi Shoufan. 2023. Can Students without Prior Knowledge Use ChatGPT to Answer Test Questions? An Empirical Study.ACM Trans. Comput. Educ.23, 4 (Dec. 2023), 45:1–45:29. doi:10.1145/3628162
2023 doi
-
[69]
Nancy Sommers. 1980. Revision Strategies of Student Writers and Experienced Adult Writers.College Composition & Communication31, 4 (1980), 378–388. doi:10.58680/ccc198015930 Type: Journal Article
1980 doi
-
[70]
Nancy Sommers. 1982. Responding to Student Writing.College Composition & Communication33, 2 (1982), 148–156. doi:10.58680/ccc198215854 Type: Journal Article
1982 doi
-
[71]
Son and Rajiv Sethi
Lisa K. Son and Rajiv Sethi. 2006. Metacognitive Control and Optimal Learning.Cognitive Science30, 4 (July 2006), 759–774. doi:10.1207/ s15516709cog0000_74
2006
-
[72]
Phil Steinhorst, Andrew Petersen, and Jan Vahrenhold. 2020. Revisiting Self-Efficacy in Introductory Programming. InProceedings of the 2020 ACM Conference on International Computing Education Research. ACM, Virtual Event New Zealand, 158–169. doi:10.1145/3372782.3406281
2020 doi
-
[73]
John Sweller. 2011. Cognitive load theory. InPsychology of learning and motivation. Vol. 55. Elsevier, 37–76. https://www.sciencedirect.com/science/ article/pii/B9780123876911000028
2011
-
[74]
Ritzhaupt
Karthikeyan Umapathy and Albert D. Ritzhaupt. 2017. A meta-analysis of pair-programming in computer programming courses: Implications for educational practice.ACM Transactions on Computing Education (TOCE)17, 4 (2017), 1–13
2017
-
[75]
Smith IV, Mounika Padala, Christine Alvarado, Jamie Gorson Benario, and Leo Porter
Annapurna Vadaparty, Daniel Zingaro, David H. Smith IV, Mounika Padala, Christine Alvarado, Jamie Gorson Benario, and Leo Porter. 2024. CS1-LLM: Integrating LLMs into CS1 Instruction. InProceedings of the 2024 on Innovation and Technology in Computer Science Education V. 1 (IT...
2024 doi
-
[76]
Marcel Valový and Alena Buchalcevova. 2023. The Psychological Effects of AI-Assisted Programming on Students and Professionals. In2023 IEEE International Conference on Software Maintenance and Evolution (ICSME). 385–390. doi:10.1109/ICSME58846.2023.00050 ISSN: 2576-3148
2023 doi
-
[77]
Vygotsky
Lev S. Vygotsky. 1978.Mind in Society: The Development of Higher Psychological Processes. Harvard University Press, Cambridge, MA
1978
-
[78]
Matthew D. Webb. 2023. Reworking wild bootstrap-based inference for clustered errors.Canadian Journal of Economics/Revue canadienne d’économique56, 3 (Aug. 2023), 839–858. doi:10.1111/caje.12661
2023 doi
-
[79]
Weisz, Michael Muller, Michael Muller, Michael Muller, Stephanie Houde, John T
Justin D. Weisz, Michael Muller, Michael Muller, Michael Muller, Stephanie Houde, John T. Richards, Steven I. Ross, Fernando Martinez, Mayank Agarwal, and Kartik Talamadupula. 2021. Perfection Not Required? Human-AI Partnerships in Code Translation. 402–412. doi:10.1145/339748...
2021 doi
-
[80]
Bai, Robert Tairas, and Yu Huang
Yuankai Xue, Hanlin Chen, Gina R. Bai, Robert Tairas, and Yu Huang. 2024. Does ChatGPT Help With Introductory Programming?An Experiment of Students Using ChatGPT in CS1. InProceedings of the 46th International Conference on Software Engineering: Software Engineering Education ...
2024 doi
-
[81]
Burak Yetistiren, Isik Ozsoy, and Eray Tuzun. 2022. Assessing the quality of GitHub copilot’s code generation. InProceedings of the 18th International Conference on Predictive Models and Data Analytics in Software Engineering. 62–71
2022
-
[82]
Ramazan Yilmaz and Fatma Gizem Karaoglan Yilmaz. 2023. The effect of generative artificial intelligence (AI)-based tool use on students’ computational thinking skills, programming self-efficacy and motivation.Computers and Education: Artificial Intelligence4 (Jan. 2023), 10014...
2023 doi
-
[83]
Achim Zeileis. 2004. Econometric computing with HC and HAC covariance matrix estimators.Journal of statistical software11 (2004), 1–17. https://www.jstatsoft.org/article/view/v011i10/0
2004
-
[84]
Achim Zeileis and Torsten Hothorn. 2002. Diagnostic Checking in Regression Relationships.R News2, 3 (2002), 7–10. https://CRAN.R- project.org/doc/Rnews/
2002
-
[85]
Achim Zeileis, Susanne Köll, and Nathaniel Graham. 2020. Various versatile variances: an object-oriented implementation of clustered covariances in R.Journal of Statistical Software95 (2020), 1–36. https://www.jstatsoft.org/article/view/v095i01/0 Received 20 April 2026; revise...
2020
Reviewed July 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.