Pith. sign in

REVIEW 3 major objections 9 minor 85 references

OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning

T0 review · 3 major / 9 minor · reviewed 2026-07-05 · glm-5.2

Pith's one-line read AI pair programmer boosts speed but flattens emotion and learning

desk verdict Well-designed study of human vs. AI pair programming finds AI boosts performance but humans boost emotion; sample size limits retention claims. read the letter →

arxiv 2604.18530 v2 pith:HKCEZK66 submitted 2026-04-20 cs.AI

classification cs.AI
keywords pairprogrammingGitHubCopilotcomputingeducationemotionlearningretentionControl-ValueTheoryNASA-TLXwithin-subjectsexperiment
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper reports a controlled within-subjects experiment in which 22 novice programmers completed coding tasks both with a human teammate and with GitHub Copilot. The central finding is a trade-off: participants performed significantly better and reported lower mental workload when paired with Copilot, but experienced significantly less positive emotion and less arousal compared to working with a human partner. One week later, there was a trend (marginally significant relative to initial performance) toward worse retention of tasks originally learned with Copilot, particularly for the stronger member of each human pair. The authors interpret these results through Control-Value Theory, arguing that AI assistance raises outcome value (better results) while reducing the programmer's sense of control and activity value, which in turn diminishes emotional engagement. The paper recommends that educators preserve human pair programming as a pedagogical tool rather than allowing AI to fully replace it, because the social, conflict-driven, and emotionally engaging qualities of human collaboration appear to support deeper cognitive processing even when raw performance is lower.

What carries the argument

The experiment uses a within-subjects counterbalanced design where each participant programs in both conditions (human teammate and Copilot) across two task pools drawn from HumanEval, with a one-week retention interval followed by a solo retest. The analytical machinery is Ordinary Least Squares regression with Wild Cluster Bootstrapping to handle the small sample and nested data structure (participants within pairs, repeated measures). The theoretical interpretive frame is Pekrun's Control-Value Theory, which predicts that emotions follow from appraisals of both control over an activity and the value of its outcome, explaining why Copilot's high competence raises outcome value while lowing

What would settle it

A larger replication with more participants and independently validated task difficulty equivalence that finds no retention difference between conditions, or that finds the performance gap disappears when task pools are more rigorously matched.

Watch

Extended reading notes

Core claim

The paper's central discovery is that AI-assisted programming and human pair programming produce opposite profiles across four measured dimensions. Copilot yielded higher task performance (approximately 14 points on a 100-point scale, a large effect) and lower mental demand, temporal demand, and effort (all large effects), but the human condition produced significantly more positive emotional valence and higher arousal (both large effects). Retention performance one week later was statistically comparable in absolute terms between conditions, but the drop from initial team performance to solo retest was significantly larger for tasks learned with Copilot, suggesting that the higher initialAI

Load-bearing premise

The paper assumes that the two sets of programming tasks (Pool X and Pool Y) are equivalent in difficulty, and with only 22 participants and 11 pairs, the study has limited statistical power to detect a task-pool-by-condition interaction that could confound the results.

Editorial extensions

If this is right

  • Educators who replace pair programming with AI-assisted programming may gain short-term productivity but risk eroding the emotional engagement and social learning dynamics that support deeper processing.
  • AI tool designers could explore rate-limiting or scaffolding features that withhold full solutions, forcing the learner to maintain cognitive control and potentially preserving both engagement and retention.
  • The finding that stronger teammates learned less when their first exposure was through Copilot suggests that the act of teaching a peer may itself be a learning mechanism that AI assistance bypasses.
  • The emotional flatness associated with Copilot, if generalizable, could have cumulative effects on motivation and persistence in programming courses where students increasingly default to AI tools.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the emotional and retention differences hold in larger samples, the result would imply an inverse relationship between immediate task performance and durable learning in AI-assisted contexts, a pattern reminiscent of 'desirable difficulties' in cognitive psychology.
  • The marginally significant retention effects, combined with the small sample, suggest that a larger replication could either confirm a genuine learning cost of AI assistance or reveal that the effect is unstable; either outcome would be informative.
  • The finding that participants did not credit themselves for Copilot-enhanced performance raises the possibility that AI tools may undermine self-efficacy formation in ways not captured by standard self-report measures.
  • If task difficulty equivalence between the two pools is not perfect, some portion of the performance gap could reflect task-Copilot interaction rather than the teammate condition alone, though the paper's Wald tests found no significant confound.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 9 minor

Summary. This manuscript reports a within-subjects controlled experiment (n=22, 11 pairs) comparing human-human pair programming to human-AI pair programming (GitHub Copilot) among novice and intermediate programmers. The study measures four outcomes: programming performance (RQ1), learning retention after one week (RQ2), subjective workload via NASA-TLX (RQ3), and emotional impact via valence/arousal change scores (RQ4). The central findings are that AI assistance significantly improves performance and reduces workload but diminishes emotional engagement compared to human pair programming. Learning retention differences were mostly non-significant, with only a marginally significant relative performance decrement in the AI condition. The study is grounded in Control-Value Theory of emotion and Cognitive Load Theory. The experimental design is careful: counterbalancing across eight permutations, deception with debriefing, pre-registered-style task pools, and Wild Cluster Bootstrapping for inference given the small clustered sample. The paper also provides qualitative triangulation from exit questionnaires. Note: the abstract and arXiv metadata describe an RL/LLM paper (OGER), but the full text is an HCI/computing education paper; this review addresses the actual full-text manuscript.

Significance. The paper makes a timely contribution to computing education research by directly comparing human-human and human-AI pair programming on a multi-dimensional outcome set (performance, learning, workload, emotion) within the same participants. The emotion finding (RQ4) is the most distinctive contribution, as prior work has focused primarily on performance and workload. The within-subjects design with counterbalancing, the use of WCB inference appropriate for few-cluster data, and the BH correction within RQ families are methodological strengths. The authors provide open appendices and task materials via Zenodo, supporting reproducibility. The theoretical framing via CVT is well-motivated and the qualitative data triangulates the quantitative findings effectively.

major comments (3)
  1. §4.4, Table 4: The emotion findings (valence change B=-1.18, g=-1.61; arousal change B=-2.09, g=-1.74) are the paper's most novel contribution, but they rely on sequential within-subjects change scores where differential carryover could inflate the apparent AI-vs-human gap. If the human condition produces a strong positive emotional response (which is the finding), participants who experience it first carry an elevated baseline into the subsequent AI trial, making the AI condition's change score artificially negative. This is an asymmetric carryover effect that counterbalancing cannot fully address. The paper acknowledges this in §6 as a caveat, but the Wald test for ordering effects (p>.1, §3.7) has very low power at 11 clusters to detect such an interaction. The paper should either (a) report the emotion results stratified by ordering (human-first vs. AI-first) so readers can assess to
  2. §3.3: The assumption that the two task pools (Pool X: tasks 14, 52, 72, 0; Pool Y: tasks 35, 9, 24, 18) are equivalent in difficulty is load-bearing for all four RQs. The paper states pools were 'as similar as possible in difficulty' and reports Wald tests (p>.1), but with only 11 pairs, the power to detect a task-pool-by-condition interaction is very low. If one pool happened to be more amenable to Copilot's strengths (e.g., tasks requiring more boilerplate vs. algorithmic reasoning), the performance difference could be partially attributable to task selection. The paper should report descriptive statistics (mean scores) for each pool separately by condition, or at minimum acknowledge this power limitation more explicitly in §6.
  3. §4.2, Table 2: The learning retention analysis (RQ2) reports a marginally significant relative retest decrement (p_adj=.054, g=-1.13) but non-significant absolute retest performance (p_adj=.529, g=-0.27). The large discrepancy between the relative and absolute effect sizes (g=-1.13 vs. g=-0.27) is driven entirely by the higher Session 1 performance with AI. This means the 'large' effect for relative retest change is a mathematical consequence of the RQ1 performance finding, not independent evidence of a learning difference. The paper should clarify that the relative retest result does not provide independent evidence of differential learning beyond what RQ1 already establishes.
minor comments (9)
  1. The arXiv abstract and metadata describe a completely different paper (OGER: Offline-Guided Exploration Reward for RL). The title, abstract, and metadata should be corrected to match the actual manuscript content.
  2. §3.7: 'Benjimini-Hochberg' is misspelled (should be 'Benjamini-Hochberg') in multiple places, including Tables 1-4 captions and the main text.
  3. §3.6.1: The emotion measure uses 7-point Likert scales for valence and arousal, but the paper does not specify the exact anchor labels. Including these would help readers interpret the magnitude of the changes.
  4. Fig. 2a/2b: The figure captions reference subfigures (a) and (b) but the timeline details are small and difficult to read. Consider enlarging or providing a text-based timeline in addition to the figure.
  5. §4.1: The claim that the performance advantage corresponds to '7/10 of a task or the same number of tasks but faster by a margin exceeding half of the total time limit' is somewhat informal. A more precise statement of the effect in terms of the scoring formula (Eq. 2) would be clearer.
  6. §5.2: The discussion of 'near-peer teaching' as an explanation for stronger teammates' lower AI-condition retest performance is speculative. The paper should either flag this more clearly as post-hoc or provide additional supporting evidence.
  7. §3.1: The demographic skew (16 male, 15 Asian out of 22) is noted but not discussed as a potential limitation on generalizability. A brief mention in §6 would be appropriate.
  8. Table 3 caption states 'Correcting p-values for multiple testing would not change the conclusions here, even with the very conservative Bonferroni adjustment.' This is an editorial claim that should be verified and, if accurate, stated more precisely.
  9. §4.4: The paper notes that 'valence and arousal simply tended to move together.' A correlation between valence and arousal change scores could be reported to quantify this.

Simulated Author's Rebuttal

3 responses · 0 unresolved

We thank the referee for a careful and constructive review. The referee raises three substantive points concerning (1) potential asymmetric carryover in the emotion change scores, (2) the equivalence of the two task pools, and (3) the interpretation of the relative retest decrement. We agree with all three points and will revise the manuscript accordingly. Below we address each in turn.

read point-by-point responses
  1. Referee: §4.4, Table 4: The emotion findings rely on sequential within-subjects change scores where differential carryover could inflate the apparent AI-vs-human gap. The paper should report emotion results stratified by ordering (human-first vs. AI-first) or otherwise address this concern.

    Authors: The referee is correct that asymmetric carryover is a genuine threat to the emotion change-score analysis, and we agree that the Wald test for ordering effects is underpowered at 11 clusters to detect such an interaction. We will address this in revision by reporting the emotion results (valence change and arousal change) stratified by ordering group (human-first vs. AI-first) in a supplementary table, so readers can directly assess whether the pattern is consistent across orderings. We will also expand the caveat in §6 to explicitly state that the low power of the ordering test means it cannot rule out asymmetric carryover, and that the stratified results should be interpreted accordingly. We want to be transparent: if the stratified analysis reveals that the effect is concentrated in one ordering group, this would weaken the causal interpretation of the emotion finding, and we will say so plainly. revision: partial

  2. Referee: §3.3: The assumption that the two task pools are equivalent in difficulty is load-bearing for all four RQs. The paper should report descriptive statistics for each pool separately by condition, or at minimum acknowledge this power limitation more explicitly in §6.

    Authors: We agree that task-pool equivalence is a critical assumption and that the Wald test is underpowered to detect a task-pool-by-condition interaction at this sample size. We will add a supplementary table reporting mean scores (performance, retest, workload subscales, and emotion change) for each pool separately by condition. This will allow readers to assess whether any pool appears systematically more or less amenable to AI assistance. We will also strengthen the discussion in §6 to explicitly acknowledge that the non-significant Wald test does not constitute strong evidence of pool equivalence given the limited power, and that the task-pool confound remains a threat to internal validity that cannot be fully resolved with our sample size. revision: partial

  3. Referee: §4.2, Table 2: The large discrepancy between the relative and absolute retest effect sizes (g=-1.13 vs. g=-0.27) is driven entirely by the higher Session 1 performance with AI. The paper should clarify that the relative retest result does not provide independent evidence of differential learning beyond what RQ1 already establishes.

    Authors: The referee is entirely correct. The large effect size for the relative retest change (g=-1.13) is a mathematical consequence of the higher Session 1 performance in the AI condition (established in RQ1), combined with comparable absolute retest performance across conditions (g=-0.27, non-significant). The relative retest result therefore does not constitute independent evidence of differential learning; it is a restatement of the performance difference already captured by RQ1. We will revise §4.2 and the corresponding discussion in §5.2 to make this explicit, removing any framing that could be read as the relative retest decrement providing additional evidence of a learning deficit. We will state clearly that the only learning-related finding is the non-significant trend toward lower absolute retest performance in the AI condition, and that even this should be interpreted cautiously given the sample size. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: empirical study with external theoretical framework and independently collected data

full rationale

This is an empirical HCI study comparing human-human pair programming to human-AI pair programming. The theoretical frameworks (Control-Value Theory by Pekrun, Cognitive Load Theory by Sweller, Russell's circumplex model of affect) are all external to the authors and not self-cited. The hypotheses (AI improves performance, human provides better emotional experience) are tested against data collected through a controlled within-subjects experiment with counterbalancing. No prediction or result is defined in terms of the data it claims to explain. The performance metric (Eq. 2) is defined by task completion and time, not fitted to outcomes. The emotion change scores (after - before) are computed from independent Likert ratings, not from a model fitted to the same data. Self-citation is minimal (reference [29] is a prior study by some of the same authors, used for task selection and comparison, not as a load-bearing premise). The Wald tests for confounds and the Wild Cluster Bootstrapping are standard statistical methods. The derivation chain from data collection to analysis to conclusions contains no step that reduces to its own inputs by construction.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

No new entities are invented. The free parameters are design choices (task selection, time limits, scoring weights) rather than fitted model parameters. The axioms are standard domain assumptions from established theories (CVT, cognitive load theory) and methodological choices, not ad hoc constructions.

free parameters (3)
  • Task pool assignment = Pool X: tasks 14,52,72,0; Pool Y: tasks 35,9,24,18
    The specific HumanEval tasks chosen for each pool determine difficulty equivalence; chosen by the authors, not empirically validated for equivalence in this population.
  • 20-minute time limit = 20 minutes per trial
    The time limit constrains the performance score formula and affects how much spare time participants have for review; chosen by design.
  • Scoring formula weights = 20 pts/task + 5 pts/task * (%time to spare)
    The scoring formula (Eq. 2) weights task completion vs. speed in a specific ratio chosen by the authors.
assumptions (4)
  • domain assumption Control-Value Theory of Achievement Emotions (Pekrun 2006, 2024) accurately describes emotional responses in programming contexts
    The theoretical framework is invoked throughout (Sections 2.4, 5) to interpret emotional findings, but CVT has not been extensively validated in programming education contexts.
  • domain assumption The two HumanEval task pools are equivalent in difficulty
    Stated in Section 3.3; supported by Wald tests (p>.1) but underpowered with n=22.
  • domain assumption Self-reported emotion on 7-point Likert scales validly captures valence and arousal changes
    Section 3.6.1; uses Russell's circumplex model. Single-item scales have known reliability limitations.
  • domain assumption One-week retention test measures learning from the team condition
    Section 3.6.3; assumes no external Python study during the gap (enforced by protocol but not verified).

how reviews work

0 comments
Cite this review

Pith. "Pith review of OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning." pith.science (2026). https://pith.science/paper/HKCEZK66

@misc{pith2026260418530,
  author       = {Pith},
  title        = {Pith review of: OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HKCEZK66}},
  note         = {Machine review of arXiv:2604.18530}
}
read the original abstract

Recent advancements in Reinforcement Learning with Verifiable Rewards (RLVR) have significantly improved Large Language Model (LLM) reasoning, yet models often struggle to explore novel trajectories beyond their initial policy distribution. While offline teacher guidance and entropy-driven strategies have been proposed to address this, they often lack deep integration or are constrained by the model's inherent capacity. In this paper, we propose OGER (Offline-Guided Exploration Reward), a novel framework that unifies offline teacher guidance and online reinforcement learning through a specialized reward modeling lens. OGER employs multi-teacher collaborative training and constructs an auxiliary exploration reward that leverages both offline trajectories and the model's own entropy to incentivize autonomous exploration. Extensive experiments across mathematical and general reasoning benchmarks demonstrate that OGER consistently outperforms competitive baselines, achieving substantial gains in mathematical reasoning while maintaining robust generalization to out-of-domain tasks. We provide a comprehensive analysis of training dynamics and conduct detailed ablation studies to validate the effectiveness of our entropy-aware reward modulation. Our code is available at https://github.com/ecoli-hit/OGER.git.

Figures

Figures reproduced from arXiv: 2604.18530 by the authors.

Figure 1
Figure 1. The overall architecture of the OGER framework. We first construct a comprehensive, high-quality offline [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Distribution of sequence lengths for trajec [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Comparative analysis of training dynamics across OGER, its variant OGER w/o Refinement, and baselines [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: pass@k performance on AIME 2024 and AIME 2025 using 256 rollouts. Our proposed OGER method consistently outperforms all baselines across various k values, demonstrating a significantly higher convergence rate in solvability coverage. These results demonstrate that our …
Figure 5
Figure 5. Figure 5: The evolution of the OGER exploration re [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]
Figure 6
Figure 6. Figure 6: The pass@8 performance across different inference temperatures on AIME 2024 and AIME 2025, we illustrate the average score. high-quality reasoning patterns from the offline teacher trajectories. As training progresses into the mid-to-late stages, the model’s intrinsic …

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

85 extracted references · 85 canonical work pages

  1. [1]

    Kazunori Akizuki and Yukari Ohashi. 2015. Measurement of functional task difficulty during motor learning: What level of difficulty corresponds to the optimal challenge point?Human Movement Science43 (Oct. 2015), 107–117. doi:10.1016/j.humov.2015.07.007

  2. [2]

    Matin Amoozadeh, Daye Nam, Daniel Prol, Ali Alfageeh, James Prather, Michael Hilton, Sruti Srinivasa Ragavan, and Amin Alipour. 2024. Student-AI Interaction: A Case Study of CS1 students. InProceedings of the 24th Koli Calling International Conference on Computing Education Research (Koli Calling ’24). Association for Computing Machinery, New York, NY, US...

  3. [3]

    Lisanne Bainbridge. 1983. Ironies of automation. InAnalysis, design and evaluation of man–machine systems. Elsevier, 129–135

  4. [4]

    André Barcaui. 2025. ChatGPT as a cognitive crutch: Evidence from a randomized controlled trial on knowledge retention.Social Sciences & Humanities Open12 (2025), 102287. doi:10.1016/j.ssaho.2025.102287

  5. [5]

    James, and Nadia Polikarpova

    Shraddha Barke, Michael B. James, and Nadia Polikarpova. 2023. Grounded copilot: How programmers interact with code-generating models. Proceedings of the ACM on Programming Languages7, OOPSLA1 (2023), 85–111. Number: OOPSLA1

  6. [6]

    Basawapatna, Alexander Repenning, Kyu Han Koh, and Hilarie Nickerson

    Ashok R. Basawapatna, Alexander Repenning, Kyu Han Koh, and Hilarie Nickerson. 2013. The zones of proximal flow: guiding students through a space of computational thinking skills and challenges. InProceedings of the ninth annual international ACM conference on International computing education research. ACM, San Diego San California USA, 67–74. doi:10.114...

  7. [7]

    2004.Extreme Programming Explained: Embrace Change

    Kent Beck, Cynthia Andres, and O’Reilly Online Learning: Academic/Public Library Edition. 2004.Extreme Programming Explained: Embrace Change. Addison Wesley Professional, Boston. http://RE5QY4SB7X.search.serialssolutions.com/?V=1.0&L=RE5QY4SB7X&S=JCs&C=TC0000074886&T=marc

  8. [8]

    Becker, Paul Denny, James Finnie-Ansley, Andrew Luxton-Reilly, James Prather, and Eddie Antonio Santos

    Brett A. Becker, Paul Denny, James Finnie-Ansley, Andrew Luxton-Reilly, James Prather, and Eddie Antonio Santos. 2023. Programming Is Hard - Or at Least It Used to Be: Educational Opportunities and Challenges of AI Code Generation. InProceedings of the 54th ACM Technical Symposium on Computer Science Education V. 1 (SIGCSE 2023). Association for Computing...

Show all 85 references
  1. [9]

    Ben-Shachar, Daniel Lüdecke, and Dominique Makowski

    Mattan S. Ben-Shachar, Daniel Lüdecke, and Dominique Makowski. 2020. effectsize: Estimation of Effect Size Indices and Standardized Parameters. Journal of Open Source Software5, 56 (2020), 2815. doi:10.21105/joss.02815

  2. [10]

    Yoav Benjamini and Yosef Hochberg. 1995. Controlling the false discovery rate: a practical and powerful approach to multiple testing.Journal of the Royal statistical society: series B (Methodological)57, 1 (1995), 289–300. Number: 1

  3. [11]

    Simon Berger and Julian Voigt. 2025. My Way or the AI Way? How Autonomy Frustration Undermines Creative Performance Through Editing AI Advice.How Autonomy Frustration Undermines Creative Performance Through Editing AI Advice (April 28, 2025)(2025)

  4. [12]

    Christian Bird, Denae Ford, Thomas Zimmermann, Nicole Forsgren, Eirini Kalliamvakou, Travis Lowdermilk, and Idan Gazit. 2023. Taking flight with copilot.Commun. ACM66, 6 (2023), 56–62

  5. [13]

    Robert A. Bjork. 1994. Memory and Metamemory Considerations in the Training of Human Beings. InMetacognition: Knowing About Knowing, Janet Metcalfe and Arthur P. Shimamura (Eds.). MIT Press, Cambridge, MA, 185–205

  6. [14]

    Colin Cameron, Jonah B

    A. Colin Cameron, Jonah B. Gelbach, and Douglas L. Miller. 2008. Bootstrap-based improvements for inference with clustered errors.The review of economics and statistics90, 3 (2008), 414–427. https://direct.mit.edu/rest/article-abstract/90/3/414/57731

  7. [15]

    Colin Cameron and Douglas L

    A. Colin Cameron and Douglas L. Miller. 2015. A practitioner’s guide to cluster-robust inference.Journal of human resources50, 2 (2015), 317–372. https://jhr.uwpress.org/content/50/2/317.short

  8. [16]

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, and Greg Brockman. 2021. Evaluating large language models trained on code.arXiv preprint arXiv:2107.03374(2021)

  9. [17]

    Myra Cheng, Cinoo Lee, Pranav Khadpe, Sunny Yu, Dyllan Han, and Dan Jurafsky. 2025. Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence. doi:10.48550/arXiv.2510.01395 arXiv:2510.01395 [cs]

  10. [18]

    2013.Statistical power analysis for the behavioral sciences

    Jacob Cohen. 2013.Statistical power analysis for the behavioral sciences. Academic press

  11. [19]

    Livesey, and Justin A

    Ben Colagiuri, Evan J. Livesey, and Justin A. Harris. 2011. Can expectancies produce placebo effects for implicit learning?Psychonomic Bulletin & Review18, 2 (April 2011), 399–405. doi:10.3758/s13423-010-0041-1

  12. [20]

    2013.Understanding the new statistics: Effect sizes, confidence intervals, and meta-analysis

    Geoff Cumming. 2013.Understanding the new statistics: Effect sizes, confidence intervals, and meta-analysis. Routledge. https://api.taylorfrancis.com/ content/books/mono/download?identifierName=doi&identifierValue=10.4324/9780203807002&type=googlepdf Manuscript submitted to AC...

  13. [21]

    Elise Deitrick, R Benjamin Shapiro, and Brian Gravel. 2016. How do we assess equity in programming pairs? Singapore: International Society of the Learning Sciences

  14. [22]

    Becker, and Brent N

    Paul Denny, Juho Leinonen, James Prather, Andrew Luxton-Reilly, Thezyrie Amarouche, Brett A. Becker, and Brent N. Reeves. 2024. Prompt Problems: A New Programming Exercise for the Generative AI Era. InProceedings of the 55th ACM Technical Symposium on Computer Science Educatio...

  15. [23]

    Becker, James Finnie-Ansley, Arto Hellas, Juho Leinonen, Andrew Luxton-Reilly, Brent N

    Paul Denny, James Prather, Brett A. Becker, James Finnie-Ansley, Arto Hellas, Juho Leinonen, Andrew Luxton-Reilly, Brent N. Reeves, Eddie Antonio Santos, and Sami Sarsa. 2024. Computing Education in the Era of Generative AI.Commun. ACM67, 2 (Jan. 2024), 56–67. doi:10.1145/3624720

  16. [24]

    Maria TM Dijkstra and Astrid C. Homan. 2016. Engaging in rather than disengaging from stress: Effective coping and perceived control.Frontiers in psychology7 (2016), 1415. https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2016.01415/full

  17. [25]

    Guangrui Fan, Dandan Liu, Rui Zhang, and Lihu Pan. 2025. The impact of AI-assisted pair programming on student motivation, programming anxiety, collaborative learning, and programming performance: a comparative study with traditional pair programming and individual approaches....

  18. [26]

    James Finnie-Ansley, Paul Denny, Andrew Luxton-Reilly, Eddie Antonio Santos, James Prather, and Brett A. Becker. 2023. My AI Wants to Know if This Will Be on the Exam: Testing OpenAI’s Codex on CS2 Programming Exercises. InProceedings of the 25th Australasian Computing Educati...

  19. [27]

    2004.The Midnight Disease: The Drive to Write, Writer’s Block, and the Creative Brain

    Alice Weaver Flaherty. 2004.The Midnight Disease: The Drive to Write, Writer’s Block, and the Creative Brain. Houghton Mifflin, Boston

  20. [28]

    Nat Friedman. 2021. Introducing GitHub Copilot: your AI pair programmer. https://github.blog/2021-06-29-introducing-github-copilot-ai-pair- programmer/

  21. [29]

    Nicholas Gardella, Raymond Pettit, and Sara L. Riggs. 2024. Performance, Workload, Emotion, and Self-Efficacy of Novice Programmers Using AI Code Generation. InProceedings of the 2024 on Innovation and Technology in Computer Science Education V. 1. ACM, Milan Italy, 290–296. d...

  22. [30]

    Nicholas Gardella, Joseph Shelton, Isabella Graßl, and Sara Riggs. 2025. HBCU Student Perspectives on Identity, Persistence, and Code-Generating AI in CS Education: A Case Study. InProceedings of the 25th Koli Calling International Conference on Computing Education Research. A...

  23. [31]

    Gian Luca Scoccia. 2023. Exploring Early Adopters’ Perceptions of ChatGPT as a Code Generation Tool. (Sept. 2023). doi:10.1109/asew60602.2023.00016 MAG ID: 4388214696

  24. [32]

    Brian Hanks, Sue Fitzgerald, Renée McCauley, Laurie Murphy, and Carol Zander. 2011. Pair programming in education: A literature review.Computer Science Education21, 2 (2011), 135–173

  25. [33]

    Hannay, Tore Dybå, Erik Arisholm, and Dag I

    Jo E. Hannay, Tore Dybå, Erik Arisholm, and Dag I. K. Sjøberg. 2009. The effectiveness of pair programming: A meta-analysis.Information and Software Technology51, 7 (July 2009), 1110–1122. doi:10.1016/j.infsof.2009.02.001

  26. [34]

    Hart and Lowell E

    Sandra G. Hart and Lowell E. Staveland. 1988. Development of NASA-TLX (Task Load Index): Results of empirical and theoretical research. In Advances in psychology. Vol. 52. Elsevier, 139–183

  27. [35]

    Anja Hawlitschek, Sarah Berndt, and Sandra Schulz. 2023. Empirical research on pair programming in higher education: a literature review. Computer Science Education33, 3 (July 2023), 400–428. doi:10.1080/08993408.2022.2039504

  28. [36]

    Toni Honicke and Jaclyn Broadbent. 2016. The influence of academic self-efficacy on academic performance: A systematic review.Educational research review17 (2016), 63–84. https://www.sciencedirect.com/science/article/pii/S1747938X15000639

  29. [37]

    Irene Hou, Owen Man, Kate Hamilton, Srishty Muthusekaran, Jeffin Johnykutty, Leili Zadeh, and Stephen MacNeil. 2025. ’All Roads Lead to ChatGPT’: How Generative AI is Eroding Social Interactions and Student Learning Communities. InProceedings of the 30th ACM Conference on Inno...

  30. [38]

    Irene Hou, Sophia Mettille, Owen Man, Zhuo Li, Cynthia Zastudil, and Stephen MacNeil. 2024. The Effects of Generative AI on Computing Students’ Help-Seeking Preferences. InProceedings of the 26th Australasian Computing Education Conference (ACE ’24). Association for Computing ...

  31. [39]

    Ross Ihaka and Robert Gentleman. 1996. R: A Language for Data Analysis and Graphics.Journal of Computational and Graphical Statistics5, 3 (Sept. 1996), 299–314. doi:10.1080/10618600.1996.10474713

  32. [40]

    Sumeet Jeswani, Akshay Mittal, and Soni Sewlani. 2025. Premature Trust: How Student Overreliance on GenAI is Seeding Tomorrow’s Security Gaps. InProceedings of the 26th ACM Annual Conference on Cybersecurity & Information Technology Education (SIGCITE ’25). Association for Com...

  33. [41]

    Martin Jonsson and Jakob Tholander. 2022. Cracking the code: Co-coding with AI in creative programming education. InProceedings of the 14th Conference on Creativity and Cognition (C&C ’22). Association for Computing Machinery, New York, NY, USA, 5–14. doi:10.1145/3527927.3532801

  34. [42]

    Gregor Jošt, Viktor Taneski, and Sašo Karakatič. 2024. The Impact of Large Language Models on Programming Education and Student Learning Outcomes.Applied Sciences14, 10 (Jan. 2024), 4115. doi:10.3390/app14104115 Number: 10

  35. [43]

    Maria Kallia. 2025. To Be, or to Be Otherwise? Silicon Souls in Search of Dasein and Authentic Human Engagement in the Age of Generative AI. In Proceedings of the 2025 ACM Conference on International Computing Education Research V.1 (ICER ’25). Association for Computing Machin...

  36. [44]

    Ericson, David Weintrop, and Tovi Grossman

    Majeed Kazemitabaar, Justin Chow, Carl Ka To Ma, Barbara J. Ericson, David Weintrop, and Tovi Grossman. 2023. Studying the effect of AI Code Generators on Supporting Novice Learners in Introductory Programming. InProceedings of the 2023 CHI Conference on Human Factors in Compu...

  37. [45]

    Majeed Kazemitabaar, Xinying Hou, Austin Henley, Barbara Jane Ericson, David Weintrop, and Tovi Grossman. 2024. How Novices Use LLM-based Code Generators to Solve CS1 Coding Tasks in a Self-Paced Learning Environment. InProceedings of the 23rd Koli Calling International Confer...

  38. [46]

    Majeed Kazemitabaar, Runlong Ye, Xiaoning Wang, Austin Zachary Henley, Paul Denny, Michelle Craig, and Tovi Grossman. 2024. CodeAid: Evaluating a Classroom Deployment of an LLM-based Programming Assistant that Balances Student and Educator Needs. InProceedings of the 2024 CHI ...

  39. [47]

    András Komócsi, Gergely Csaba, Eszter Veronika Csöngei, Krisztina Fischer, Kristóf Filipánits, László Czopf, and Andrea Tamás. 2026. Academic performance and progression among near-peer tutors: A comparative analysis in undergraduate medical education.Medical Teacher48, 3 (Mar...

  40. [48]

    Sandeep Kaur Kuttal, Bali Ong, Kate Kwasny, and Peter Robe. 2021. Trade-offs for Substituting a Human with an Agent in a Pair Programming Context: The Good, the Bad, and the Ugly. InProceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI ’21). Associa...

  41. [49]

    Liang, Chenyang Yang, and Brad A

    Jenny T. Liang, Chenyang Yang, and Brad A. Myers. 2024. A Large-Scale Survey on the Usability of AI Programming Assistants: Successes and Challenges. InProceedings of the IEEE/ACM 46th International Conference on Software Engineering (ICSE ’24). Association for Computing Machi...

  42. [50]

    Mark Liffiton, Brad E Sheese, Jaromir Savelka, and Paul Denny. 2023. CodeHelp: Using Large Language Models with Guardrails for Scalable Support in Programming Classes. InProceedings of the 23rd Koli Calling International Conference on Computing Education Research. ACM, Koli Fi...

  43. [51]

    Wenhan Lyu, Yimeng Wang, Yifan Sun, and Yixuan Zhang. 2025. Will Your Next Pair Programming Partner Be Human? An Empirical Evaluation of Generative AI as a Collaborative Teammate in a Semester-Long Classroom Setting. InProceedings of the Twelfth ACM Conference on Learning @ Sc...

  44. [52]

    MacKinnon, Morten Ørregaard Nielsen, and Matthew D

    James G. MacKinnon, Morten Ørregaard Nielsen, and Matthew D. Webb. 2023. Fast and reliable jackknife and bootstrap methods for cluster-robust inference.Journal of Applied Econometrics38, 5 (Aug. 2023), 671–694. doi:10.1002/jae.2969

  45. [53]

    Margulieux, James Prather, Brent N

    Lauren E. Margulieux, James Prather, Brent N. Reeves, Brett A. Becker, Gozde Cetin Uzun, Dastyni Loksa, Juho Leinonen, and Paul Denny. 2024. Self-Regulation, Self-Efficacy, and Fear of Failure Interactions with How Novices Use LLMs to Solve Programming Problems. InProceedings ...

  46. [54]

    Daye Nam, Andrew Macvean, Vincent Hellendoorn, Bogdan Vasilescu, and Brad Myers. 2024. Using an LLM to Help With Code Understanding. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering. ACM, Lisbon Portugal, 1–13. doi:10.1145/3597503.3639187

  47. [55]

    Akilan, P

    Palanichamy Naveen, T. Akilan, P. Manikandan, C. Swedheetha, and M. Saravanan. 2025. Comparative Review of Large Language Models. In2025 5th International Conference on Soft Computing for Security Applications (ICSCSA). IEEE, 2120–2124. https://ieeexplore.ieee.org/abstract/doc...

  48. [56]

    2021.The Extended Mind: The Power of Thinking Outside the Brain

    Annie Murphy Paul. 2021.The Extended Mind: The Power of Thinking Outside the Brain. Houghton Mifflin Harcourt, Boston, MA

  49. [57]

    Reinhard Pekrun. 2006. The Control-Value Theory of Achievement Emotions: Assumptions, Corollaries, and Implications for Educational Research and Practice.Educational Psychology Review18, 4 (Nov. 2006), 315–341. doi:10.1007/s10648-006-9029-9

  50. [58]

    Reinhard Pekrun. 2024. Control-Value Theory: From Achievement Emotion to a General Theory of Human Emotions.Educational Psychology Review36, 3 (Sept. 2024), 83. doi:10.1007/s10648-024-09909-7

  51. [59]

    Jacob Penney, Pawan Acharya, Peter Hilbert, Priyanka Parekh, Anita Sarma, Igor Steinmacher, and Marco Gerosa. 2025. Understanding Programming Students’ Help-Seeking Preferences in the Era of Generative AI. InProceedings of the ACM Global Computing Education Conference 2025 - V...

  52. [60]

    Reeves, Jaromir Savelka, IV Smith, David H., Sven Strickroth, and Daniel Zingaro

    James Prather, Juho Leinonen, Natalie Kiesler, Jamie Gorson Benario, Sam Lau, Stephen MacNeil, Narges Norouzi, Simone Opel, Vee Pettit, Leo Porter, Brent N. Reeves, Jaromir Savelka, IV Smith, David H., Sven Strickroth, and Daniel Zingaro. 2025. Beyond the Hype: A Comprehensive...

  53. [61]

    It’s Weird That it Knows What I Want

    James Prather, Brent N. Reeves, Paul Denny, Brett A. Becker, Juho Leinonen, Andrew Luxton-Reilly, Garrett Powell, James Finnie-Ansley, and Eddie Antonio Santos. 2024. “It’s Weird That it Knows What I Want”: Usability and Interactions with Copilot for Novice Programmers.ACM Tra...

  54. [62]

    Becker, Bailey Kimmel, Jared Wright, and Ben Briggs

    James Prather, Brent N Reeves, Juho Leinonen, Stephen MacNeil, Arisoa S Randrianasolo, Brett A. Becker, Bailey Kimmel, Jared Wright, and Ben Briggs. 2024. The Widening Gap: The Benefits and Harms of Generative AI for Novice Programmers. InProceedings of the 2024 ACM Conference...

  55. [63]

    Ira J. Roseman. 1996. Appraisal Determinants of Emotions: Constructing a More Accurate and Comprehensive Theory.Cognition & Emotion10, 3 (May 1996), 241–278. doi:10.1080/026999396380240 Manuscript submitted to ACM 22 Gardella et al

  56. [64]

    Ross, Fernando Martinez, Stephanie Houde, Michael Muller, and Justin D

    Steven I. Ross, Fernando Martinez, Stephanie Houde, Michael Muller, and Justin D. Weisz. 2023. The Programmer’s Assistant: Conversational Interaction with a Large Language Model for Software Development. InProceedings of the 28th International Conference on Intelligent User In...

  57. [65]

    James A. Russell. 1980. A circumplex model of affect.Journal of personality and social psychology39, 6 (1980), 1161. https://psycnet.apa.org/journals/ psp/39/6/1161/

  58. [66]

    Norsaremah Salleh, Emilia Mendes, and John Grundy. 2011. Empirical Studies of Pair Programming for CS/SE Teaching in Higher Education: A Systematic Literature Review.IEEE Transactions on Software Engineering37, 4 (July 2011), 509–525. doi:10.1109/TSE.2010.59

  59. [67]

    Sharpe and Ian Tyndall

    Benjamin T. Sharpe and Ian Tyndall. 2025. The Sustained Attention Paradox: A Critical Commentary on the Theoretical Impossibility of Perfect Vigilance.Cognitive Science49, 4 (2025), e70061. doi:10.1111/cogs.70061 _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/cogs.70061

  60. [68]

    Abdulhadi Shoufan. 2023. Can Students without Prior Knowledge Use ChatGPT to Answer Test Questions? An Empirical Study.ACM Trans. Comput. Educ.23, 4 (Dec. 2023), 45:1–45:29. doi:10.1145/3628162

  61. [69]

    Nancy Sommers. 1980. Revision Strategies of Student Writers and Experienced Adult Writers.College Composition & Communication31, 4 (1980), 378–388. doi:10.58680/ccc198015930 Type: Journal Article

  62. [70]

    Nancy Sommers. 1982. Responding to Student Writing.College Composition & Communication33, 2 (1982), 148–156. doi:10.58680/ccc198215854 Type: Journal Article

  63. [71]

    Son and Rajiv Sethi

    Lisa K. Son and Rajiv Sethi. 2006. Metacognitive Control and Optimal Learning.Cognitive Science30, 4 (July 2006), 759–774. doi:10.1207/ s15516709cog0000_74

  64. [72]

    Phil Steinhorst, Andrew Petersen, and Jan Vahrenhold. 2020. Revisiting Self-Efficacy in Introductory Programming. InProceedings of the 2020 ACM Conference on International Computing Education Research. ACM, Virtual Event New Zealand, 158–169. doi:10.1145/3372782.3406281

  65. [73]

    John Sweller. 2011. Cognitive load theory. InPsychology of learning and motivation. Vol. 55. Elsevier, 37–76. https://www.sciencedirect.com/science/ article/pii/B9780123876911000028

  66. [74]

    Ritzhaupt

    Karthikeyan Umapathy and Albert D. Ritzhaupt. 2017. A meta-analysis of pair-programming in computer programming courses: Implications for educational practice.ACM Transactions on Computing Education (TOCE)17, 4 (2017), 1–13

  67. [75]

    Smith IV, Mounika Padala, Christine Alvarado, Jamie Gorson Benario, and Leo Porter

    Annapurna Vadaparty, Daniel Zingaro, David H. Smith IV, Mounika Padala, Christine Alvarado, Jamie Gorson Benario, and Leo Porter. 2024. CS1-LLM: Integrating LLMs into CS1 Instruction. InProceedings of the 2024 on Innovation and Technology in Computer Science Education V. 1 (IT...

  68. [76]

    Marcel Valový and Alena Buchalcevova. 2023. The Psychological Effects of AI-Assisted Programming on Students and Professionals. In2023 IEEE International Conference on Software Maintenance and Evolution (ICSME). 385–390. doi:10.1109/ICSME58846.2023.00050 ISSN: 2576-3148

  69. [77]

    Vygotsky

    Lev S. Vygotsky. 1978.Mind in Society: The Development of Higher Psychological Processes. Harvard University Press, Cambridge, MA

  70. [78]

    Matthew D. Webb. 2023. Reworking wild bootstrap-based inference for clustered errors.Canadian Journal of Economics/Revue canadienne d’économique56, 3 (Aug. 2023), 839–858. doi:10.1111/caje.12661

  71. [79]

    Weisz, Michael Muller, Michael Muller, Michael Muller, Stephanie Houde, John T

    Justin D. Weisz, Michael Muller, Michael Muller, Michael Muller, Stephanie Houde, John T. Richards, Steven I. Ross, Fernando Martinez, Mayank Agarwal, and Kartik Talamadupula. 2021. Perfection Not Required? Human-AI Partnerships in Code Translation. 402–412. doi:10.1145/339748...

  72. [80]

    Bai, Robert Tairas, and Yu Huang

    Yuankai Xue, Hanlin Chen, Gina R. Bai, Robert Tairas, and Yu Huang. 2024. Does ChatGPT Help With Introductory Programming?An Experiment of Students Using ChatGPT in CS1. InProceedings of the 46th International Conference on Software Engineering: Software Engineering Education ...

  73. [81]

    Burak Yetistiren, Isik Ozsoy, and Eray Tuzun. 2022. Assessing the quality of GitHub copilot’s code generation. InProceedings of the 18th International Conference on Predictive Models and Data Analytics in Software Engineering. 62–71

  74. [82]

    Ramazan Yilmaz and Fatma Gizem Karaoglan Yilmaz. 2023. The effect of generative artificial intelligence (AI)-based tool use on students’ computational thinking skills, programming self-efficacy and motivation.Computers and Education: Artificial Intelligence4 (Jan. 2023), 10014...

  75. [83]

    Achim Zeileis. 2004. Econometric computing with HC and HAC covariance matrix estimators.Journal of statistical software11 (2004), 1–17. https://www.jstatsoft.org/article/view/v011i10/0

  76. [84]

    Achim Zeileis and Torsten Hothorn. 2002. Diagnostic Checking in Regression Relationships.R News2, 3 (2002), 7–10. https://CRAN.R- project.org/doc/Rnews/

  77. [85]

    Achim Zeileis, Susanne Köll, and Nathaniel Graham. 2020. Various versatile variances: an object-oriented implementation of clustered covariances in R.Journal of Statistical Software95 (2020), 1–36. https://www.jstatsoft.org/article/view/v095i01/0 Received 20 April 2026; revise...

Pith tools

Reviewed July 5, 2026 · model on record in the stance chip above.