REVIEW 3 major objections 6 minor 31 references
LLM tutor response styles track whether students continue productively, with only small global differences but much larger gaps under high cognitive load and in debugging.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 14:35 UTC pith:3RQ5YZCU
load-bearing objection Solid observational map of which LLM response styles track next-turn student engagement in real programming help dialogues; small global effects, clearer context variation, and one real but not fatal labeling caveat. the 3 major comments →
When LLM Tutoring Responses Work: Evidence from Student Programming Conversations
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Assistant response style is significantly associated with productive continuation and with unresolved continuation across the full set of interactions, though effect sizes are small. Verification feedback shows the highest productive-continuation rate (82.4 percent) and direct answers the lowest (62.7 percent). Effectiveness-score ranges are smallest in low-confusion conceptual contexts and largest in high-cognitive-load contexts; detailed comparisons further show that the same style is not followed by the same outcomes in every help-seeking situation.
What carries the argument
The effectiveness score ES = productive-continuation rate minus unresolved-continuation rate, computed per response style inside each help-seeking context (situation combined with student state), plus present-versus-absent risk differences tested with FDR-corrected Fisher exact tests. That score range is what reveals where style differences are small versus large.
Load-bearing premise
The central claim treats next-turn outcome labels produced mainly by an automated annotator, checked at 82 percent agreement on only one hundred samples, as reliable proxies for whether a reply actually helped the student continue productively.
What would settle it
Randomly assign response styles within matched student situations and states in a controlled tutoring study; if productive continuation, confusion decrease, and longer-term learning or debugging success show no reliable differences between verification feedback and direct answers—especially under high cognitive load and debugging—the claimed context-dependent association collapses.
If this is right
- Global rankings of LLM reply styles are weak guides for programming tutors; context must be part of evaluation.
- High cognitive load, high confusion, debugging, and code-request moments are where style choice is most consequential.
- Verification feedback and explanation-before-answer are more often followed by productive continuation than bare direct answers.
- Adaptive tutors that detect situation and state can select strategies rather than apply one fixed style.
- Next-turn engagement metrics should sit alongside answer correctness when judging LLM tutoring quality.
Where Pith is reading between the lines
- If next-turn productive continuation is a usable short proxy, real-time detectors of load and confusion could route high-load debugging turns away from direct answers toward stepwise or diagnostic replies.
- Small global effect sizes imply most practical gain will come from context-sensitive routing, not from crowning a single best style for every student.
- The same annotation-and-score pipeline could be run on other STEM tutoring dialogues to test whether high-load contexts again show the largest style ranges.
- Assignment design that surfaces debugging or code-generation modes could trigger a temporary shift in the tutor’s default reply policy.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper studies how ChatGPT tutoring response styles relate to students’ immediate next-turn engagement in authentic programming help-seeking dialogues. Using StudyChat, the authors construct 16,851 assistant-response interactions from 203 students, annotate student situation/state, assistant response style, and next-turn outcomes with Gemma 4 (human agreement 82%, κ=.74), and report global associations of response style with productive and unresolved continuation (χ² tests, small Cramer’s V) plus larger descriptive effectiveness-score ranges and FDR-corrected present/absent comparisons in high-load, high-confusion, debugging, and code-request contexts. The central claim is that response-style differences are statistically detectable but small globally, and that effectiveness is situation-dependent, supporting context-aware evaluation of AI tutors.
Significance. The question—what kind of LLM response is followed by productive student continuation in programming help—is timely and practically important for CS education and adaptive tutoring design. Strengths include use of a large naturalistic, course-embedded public dataset rather than synthetic prompts; transparent observational framing with reported effect sizes rather than overstated causal claims; FDR correction for many Fisher comparisons; and a clear move from global rankings to context-conditioned patterns (e.g., stepwise guidance vs. direct answers under high load/debugging). If the associations hold under stronger reliability and dependence-aware analysis, the work provides useful empirical priors for designing and evaluating context-sensitive LLM tutors. The contribution is primarily empirical association mining, not a new theory or causal identification strategy.
major comments (3)
- [Methods, Data Preparation; Results §4] All primary quantitative claims (global χ² associations, ranked productive/unresolved rates in Fig. 2, effectiveness-score ranges in Fig. 3, and Table 1 risk differences) rest on multi-class Gemma-4 labels for response style and next-turn outcomes. Human validation covers only 100 of 16,851 interactions (~0.6%), reports overall 82% agreement and κ=.74, and does not provide per-category agreement, confusion matrices, human–human reliability, or stratification by rare styles or high-load/debugging cells. Systematic category-specific noise could reorder the verification-feedback vs. direct-answer ranking and distort score ranges. Please expand validation (larger stratified sample), report category-level reliability, and ideally add a sensitivity check showing whether the main rankings and Table 1 patterns are stable under plausible label error.
- [Methods, Data Analysis; Results, Global Response-Style Patterns] The global χ² tests treat 16,851 assistant-response interactions as independent observations, but they come from only 203 students and 2,214 conversations. Within-student and within-conversation dependence will inflate χ² and understate uncertainty; with already small V (.078/.087), significance may be overstated. Please re-estimate associations with a dependence-aware approach (e.g., mixed-effects logistic models with student/conversation random effects, cluster-robust inference, or student-level bootstrap) and report whether the global style associations and key Table 1 contrasts remain after accounting for clustering.
- [Methods, Analysis Variables and Outcome Construction] Productive continuation is defined broadly (continuing the task, meaningful follow-up, applying the response, requesting verification, or indicating progress). Combined with generally high productive rates (often ~75–82%) and low unresolved rates, this breadth may compress style differences and make the outcome partly a measure of continued chat engagement rather than tutoring effectiveness. Please clarify coding decision rules with examples, report the prevalence of each productive sub-signal if available, and discuss how alternative, stricter operationalizations would change Fig. 2 and the effectiveness scores.
minor comments (6)
- [Methods, Data Analysis] State explicitly how many interactions lacked a following student turn and were excluded from outcome-rate calculations; selection into continued dialogue may bias productive/unresolved rates.
- [Figure 2] Figure 2 plots unresolved rates below zero for visual contrast; add a note in the caption that the negative axis is a display convention, not a signed rate.
- [Abstract; References [17]] The abstract and introduction cite StudyChat as UMass 2026 / McNichols et al. 2026 while the paper is framed for SIGCSE TS 2027; ensure citation status and dataset version are consistent and citable.
- [Methods, Data Preparation] Define or briefly exemplify secondary state labels (cognitive-load signal, context richness, answer-seeking level) so readers can judge construct validity without the annotation prompts.
- [Results, Global Response-Style Patterns / Figure 2] Report cell counts (n) for each response style in the global rate plot or an appendix table so readers can assess precision of the 82.4% vs. 62.7% comparison.
- [Abstract; Methods] Minor wording: abstract says “Gemma 4” while Methods says “Gemma4 with 26B parameters”—standardize the model name and cite the exact checkpoint if possible.
Circularity Check
No circularity: purely observational associations; ES = PCR − UCR is a transparent descriptive summary, not a forced prediction.
full rationale
The paper reports chi-square associations and descriptive rates between Gemma-4-annotated assistant response styles and next-turn student outcomes on the external StudyChat corpus. The only constructed quantities are the effectiveness score ES = PCR − UCR and the score range SR = max(ES) − min(ES); both are explicit arithmetic summaries of observed rates, not parameters fitted to data and then re-presented as predictions. There is no self-definitional equation, no fitted input called a prediction, no uniqueness theorem imported from the authors’ prior work, and no ansatz smuggled via self-citation. Related-work citations (including StudyChat itself) are to independent authors and serve as background, not load-bearing premises that force the reported contingency tables. The derivation chain is therefore self-contained observational statistics; any weakness lies in label reliability, not circular construction.
Axiom & Free-Parameter Ledger
free parameters (2)
- minimum cell size for score-range calculation =
50
- FDR correction threshold for Fisher comparisons
axioms (3)
- domain assumption Next-turn student message is a valid short-term proxy for whether an assistant response was productively helpful or left the issue unresolved.
- domain assumption Gemma-4 structured JSON labels for situation, state, response style, and next-turn outcome are sufficiently accurate for aggregate statistical analysis after 82%/κ=.74 human validation on 100 samples.
- ad hoc to paper The six focal help-seeking situations and seven response styles exhaust the pedagogically meaningful distinctions for the reported comparisons.
invented entities (2)
-
effectiveness score ES = productive-continuation rate − unresolved-continuation rate
no independent evidence
-
help-seeking context (situation × student-state combination)
no independent evidence
read the original abstract
As students increasingly use LLM tutors in computer science education, one question becomes especially important: what kind of response helps a student continue productively? Prior work has studied how students use LLMs in computer science education, but less is known about how tutoring response styles are associated with student follow-up across programming help-seeking contexts. This paper analyzes StudyChat (UMass, 2026), a public dataset of student and ChatGPT tutoring conversations from an artificial intelligence course. We transformed StudyChat into 16,851 assistant-response interactions from 203 students and 2,214 conversations. Using local LLM-assisted annotation with Gemma 4, we labeled student help-seeking situations, student state, assistant response style, and student next-turn outcome. Human validation showed 82\% agreement with the LLM-assisted labels (Cohen's $\kappa=.74$). We analyzed productive continuation and unresolved continuation across the full dataset and across help-seeking contexts. Globally, response style was significantly associated with productive continuation, $\chi^2(7)=100.39$, $p<.001$, $V=.078$, and unresolved continuation, $\chi^2(7)=125.77$, $p<.001$, $V=.087$, though effect sizes were small. Verification feedback had the highest productive-continuation rate (82.4\%), while direct answers had the lowest (62.7\%). Descriptively, response-style score ranges were smallest in low-confusion conceptual contexts (.017) and largest in high-cognitive-load contexts (.203). More detailed comparisons showed situation-dependent response patterns. For example, stepwise guidance was followed by greater confusion decrease in high-cognitive-load code requests, while direct answers were followed by more unresolved continuation in high-load debugging. These findings support context-aware evaluation and design of AI tutoring responses for programming education.
Figures
Reference graph
Works this paper leans on
-
[1]
Vincent Aleven, Ido Roll, Bruce M. McLaren, and Kenneth R. Koedinger. 2016. Help Helps, But Only So Much: Research on Help Seeking with Intelligent Tu- toring Systems.International Journal of Artificial Intelligence in Education26, 1 (2016), 205–223. doi:10.1007/s40593-015-0089-1
-
[2]
Isaac Alpizar-Chacon and Hieke Keuning. 2025. Student’s Use of Generative AI as a Support Tool in an Advanced Web Development Course. InProceedings of the 30th ACM Conference on Innovation and Technology in Computer Science Education V. 1 (ITiCSE 2025). 312–318. doi:10.1145/3724363.3729106
-
[3]
Patrick Bassner, Eduard Frankford, and Stephan Krusche. 2024. Iris: An AI- Driven Virtual Tutor for Computer Science Education. InProceedings of the 2024 Innovation and Technology in Computer Science Education V. 1 (ITiCSE 2024). 394–400. doi:10.1145/3649217.3653543
-
[4]
Anael Kuperwajs Cohen, Alannah Oleson, and Amy J. Ko. 2024. Factors Influenc- ing the Social Help-Seeking Behavior of Introductory Programming Students in a Competitive University Environment.ACM Transactions on Computing Education 24, 1 (2024), 11:1–11:27. doi:10.1145/3639059
-
[5]
Paul Denny, Viraj Kumar, and Nasser Giacaman. 2023. Conversing with Copilot: Exploring Prompt Engineering for Solving CS1 Problems Using Natural Lan- guage. InProceedings of the 54th ACM Technical Symposium on Computer Science Education V. 1 (SIGCSE ’23). 1136–1142. doi:10.1145/3545945.3569823
-
[6]
Paul Denny, Stephen MacNeil, Jaromir Savelka, Leo Porter, and Andrew Luxton- Reilly. 2024. Desirable Characteristics for AI Teaching Assistants in Programming Education. InProceedings of the 2024 Innovation and Technology in Computer Science Education V. 1 (ITiCSE 2024). 408–414. doi:10.1145/3649217.3653574
-
[7]
Sidney D’Mello, Blair Lehman, Reinhard Pekrun, and Art Graesser. 2014. Confu- sion Can Be Beneficial for Learning.Learning and Instruction29 (2014), 153–170. doi:10.1016/j.learninstruc.2012.05.003
-
[8]
Matthew Frazier, Kostadin Damevski, and Lori Pollock. 2024. Customizing Chat- GPT to Help Computer Science Principles Students Learn Through Conversation. InProceedings of the 2024 Innovation and Technology in Computer Science Education V. 1 (ITiCSE 2024). 633–639. doi:10.1145/3649217.3653570
-
[9]
John Hattie and Helen Timperley. 2007. The Power of Feedback.Review of Educational Research77, 1 (2007), 81–112. doi:10.3102/003465430298487
-
[10]
Arto Hellas, Juho Leinonen, Sami Sarsa, Charles Koutcheme, Lilja Kujanpää, and Juha Sorva. 2023. Exploring the responses of large language models to beginner programmers’ help requests. InProceedings of the 2023 ACM Conference on International Computing Education Research-Volume 1. 93–105. doi:10.1145/ 3568813.3600139
arXiv 2023
-
[11]
Henley, Paul Denny, Michelle Craig, and Tovi Grossman
Majeed Kazemitabaar, Runlong Ye, Xiaoning Wang, Austin Z. Henley, Paul Denny, Michelle Craig, and Tovi Grossman. 2024. CodeAid: Evaluating a Classroom Deployment of an LLM-based Programming Assistant that Balances Student and Educator Needs. InProceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24). 650:1–650:20. doi:10.1145/3613...
-
[12]
Natalie Kiesler, Dominic Lohr, and Hieke Keuning. 2023. Exploring the potential of large language models to generate formative programming feedback. In2023 IEEE frontiers in education conference (FIE). IEEE, 1–5. doi:10.1109/FIE58773.2023. 10343457
-
[13]
Natalie Kiesler and Daniel Schiffner. 2023. Large Language Models in Introductory Programming Education: ChatGPT’s Performance and Implications for Assess- ments.arXiv preprint arXiv:2308.08572(2023). doi:10.48550/arXiv.2308.08572
-
[14]
Oka Kurniawan, Cyrille Jégourel, Norman Tiong Seng Lee, Matthieu De Mari, and Christopher M. Poskitt. 2022. Steps Before Syntax: Helping Novice Programmers Solve Problems Using the PCDIT Framework. InProceedings of the Annual Hawaii International Conference on System Sciences. Hawaii International Conference on System Sciences. doi:10.24251/HICSS.2022.121
-
[15]
Rongxin Liu, Carter Zenke, Charlie Liu, Andrew Holmes, Patrick Thornton, and David J. Malan. 2024. Teaching CS50 with AI: Leveraging Generative Artificial Intelligence in Computer Science Education. InProceedings of the 55th ACM Technical Symposium on Computer Science Education V. 1 (SIGCSE 2024). 750–756. doi:10.1145/3626252.3630938
-
[16]
Jakub Macina, Nico Daheim, Sankalan Pal Chowdhury, Tanmay Sinha, Manu Kapur, Iryna Gurevych, and Mrinmaya Sachan. 2023. MathDial: A Dialogue Tutoring Dataset with Rich Pedagogical Properties Grounded in Math Reasoning Problems. InFindings of the Association for Computational Linguistics: EMNLP
2023
-
[17]
doi:10.18653/v1/2023.findings-emnlp.372
5602–5621. doi:10.18653/v1/2023.findings-emnlp.372
-
[18]
Hunter McNichols, Fareya Ikram, and Andrew S. Lan. 2026. The StudyChat Dataset: Analyzing Student Dialogues With ChatGPT in an Artificial Intelligence Course. InProceedings of the 16th International Learning Analytics and Knowledge Conference (LAK 2026). 53–63. doi:10.1145/3785022.3785029
-
[19]
Valeria Ramirez Osorio, Angela Zavaleta Bernuy, Bogdan Simion, and Michael Liut. 2025. Understanding the Impact of Using Generative AI Tools in a Database Course. InProceedings of the 56th ACM Technical Symposium on Computer Science Education V. 1 (SIGCSE TS 2025). 959–965. doi:10.1145/3641554.3701785
-
[20]
Fred Paas, Alexander Renkl, and John Sweller. 2003. Cognitive Load Theory and Instructional Design: Recent Developments.Educational Psychologist38, 1 (2003), 1–4. doi:10.1207/S15326985EP3801_1
-
[21]
Jacob Penney, Pawan Acharya, Peter Hilbert, Priyanka Parekh, Anita Sarma, Igor Steinmacher, and Marco Aurélio Gerosa. 2025. Understanding Programming Students’ Help-Seeking Preferences in the Era of Generative AI. InCompEd 2025: Proceedings of the ACM Global Computing Education Conference. 15–21. doi:10.1145/3736181.3747165
-
[22]
Tung Phung, Heeryung Choi, Mengyan Wu, Christopher Brooks, Sumit Gulwani, and Adish Singla. 2026. Closing the Loop: An Instructor-in-the-Loop AI Assis- tance System for Supporting Student Help-Seeking in Programming Education. In Proceedings of the 57th ACM Technical Symposium on Computer Science Education V. 1 (SIGCSE 2026). 852–858. doi:10.1145/3770762.3772612
-
[23]
Price, Zhongxiu Liu, Veronica Cateté, and Tiffany Barnes
Thomas W. Price, Zhongxiu Liu, Veronica Cateté, and Tiffany Barnes. 2017. Factors Influencing Students’ Help-Seeking Behavior while Programming with Human and Computer Tutors. InProceedings of the 2017 ACM Conference on International Computing Education Research (ICER ’17). 127–135. doi:10.1145/ 3105726.3106179
arXiv 2017
-
[24]
Lianne Roest, Hieke Keuning, and Johan Jeuring. 2024. Next-step hint generation for introductory programming using large language models. InProceedings of the 26th Australasian Computing Education Conference. 144–153. doi:10.1145/3636243. 3636259
-
[25]
Griswold, and Adalbert Gerald Soosai Raj
Anshul Shah, Anya Chernova, Elena Tomson, Leo Porter, William G. Griswold, and Adalbert Gerald Soosai Raj. 2025. Students’ Use of GitHub Copilot for Working with Large Code Bases. InProceedings of the 56th ACM Technical Symposium on Computer Science Education V. 1 (SIGCSE TS 2025). 1050–1056. doi:10.1145/3641554.3701800
-
[26]
Brad Sheese, Mark Liffiton, Jaromir Savelka, and Paul Denny. 2024. Patterns of student help-seeking when using a large language model-powered programming assistant. InProceedings of the 26th Australasian computing education conference. 49–57. doi:10.1145/3636243.3636249
-
[27]
Valerie J. Shute. 2008. Focus on Formative Feedback.Review of Educational Research78, 1 (2008), 153–189. doi:10.3102/0034654307313795
-
[28]
John Sweller. 1988. Cognitive Load During Problem Solving: Effects on Learning. Cognitive Science12, 2 (1988), 257–285. doi:10.1207/s15516709cog1202_4
-
[29]
David Wood, Jerome S. Bruner, and Gail Ross. 1976. The Role of Tutoring in Problem Solving.Journal of Child Psychology and Psychiatry17, 2 (1976), 89–100. doi:10.1111/j.1469-7610.1976.tb00381.x
-
[30]
Stephanie Yang, Hanzhang Zhao, Yudian Xu, Karen Brennan, and Bertrand Schnei- der. 2024. Debugging with an AI Tutor: Investigating Novice Help-seeking Behaviors and Perceived Learning. InProceedings of the 2024 ACM Confer- ence on International Computing Education Research V. 1 (ICER 2024). 84–94. doi:10.1145/3632620.3671092
-
[31]
J. D. Zamfirescu-Pereira, Laryn Qi, Björn Hartmann, John DeNero, and Narges Norouzi. 2025. 61A Bot Report: AI Assistants in CS1 Save Students Homework Time and Reduce Demands on Staff. (Now What?). InProceedings of the 56th ACM Technical Symposium on Computer Science Education V. 1 (SIGCSE TS 2025). 1309–1315. doi:10.1145/3641554.3701864
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.