{"id":"3f55bb14-3c8f-41a6-b9ea-5e48e9c1da8f","arxiv_id":"2506.20433","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In an eye-tracking study of 11 students, novices fixated twice as long on AI-generated feedback as experienced peers, relied on it instead of compiler output, and could not comprehend about 20% of the AI feedback they received.","lead":"Eleven students solved Python tasks in a custom tutor that offers both AI-generated and compiler feedback, while eye-tracking and think-aloud recordings captured what they read. Beginners spent about twice as long looking at the AI feedback, often skipped the compiler messages, and failed to understand about one in five AI responses.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Gaze-to-AoI mapping is the least secure link in the evidence chain: for three of eleven participants the mapping was manual frame-by-frame (Section 4.2.1), so the headline fixation percentages and the read/not-read classifications that underlie both RQs could shift with small mapping errors.","rationale":"The paper is an honest, triangulated exploratory study, but its most novel and consequential claims are quantitative: the large visual attention to GenAI feedback, the twofold difference between experience groups, and the markedly higher comprehension failures in the inexperienced group. These claims all flow through the eye-tracking pipeline, and the weakest point in that pipeline is the mapping of gaze to screen regions. Three of eleven participants (S01, S02, and S10) had to be mapped manually frame-by-frame because an equipment switch forced the use of a different tracker. The paper does not report any reliability check for those mappings, and the coding of helpfulness explicitly conditions on 'read (gaze visualization)'—so any mapping error cascades into the helpfulness and comprehension numbers, not just the attention percentages. I agree with the reader's identification of this as the weakest assumption. I considered two alternative concerns: the lack of inferential statistics (which is less damaging for an exploratory descriptive study) and the possibility that fixation times were not normalized by feedback-display duration (which would affect the attention comparison but not the event-level read/not-read analysis). The manual mapping concern is more load-bearing because it threatens both RQs and is concrete. A second-coder re-annotation plus a perturbation sensitivity analysis would settle it. If the group differences persist under that check, the central claims stand on much firmer ground; if they vanish, the paper's headline should be substantially softened. Since the study is explicitly exploratory and the authors already list threats, CONDITIONAL remains the appropriate verdict.","tokens_in":11955,"tokens_out":7976,"duration_ms":85377,"concrete_test":"Re-annotate the gaze recordings of S01, S02, and S10 with a second, independent coder using the same frame-by-frame procedure, and additionally report a sensitivity analysis in which 5% and 10% of fixations in these sessions are randomly re-assigned to the adjacent AoIs. If, under either the second annotation or the perturbation, the group-level GenAI fixation-time gap (30.71% vs 15.49%) and the comprehension-failure counts (22 vs 1) move by more than the original between-group gap, then the central claims are not robust to mapping uncertainty.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central descriptive claims—that GenAI feedback received 23.79% of fixation time, that inexperienced students fixated nearly twice as long (30.71% vs 15.49%), and that 22/107 vs 1/64 GenAI messages were not comprehended—all depend on mapping gaze recordings to the four Areas of Interest. Section 4.2.1 reports that for participants S01, S02, and S10 the gaze data were mapped manually, frame by frame, because a different eye tracker (Tobii Eye Tracker 5 vs Tobii Pro Glasses 3) had to be used. Manual frame-by-frame mapping of video recordings to screen regions is subject to drift, misalignment, and coder judgment, especially if the participant moves or the tracker coordinate systems differ. No validation, inter-coder agreement, or sensitivity analysis for these three participants is reported. The consequence is not limited to the attention percentages: the coding scheme in Section 4.2.2 defines 'Has helped,' 'Has not helped,' and 'Has not been read' with the precondition that the feedback 'was read (gaze visualization).' Thus any mapping error can reclassify an instance across these categories, changing the helpfulness rates (60.9% vs 43.0%) and the comprehension-failure counts (22 vs 1). Because S01 is in the inexperienced group and S02/S10 are in the experienced group, systematic mapping errors would not cancel across the comparison groups. The paper frames the finding as a robust group difference, but with n=6 vs n=5 and one manual participant in one group and two in the other, the result is not yet protected against this specific threat. This is the load-bearing concern because it sits upstream of both RQs.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a mixed-methods lab study with 11 undergraduate students using the authors' custom web application, Tutor Kai, which presents Python programming tasks, a code editor, GenAI feedback from GPT-4 Turbo, and compiler feedback. Eye-tracking, think-aloud protocols, and semi-structured interviews are used to investigate (RQ1) how much attention learners pay to GenAI feedback relative to compiler feedback and (RQ2) to what extent the GenAI feedback is helpful, with comparisons between inexperienced (n=6) and experienced (n=5) students. The headline descriptive findings are that GenAI feedback received 23.79% of overall fixation time, that inexperienced students fixated on it nearly twice as long as experienced students (30.71% vs. 15.49%), that inexperienced students requested GenAI feedback 107 times vs. 64 times for experienced students, and that GenAI feedback was coded as helpful in 43.0% of cases for inexperienced students vs. 60.9% for experienced students. The authors also report that inexperienced students failed to comprehend 22 of 107 GenAI feedback messages, versus 1 of 64 for experienced students, and that compiler feedback was frequently not read by inexperienced students. The paper concludes with concerns about over-reliance on GenAI feedback and suggestions for adaptive feedback design.","tokens_in":12256,"tokens_out":5888,"duration_ms":63325,"significance":"If the descriptive findings hold, the paper makes a useful empirical contribution to computing education research by directly measuring visual attention to GenAI feedback in a realistic programming-task environment and by triangulating gaze with think-aloud and code-change evidence. Strengths include a transparent coding scheme that links gaze, verbalization, and code improvements; the availability of an open repository with the prompt, tasks, and interview guide; and an explicit threats-to-validity section that acknowledges think-aloud disruption, eye-tracker distraction, and LLM nondeterminism. The main limitation is that the gaze-to-AoI mapping, which underpins both research questions, is insecure for three of eleven participants and is not validated; this tempers the strength of the quantitative comparisons. The paper is best read as an exploratory descriptive study rather than as providing robust statistical evidence of group differences.","major_comments":[{"comment":"The central quantitative claims depend on the gaze-to-AoI mapping, but for participants S01, S02, and S10 this mapping was done manually, frame by frame, after a hardware change to a different Tobii tracker, and no validation, inter-coder agreement, drift correction, or sensitivity analysis is reported. Because the Section 4.2.2 coding scheme defines 'Has helped,' 'Has not helped,' and 'Has not been read' using the precondition that the feedback 'was read (gaze visualization),' any mapping error can reclassify an instance across the helpfulness categories. The three manually mapped participants are distributed across both comparison groups (S01 in the inexperienced group; S02 and S10 in the experienced group), so systematic error would not cancel. Please report a robustness check (e.g., recomputing the headline fixation percentages and helpfulness rates with S01/S02/S10 excluded) or provide evidence of mapping reliability.","section":"Section 4.2.1; Section 4.2.2; Section 5.2; Section 5.3"},{"comment":"The count of 'has not helped' instances is internally inconsistent. The text states that the analysis covers 77 such instances, but the four listed categories sum to 78 (25 + 23 + 17 + 13). In addition, the relationship between this count, the stated 49.7% helpfulness rate for 171 GenAI feedback outputs, and any 'has not been read' classifications is not fully specified. Please reconcile these numbers and clarify the denominator for each category.","section":"Section 5.3, paragraph beginning 'Our analysis of the 77 instances'"},{"comment":"The helpfulness coding is load-bearing for RQ2, but no inter-rater reliability or second-coder check is reported. The categories require judgment about whether a code change 'directly related' to provided information and whether a verbalization indicates help, so coder subjectivity is nontrivial. Please report agreement on at least a subset of the 171 GenAI feedback instances and the 287 compiler feedback instances, or justify why a single-coder scheme is sufficient in this context.","section":"Section 4.2.2"},{"comment":"The paper makes comparative claims such as 'nearly twice as much fixation time,' 'almost four times more helpful,' and 'substantial differences between the two groups' without any confidence intervals or inferential statistics. With n=6 versus n=5 and highly variable per-student behavior, these percentage differences may be fragile. At minimum, the claims should be reframed as exploratory descriptive observations, or accompanied by effect sizes and uncertainty estimates.","section":"Section 5.2; Section 5.3"}],"minor_comments":[{"comment":"The text in Section 5.3 reports an average feedback helpfulness rating of 8.59, but the eleven values in Table 2 sum to 95, giving a mean of 8.64; please correct the stated average or the table.","section":"Table 2"},{"comment":"The caption contains a typo: 'Fixiated' should be 'Fixated'.","section":"Table 1 caption"},{"comment":"The sentence 'They took about 40 minutes to complete all the tasks' has an unclear antecedent; specify that the two student tutors took about 40 minutes.","section":"Section 4.1.2"},{"comment":"The statement 'With the sample size of 11 students and 171 GenAI feedbacks, we have reached a large enough sample [4]' is overgeneralized. The cited source supports adequacy for qualitative analysis, not for the quantitative comparative claims; please temper the wording.","section":"Section 7"},{"comment":"Temperature=0 reduces but does not eliminate output variability; Section 7's statement that LLM output generation 'remains intransparent' is fine, but the earlier 'maximize output consistency' wording could be softened.","section":"Section 3.2"},{"comment":"In the last sentence of Section 5.2, 'experience students' should be 'experienced students'.","section":"Section 5.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope for a computing education research venue and the empirical setup is appropriate. The main concern is the unvalidated manual gaze mapping for three of eleven participants; this is fixable with a sensitivity analysis or explicit exclusion analysis. The count inconsistency in Section 5.3 should also be corrected. I do not see grounds for rejection, but the quantitative claims need to be either better supported or more carefully framed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look. Jacobs et al. ran an 11-student eye-tracking and think-aloud study comparing how much visual attention GenAI feedback gets versus compiler feedback in their own Tutor Kai, and whether that feedback actually helps. The central descriptive findings—novices fixate nearly twice as long on GenAI feedback (30.71% vs. 15.49%), request it more often (107 vs. 64), understand it less (22/107 not comprehended vs. 1/64), and often skip compiler feedback (39.5% not read)—are new and plausible. This is one of the first direct gaze-based comparisons of GenAI vs. compiler feedback in a programming tutor, split by experience, and that is a genuine empirical contribution.\n\nWhat the paper does well: the methodology is transparent. The coding scheme for “helped / not helped / not read” is specified, and it triangulates gaze with think-aloud and actual code changes. The authors put their data and materials on OSF, which is rare and appreciated. They also openly list threats to validity, which makes the work feel honest even where it is limited.\n\nThe soft spots are real but not disqualifying. First, the gaze-to-AoI mapping for participants S01, S02, and S10 was done manually, frame-by-frame, after an equipment switch to a different Tobii tracker. No inter-coder agreement or sensitivity analysis is reported. Since the “read” precondition drives the helpfulness coding, small mapping errors could shift the fixation percentages and the read/not-read classifications. For a group comparison with n=6 vs. n=5, and one manual participant in one group but two in the other, this is a legitimate concern—though the reported group differences are large enough that I doubt they would vanish entirely. Second, there are no inferential statistics or confidence intervals anywhere. The paper makes comparative claims (“nearly twice as much”) without quantifying uncertainty; that is fine for an exploratory study, but the authors should say so clearly and avoid over-strong wording. Third, the “large enough sample” statement in Section 7, citing Boddy, overstates things: with 11 participants and no effect sizes, “large enough” is not supported.\n\nThese are all addressable in revision. The study is a solid exploratory contribution for computing education researchers and designers of GenAI tutors. I would send it to review; it deserves referee time. My recommendation: engage with it, and ask the authors to report validation of the manual gaze mapping, add basic error bars or effect sizes, and tone down the sample-size claim. The core insight—that novices lean heavily on GenAI feedback but often cannot comprehend it—is worth having in the literature.","headline":"A small but honest eye-tracking study of GenAI vs. compiler feedback in a custom tutor; the descriptive findings are plausible and new, but the unvalidated manual gaze mapping for 3 of 11 participants is the soft spot to press on in revision.","tokens_in":12803,"tokens_out":2226,"would_cite":true,"duration_ms":26325,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"In a Python tutoring environment, students without prior programming experience spent nearly twice as much of their visual attention on feedback generated by a large language model as experienced students did, yet they understood it less…","keywords":["Programming education","Feedback","Large Language Models","Generative AI","GenAI","Eye-tracking","Think-aloud","Learner experience"],"falsifier":"Re-analyze or replicate the eye-tracking data with a single automated gaze-mapping setup and a larger sample. If, excluding the three manually mapped participants, the gap between inexperienced (30.71%) and experienced (15.49%) fixation time on GenAI feedback shrinks to statistical noise, or if a larger study finds no comprehension difference (22 of 107 versus 1 of 64), then the experience-split claim would be an artifact of the mapping procedure or the small sample.","tokens_in":11741,"feed_emoji":"👀","tokens_out":5571,"duration_ms":51068,"temperature":0.7,"pith_summary":"This paper tries to show how students actually engage with feedback generated by a large language model while solving Python tasks, rather than how good the feedback looks to experts. Using eye-tracking, think-aloud verbalization, and follow-up interviews with 11 undergraduates in the custom Tutor Kai environment, the authors measure attention to the GenAI feedback pane versus the compiler output pane. The central result is that GenAI feedback draws far more visual attention than compiler feedback, and the pattern splits by prior experience: students without prior programming experience fixated on the AI feedback nearly twice as long, requested it more often, understood it less, and benefited less from it, while often skipping the compiler output entirely. A sympathetic reader would care because this suggests that the value of GenAI feedback depends on the learner's background, and that naive reliance on it may widen, rather than close, the gap between prepared and unprepared students.","feed_headline":"Novices lean hardest on AI feedback but understand it least","feed_subtitle":"Eye-tracking of 11 students in a Python tutor: novices spent 30.7% of gaze time on AI feedback vs 15.5% for experienced peers.","key_machinery":"The apparatus that carries the argument is a purpose-built web application, the Tutor Kai, whose screen is divided into four Areas of Interest: task description, code editor, GenAI feedback, and compiler feedback. Gaze recordings from eye-tracking glasses are mapped onto these areas to yield fixation times, and the synchronized gaze video is triangulated with think-aloud transcripts to classify each feedback instance as 'has helped,' 'has not helped,' or 'has not been read.' This pairing of where students look with what they say and do is what lets the authors distinguish attention from comprehension and from actual problem-solving benefit.","core_discovery":"The study's central claim is that in the Tutor Kai, engagement with GenAI feedback is high overall but unequal across experience levels. Specifically, the GenAI feedback captured 23.79% of total fixation time compared to 7.00% for compiler feedback, and inexperienced students spent 30.71% of their fixation time on the GenAI feedback versus 15.49% for experienced students. Inexperienced students also requested the AI feedback more often (107 versus 64 requests), were helped by it less often (43.0% versus 60.9% of the time), and failed to comprehend 22 of 107 messages compared to 1 of 64 for experienced students. The paper argues these numbers show that GenAI feedback can support problem-solving, but only when learners have enough foundational knowledge to interpret it, and that feedback systems need to adapt to learners' prior knowledge.","pith_inferences":["A testable extension would be to run the same Tutor Kai study with adaptive feedback that expands unfamiliar terms and provides syntax examples, predicting that the comprehension gap between experience groups narrows and the 'has helped' rate for novices rises.","The eye-tracking result suggests that in naturalistic settings, novices may be spending their most productive problem-solving time reading fluent but under-specific AI explanations instead of iterating with the compiler, which could slow the development of debugging skills over a full course.","One implicit consequence is that LLM-generated feedback should perhaps be gated by a capability check, e.g., pop-up definitions when the feedback mentions concepts outside the student's demonstrated vocabulary, since the paper's category analysis shows 'incomprehensible' and 'missing syntax example' failures are both addressable at generation time.","The authors' claim that the sample is representative rests on a course survey in which about half of 97 respondents reported no prior knowledge; a larger replication with more than 11 participants would be needed to see whether the 30.71% versus 15.49% fixation split generalizes beyond this single course."],"forward_implications":["If the central claim is right, a student's prior programming experience is a decisive factor in whether GenAI feedback helps or confuses, so feedback tools must adapt explanations to the learner's level.","Inexperienced students' habit of requesting GenAI feedback before reading compiler output means that adding LLM-based explanations of compiler errors alone will not help the students who skip the compiler pane entirely.","The finding that 22 of 107 AI messages were incomprehensible to novices, often because they referenced concepts like loops without defining them, implies that feedback prompts should be constrained to avoid unexplained terminology and to include concrete code examples.","A feedback system that withholds or delays AI feedback until the student has engaged with compiler output could push novices toward the interpretation skills they currently bypass.","Because students rated the AI feedback highly (average 8.59/10) even when it did not objectively help, perceived helpfulness cannot be used as a proxy for actual learning support; tools should be evaluated with behavioral measures."],"supporting_citations":[{"why":"the prior study of novice programmers using GenAI whose 'illusion of competence' and widening-gap findings this work extends.","marker":"[34]"},{"why":"the study showing prior knowledge shapes novice outcomes with GenAI, which motivates the experience-group split.","marker":"[21]"},{"why":"the scoping review that identifies the gap this paper fills: how students actually engage with GenAI feedback.","marker":"[45]"},{"why":"the analysis of novice chat protocols with ChatGPT that this study complements with eye-tracking data.","marker":"[41]"},{"why":"the expert evaluation of GPT-4 feedback quality that establishes the baseline for 'correct and complete' feedback this paper's helpfulness results are compared against conceptually.","marker":"[1]"},{"why":"the survey of eye-tracking methods in programming research that justifies the fixation-time measure.","marker":"[32]"}],"fun_headline_variants":["AI feedback grabs novices' attention but escapes their grasp","Novices stare at AI feedback, then fail to use it","AI feedback: Novices look twice as long, get half the help","In AI feedback, novices spend more gaze, understand less","Novices over-rely on AI feedback they often can't parse"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole attention comparison rests on assuming that where students fix their gaze, measured by eye-tracking and mapped onto screen regions, accurately reflects what they are actually reading and processing; for three of the eleven participants that gaze mapping was done manually frame by frame after a tracker change, so small mapping errors could shift the reported fixation percentages and the 'read' versus 'not read' classifications.","fun_headline_variants_meta":{"raw":{"variants":["AI feedback grabs novices' attention but escapes their grasp","Novices stare at AI feedback, then fail to use it","AI feedback: Novices look twice as long, get half the help","In AI feedback, novices spend more gaze, understand less","Novices over-rely on AI feedback they often can't parse"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000599,"raw_usage":{"total_tokens":2816,"prompt_tokens":975,"completion_tokens":1841,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":591,"completion_tokens_details":{"reasoning_tokens":1753}},"tokens_in":591,"tokens_out":1841,"duration_ms":14970,"temperature":1.0,"reasoning_tokens":1753,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:48:07.578412+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-analyze or replicate the eye-tracking data with a single automated gaze-mapping setup and a larger sample. If, excluding the three manually mapped participants, the gap between inexperienced (30.71%) and experienced (15.49%) fixation time on GenAI feedback shrinks to statistical noise, or if a larger study finds no comprehension difference (22 of 107 versus 1 of 64), then the experience-split claim would be an artifact of the mapping procedure or the small sample.","supporting_citations":[{"cited_title":"Becker, Bailey Kimmel, Jared Wright, and Ben Briggs","cited_arxiv_id":null,"evidence_quote":"the prior study of novice programmers using GenAI whose 'illusion of competence' and widening-gap findings this work extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"the study showing prior knowledge shapes novice outcomes with GenAI, which motivates the experience-group split."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"the scoping review that identifies the gap this paper fills: how students actually engage with GenAI feedback."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"the survey of eye-tracking methods in programming research that justifies the fixation-time measure."}],"review_version":1}