{"id":"3a14e926-cdfc-4cce-89bf-a4dddc4a37f9","arxiv_id":"2412.14732","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Computing educators are adopting GenAI faster than they are formalizing policies, and both educators and developers see code reading, evaluation, and problem decomposition as rising in importance over syntax recall.","lead":"This report reviews the 2024 state of generative AI in computing education, combining a systematic review of 71 studies, surveys of 76 educators and 39 developers, and 17 interviews. It maps how instructors are integrating AI tools and how they think programming skills are changing.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline adoption percentages rest on small, self-selected convenience samples; a probability-sample replication could shift the 77.6%, 35.5%, and 79.5% estimates enough to undermine the report's central 'on the ground' claims.","rationale":"The central claim is empirical and descriptive: the report purports to say what educators and developers actually do. That claim stands or falls on whether the survey and interview samples represent the populations. I agree with the reader's weakest_assumption; my independent reading did not find a second issue more central. The SLR is carefully reported, the OSF data release for the SLR is a point in the paper's favor, and the discussion of threats is unusually candid. However, the candor does not remove the problem: the recruiting channels (Section 4.5) are exactly the channels through which GenAI-engaged educators are reached, and the small Ns make the headline percentages fragile. A probability-sample replication would settle the question. Because the paper already frames its results as a working-group snapshot and the existing verdict is CONDITIONAL, I do not see grounds to move to ACCEPT or REJECT; the condition should remain.","tokens_in":48298,"tokens_out":7088,"duration_ms":64055,"concrete_test":"Run a new probability-sample survey of computing instructors stratified by institution type (e.g., 2-year, 4-year, research university) and drawn from a comprehensive frame such as institutional CS faculty directories, asking the same ES-1 and ES-2 items. If the resulting 77.6% and 35.5% estimates fall outside the probability sample's 95% confidence intervals, the convenience-sample concern is confirmed; if they fall inside, the concern is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing premise is that the 76 educators, 39 developers, and 17 interviewees recruited through the authors' mailing lists, conference contacts, LinkedIn, and personal networks (Sections 4.5-4.6) can support descriptive claims about 'what is happening on the ground' in computing classrooms. The selection mechanism is correlated with the outcome: CSEd mailing-list subscribers, ITiCSE attendees, and people who forward a GenAI survey are more likely than typical instructors to have formed a considered GenAI policy, and the 100 incomplete responses plus 33 test responses removed before analysis (Section 4.6) introduce nonresponse selection. Headline numbers such as '77.6% do not explicitly disallow' (59/76), '35.5% actively incorporate' (27/76), and '79.5% of developers use GenAI' (31/39) are presented without confidence intervals and are treated as community-level facts in the key takeaways box. The authors' Section 7.2 limitation ('respondents may be more involved in CSEd than their peers'; 'large tech companies underrepresented') names the same worry, but the worry remains unresolved because no sensitivity analysis or weighting is provided. If a broader, less GenAI-engaged sample yields materially different adoption rates, the report's empirical core is not an accurate picture of the ground.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This ITiCSE working group report combines a systematic literature review of 71 papers, an international survey of 76 computing educators and 39 industry developers, and 17 semi-structured interviews to describe how generative AI is being integrated into computing education and how programming competencies are perceived to be changing. The report triangulates these sources around ten research questions and distills the results into a key-takeaways box covering adoption rates, assessment changes, competency shifts, and a recommendation that instructors either provide guidance on GenAI use or deploy guardrailed tools.","tokens_in":48427,"tokens_out":4359,"duration_ms":36565,"significance":"The report is a timely and methodologically transparent synthesis of a rapidly moving research area. Strengths include a detailed SLR protocol with data posted on OSF, explicit threats-to-validity coverage in Section 7, independent two-coder thematic analysis for open-ended survey items, and external comparison with the StackOverflow 2024 developer survey. The competency-shift finding (toward code reading, evaluation, and problem decomposition) is consistently corroborated across the educator survey, the developer survey, the interviews, and the literature review. The principal weakness is that the survey samples are small, self-selected convenience samples, and the report presents point estimates as community-level facts without confidence intervals or sensitivity analysis. If the results are reframed as an exploratory snapshot of a GenAI-engaged subset of the community, the report is a useful contribution; in its current form, the generalizing language overreaches the data.","major_comments":[{"comment":"The headline adoption percentages (77.6% of educators not explicitly disallowing GenAI, 35.5% actively incorporating it, 79.5% of developers using GenAI) are presented as community-level facts, yet they rest on self-selected convenience samples of N=76 and N=39 recruited through CSEd mailing lists, conference contacts, and personal networks. The authors' own Section 7.2 acknowledges that respondents 'may be more involved in CSEd than their peers,' but no confidence intervals, weighting, or sensitivity analysis are provided. Because the report's central claim is to describe 'what is happening on the ground,' these point estimates must be reframed as exploratory and reported with measures of uncertainty, and the key takeaways box should be reworded to avoid unsupported generalization.","section":"Section 5.2.1 and key takeaways box"},{"comment":"The recommendation to use instructor-guardrailed tools when students receive no guidance rests on a comparison of 19/26 (73%) positive findings for guardrailed tools versus 12/22 (55%) for general-purpose tools. These cell counts are small, no significance test is reported, and the studies are not randomized; confounding factors such as course level, study design, and tool quality are not controlled. The difference could easily arise from sampling variability, yet this recommendation is repeated verbatim in the key takeaways box ('Guide students on how to use GenAI or use custom tools'). Please add a caveat about the small cell counts, report a formal test if appropriate, or soften the conclusion.","section":"Section 3.4, Tables 7 and 8"},{"comment":"The developer survey's 79.5% adoption rate is described as 'in line with larger surveys,' citing the StackOverflow 2024 figure of 63%. A 16-percentage-point gap is substantial and may itself be evidence of the self-selected, academia-connected nature of the sample. Before using the StackOverflow data as external validation, the report should acknowledge this discrepancy and discuss whether the difference reflects sample bias, question wording, or a genuinely different population.","section":"Section 5.7 and Section 6"}],"minor_comments":[{"comment":"At least five of the ten reference papers used to validate the search string are authored or co-authored by members of this working group (e.g., [34], [86], [87], [94], [120]). Please disclose this overlap explicitly and, if feasible, add reference papers without author overlap to make the validation more independent.","section":"Section 2.1, Table 3"},{"comment":"ES-10 was answered by 17 educators who disallow GenAI, but the follow-up ES-11 reports 25 responses; please clarify the branching logic or the denominator for this question.","section":"Section 5.2.1"},{"comment":"The pie chart for developer usage frequency has labels that are visually confusing, with 'several times a day' appearing twice and small slices hard to read; please redraw with a clean legend and distinct categories.","section":"Figure 7"},{"comment":"The percentages for tool types (52% autocompletion, 48% chatbots, 29% no answer) do not sum to 100%; state that these are percentages of the 31 GenAI users and that the categories are not mutually exclusive.","section":"Section 5.7, DS-3"},{"comment":"The sentence 'Instructors and noticing a large influx' appears to be a typo; it should read 'instructors are noticing a large influx.'","section":"Section 8.2"},{"comment":"The caption says '1st column' where 'first column' is intended.","section":"Table 9 caption"},{"comment":"The response counts for developer demographics vary (country n=18, job title n=23, company type n=27); please add a note that these demographic questions were optional, leading to different response rates.","section":"Section 5.1.2"}],"recommendation":"major_revision","confidential_remarks":"This is a working-group report in a genre where the authors are also the primary community being described, and many of the cited SLR papers are by the authors' own group. That is not disqualifying, but the reference-paper validation and interview selection should be transparent about this overlap. The central empirical claims rest on very small convenience samples, and the report's generalizing language is the main risk. I would advise the editor to require the authors to add confidence intervals or ranges to the key percentages, soften the key-takeaways wording to an exploratory framing, and add a caution to the guardrailed-tool recommendation. If those changes are made, the report is a useful snapshot; as it stands, the gap between the data and the 'on the ground' claim is too large."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take: this is a solid working-group report that gives a genuinely useful snapshot of where GenAI in computing education stood in mid-2024. The survey percentages should be read as indicative, not as precise population estimates, and the report itself goes a long way toward saying that—but the key takeaways box overshoots.\n\nWhat's new: a systematic review of 71 papers through May 2024, fresh survey data from 76 educators and 39 developers, and 17 interviews. The triangulation across methods is the real strength. The most interesting analytical result is in the SLR cross-tab: when students get no guidance, guardrailed tools yield positive findings in 73% of studies versus 55% for general-purpose tools; once guidance is provided the gap mostly closes. That's actionable and worth knowing.\n\nSoft spots: the samples are small and convenience-based—recruited through CSEd mailing lists, ITiCSE contacts, LinkedIn, and author networks. That almost certainly over-represents educators who have engaged with GenAI. The headline numbers (77.6% don't disallow, 35.5% incorporate, 79.5% developer use) have no confidence intervals and would move if you sampled more broadly; the developer n=39 is especially thin. The cross-tab cells in Tables 7 and 8 are small (the 'yes guidance + guardrailed' cell has n=6). These are real limits, but for a field in flux they're acceptable as long as you don't take the percentages literally. The authors do list these threats in Section 7.2, which I credit them for.\n\nOne minor circularity: the SLR's reference-paper check includes at least five papers by the report's own authors. That's worth a footnote but doesn't compromise the survey or interview results, which are independent of the search validation.\n\nWho should read this: computing education researchers, instructors deciding on policy, and anyone new to the 2024 literature who wants a mapped overview. It deserves a serious referee—it's already peer-reviewed at ITiCSE-WGR, and I'd accept it in a journal with light revision to soften the takeaways box and add sampling caveats. I'd cite it as a snapshot, with the limitations spelled out.","headline":"Useful 2024 synthesis of GenAI in computing education; treat the survey percentages as indicative, not population estimates.","tokens_in":49148,"tokens_out":3395,"would_cite":true,"duration_ms":30139,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Most computing educators now allow generative AI in class, but only a third actively teach with it, and the skills educators say matter most are shifting from writing code to reading, evaluating, and decomposing problems.","keywords":["generative AI","computing education","large language models","pedagogical practices","teaching computing","systematic literature review","educator survey","developer survey"],"falsifier":"A large, demographically representative survey of computing educators (e.g., a random sample across institution types and countries) that found the proportion of educators who explicitly disallow GenAI is much higher than 22.4%, or that guardrailed tools show positive outcomes at rates no better than general-purpose tools when guidance is held constant, would directly contradict the report's central empirical claims.","tokens_in":47980,"feed_emoji":"🎓","tokens_out":2269,"duration_ms":21721,"temperature":0.7,"pith_summary":"This report tries to establish an accurate picture of how and why computing educators are integrating generative AI (GenAI) into their teaching, and how the competencies expected of programmers are changing. It combines a systematic literature review of 71 studies, a survey of 76 educators and 39 professional developers, and 17 interviews with educators, researchers, and tool creators. If the picture is right, the field has moved past banning toward pragmatic adoption: 77.6% of surveyed educators do not explicitly disallow student use of GenAI, yet only 35.5% actively incorporate it into courses. The report also claims that guardrailed GenAI tools produce positive learning outcomes in 73% of studies when students receive no explicit guidance, versus 55% for general-purpose tools, and that both educators and developers see a shift toward code reading, code evaluation, and problem decomposition as core skills.","feed_headline":"Most CS educators now allow AI tools in class; few teach with them","feed_subtitle":"A 71-paper review plus surveys finds guardrailed AI tools beat general chatbots when students get no guidance, and skills shift to code…","key_machinery":"The central machinery is a three-part triangulation: a systematic literature review (guided by Kitchenham and Brereton's method) that classifies 71 studies by tool type (general-purpose, task-specific, or instructor-guardrailed), purpose, and whether students received guidance; an international survey of educators and developers with both closed and open-ended questions; and semi-structured interviews with educators using GenAI, educators studying GenAI, tool creators, and one thoughtful non-user. The load-bearing cross-tabulation is Table 7, which compares tool type and guidance level against whether study findings were positive, negative, mixed, or neutral; it is this table that supports the report's recommendation that instructors either guide students on using general-purpose tools or adopt tools with pedagogical guardrails.","core_discovery":"The central claim is that the computing education community has entered a phase of partial, pragmatic integration of generative AI rather than wholesale adoption or prohibition. Triangulating a systematic literature review with educator and developer surveys and in-depth interviews, the report argues that most instructors tolerate GenAI tools (77.6% do not disallow them), fewer actively integrate them (35.5%), and those who integrate focus on teaching students to use the tools, generating course content, and providing feedback at scale. A key pattern from the literature review is that when students are not given explicit guidance on how to use GenAI, studies using tools with instructor-provided guardrails report positive findings 73% of the time, while studies using general-purpose tools report positive findings only 55% of the time; when guidance is provided, the type of tool matters less. Both surveyed educators and developers report that programming competencies are changing, with code reading, code evaluation/testing, debugging, and problem decomposition becoming more important than writing code from scratch, and prompting emerging as a new skill.","pith_inferences":["The report's recommendation to prefer guardrailed tools rests on a relatively small number of studies (26 for guardrailed without guidance), and the publication-bias caveat the authors acknowledge means the 73% figure could overstate effectiveness in unpublished settings.","The survey's convenience samples—76 educators and 39 developers, mostly from the authors' professional networks—mean the descriptive percentages (e.g., 77.6% not disallowing) should be read as characterizing these respondents, not necessarily all computing educators worldwide.","If the competency shift toward code reading and evaluation is real, a testable extension would be to compare learning outcomes in courses that explicitly assess code comprehension and evaluation versus courses that continue to assess only code writing.","The report's finding that educators underestimate developers' daily GenAI use suggests a potential lag in curriculum, but also raises the question of whether industry use is itself changing rapidly enough that current curricula may be chasing a moving target."],"forward_implications":["If the report's findings hold, computing educators should expect that most students will use GenAI regardless of policy, so explicit bans are increasingly impractical and less common.","Instructors who cannot provide detailed guidance on GenAI use should prefer guardrailed tools (e.g., tools that withhold complete solutions) over general-purpose chatbots, based on the 73% versus 55% positive-result rates.","Assessment practices are likely to keep shifting toward proctored exams, oral exams, and evaluating the process of software creation rather than the final artifact alone.","Computing curricula should increase emphasis on code reading, code evaluation, testing, debugging, and problem decomposition, while reducing emphasis on writing syntactically correct code from scratch.","Educators' perceptions of how developers use GenAI may misalign with actual industry use, so regular consultation with industry partners is needed to keep curricula relevant."],"supporting_citations":[{"why":"Liffiton et al.'s CodeHelp study is a central example of an instructor-guardrailed tool; the report's literature review classifies it and relies on its positive findings in the guardrail-versus-general-tool comparison.","marker":"[86]"},{"why":"Liu et al.'s CS50 with AI paper is one of the ten reference papers used to validate the search string and is a key example of using GenAI to support instruction at scale.","marker":"[87]"},{"why":"Kazemitabaar et al.'s CodeAid deployment provides empirical evidence for a guardrailed classroom assistant and is cited in the custom-tools list that shapes the tool-type taxonomy.","marker":"[60]"},{"why":"Denny et al.'s Prompt Problems paper defines a scaffolded exercise type that teaches prompting skills; it is a reference paper validating the search and illustrates the 'teach GenAI' category.","marker":"[34]"},{"why":"Jury et al.'s worked-examples study supplies evidence for the 'learning resources' purpose category and is a reference paper for the search string.","marker":"[57]"},{"why":"Roest et al.'s next-step hint generation study grounds the 'hints' purpose category and contributes to the half-positive findings for hint generation.","marker":"[120]"},{"why":"Prather et al.'s prior ITiCSE working group report provides the survey questions and baseline percentages that this report extends and compares against.","marker":"[111]"},{"why":"Hellas et al.'s quality metrics are used to evaluate the rigor of the 71 included studies, shaping the literature review's credibility assessment.","marker":"[47]"},{"why":"Kitchenham and Brereton's SLR guidelines provide the methodological basis for the systematic literature review process.","marker":"[72]"},{"why":"The Stack Overflow developer survey is used as the larger industry benchmark that the developer survey results are checked against for external validity.","marker":"[105]"}],"fun_headline_variants":["Most CS educators allow AI tools, but few teach with them","Guardrailed AI aids unguided students more than general chatbots","Code writing less vital; reading, debugging, and prompting on rise","Most CS ed allows AI, but active integration lags at 35.5%","Generative AI in CS: tolerance high, adoption low, skills shifting"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The most load-bearing premise is that the survey and interview participants represent the broader populations of computing educators and professional developers; the samples are small (76 educators, 39 developers, 17 interviews) and were recruited through the authors' own networks, so the reported percentages and recommendations only generalize if these participants are not systematically biased.","fun_headline_variants_meta":{"raw":{"variants":["Most CS educators allow AI tools, but few teach with them","Guardrailed AI aids unguided students more than general chatbots","Code writing less vital; reading, debugging, and prompting on rise","Most CS ed allows AI, but active integration lags at 35.5%","Generative AI in CS: tolerance high, adoption low, skills shifting"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000915,"raw_usage":{"total_tokens":3967,"prompt_tokens":1021,"completion_tokens":2946,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":637,"completion_tokens_details":{"reasoning_tokens":2852}},"tokens_in":637,"tokens_out":2946,"duration_ms":17145,"temperature":1.0,"reasoning_tokens":2852,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T11:56:39.060727+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A large, demographically representative survey of computing educators (e.g., a random sample across institution types and countries) that found the proportion of educators who explicitly disallow GenAI is much higher than 22.4%, or that guardrailed tools show positive outcomes at rates no better than general-purpose tools when guidance is held constant, would directly contradict the report's central empirical claims.","supporting_citations":[{"cited_title":"In Proceedings of the 55th ACM Technical Symposium on Computer Science Education V","cited_arxiv_id":null,"evidence_quote":"Hellas et al.'s quality metrics are used to evaluate the rigor of the 71 included studies, shaping the literature review's credibility assessment."}],"review_version":1}