{"id":"9727f4b1-0c67-478a-b4a2-b5a4d475f722","arxiv_id":"2411.11166","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A spring 2023 survey of 133 CS students at one R1 university found widespread voluntary use of LLM chatbots for writing, coding, and learning, with students mostly viewing GenAI as beneficial but divided on policy.","lead":"This paper reports a survey of computer science students at one U.S. university in spring 2023, asking how they used ChatGPT, Copilot, and image generators in their coursework. It finds that most had tried chatbots for writing, coding, and learning support, and that students generally see GenAI as beneficial, though they are split on how much to restrict it.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Non-response bias is the load-bearing risk: the 'most students' claim rests on a 12% opt-in sample and is not robust without a late-responder or weighting check.","rationale":"Reading the paper in good faith, it is an exploratory survey with transparent methods and a clear acknowledgment in Section 3.4 that the sample is not representative. The qualitative coding appears careful (IRR > 0.6, iterative codebooks), and the paper does not make causal claims. However, the abstract and introduction present the central findings without the caveat that the 12% response rate imposes. The load-bearing step is the leap from sample statistics to 'most students have tried GenAI'—a claim about the department (or beyond) that requires non-response bias to be negligible. The paper provides no evidence on this, and the subject of the survey makes self-selection likely. A wave analysis or weighting check would settle whether this leap is warranted. This does not change the reader's verdict, because they already flagged the representativeness concern; it sharpens it with a specific test.","tokens_in":11844,"tokens_out":5237,"duration_ms":51675,"concrete_test":"Request the de-identified survey dataset and run a wave analysis: order respondents by submission time (or by whether they responded before vs. after a reminder), and compare the first-half vs. second-half on 'ever used an LLM' (the everyday/regularly/once-or-twice/only-fun categories vs. never) and on the 1-10 benefit rating. If late responders are significantly less likely to have used LLMs and rate benefit lower, non-response bias is confirmed and the abstract's 'most students' claim overstates the population. If no timestamps are retained, use the department's official Spring 2023 enrollment by year/status to re-weight the sample; if the weighted 'ever used LLM' proportion falls below 50%, the claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that 'most students have tried GenAI tools' and 'tend to view GenAI as beneficial'—is stated in the abstract without qualification, but the evidence is a convenience sample of 133 respondents (12% of undergrads, 7.6% of grads) from one R1 department (Section 3.4). The paper explicitly acknowledges non-representativeness, yet the headline inference treats the sample as if it speaks for the department. The critical unexamined assumption is that non-responders resemble responders in GenAI adoption and attitudes. That assumption is implausible because the survey's subject (GenAI) is correlated with the outcome: students who have used GenAI and formed opinions are more likely to respond to a survey about it. Self-selection is likely to inflate both the adoption proportion and the benefit ratings. The paper's own data show a second gap: only 56.4% of the full sample (75 respondents) supplied descriptions of writing/coding/learning use cases, while the abstract implies the 'most' who tried did so 'for a variety of use cases'—the numerical basis for that clause is not the full sample. Without a non-response analysis or weighting, the 'most students' claim cannot be distinguished from 'most survey respondents'.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"Smith et al. report a Spring 2023 survey of computer science majors at a single small R1 U.S. university on their use and perceptions of generative AI tools. The study draws on 133 responses, presents frequencies of LLM, code-generator, and image-generator use, applies directed content analysis to open-text responses about writing, coding, and learning, and reports associations between usage frequency and benefit ratings. The central claims are that most students had tried GenAI tools for a variety of writing, coding, and learning use cases, and that students tend to view GenAI as beneficial to computing.","tokens_in":12074,"tokens_out":3854,"duration_ms":38283,"significance":"If limited to the respondent sample, this paper offers a timely and useful qualitative baseline of early GenAI adoption in computing education, with concrete student quotes and a clear coding methodology. Notable strengths are the iterative inter-rater reliability process, the public sharing of codebooks, and the honest acknowledgment of the non-representative sample. The paper's main value is as an exploratory snapshot that can inform later policy-oriented and larger-scale studies, but the population-level claims in the abstract currently outrun the evidence presented.","major_comments":[{"comment":"The abstract claims that 'most students have tried GenAI tools' and that 'students tend to view GenAI tools as beneficial,' but Section 3.4 reports that the sample comprises 12% of undergraduates and 7.6% of graduate students at one institution and explicitly acknowledges the sample is not representative. Because the survey topic is GenAI itself, non-response is likely correlated with the outcome: students who have used GenAI or formed opinions may be more likely to respond. Without a non-response analysis (for example, a late-responder comparison, weighting, or benchmarking against institutional data), the headline claims cannot be distinguished from 'most survey respondents.' Please either narrow the abstract and conclusion claims to 'respondents' or supply such an analysis.","section":"Abstract; Section 3.4"},{"comment":"The 'variety of writing, coding, and learning use cases' in the abstract is grounded in the 75 of 133 respondents (56.4%) who answered the optional free-response question about their use of GenAI. The remaining 43.6% of the sample did not provide this information, and the abstract does not disclose this. Furthermore, these 75 respondents are self-selected, so their use-case descriptions may overrepresent students with more extensive or more opinionated GenAI experience. Please add this qualification to the abstract and RQ1 summary, and consider reporting a sensitivity count that treats non-responders to the optional question as missing.","section":"Section 4.1; Abstract"}],"minor_comments":[{"comment":"The ACM Reference Format line lists the year as '2014' instead of '2024'.","section":"Header; ACM Reference Format"},{"comment":"The sentence 'only 36.1% have ever reported trying an image generator' is not supported by the disaggregated counts in the text; please report the image-generator frequency counts shown in Table 1 to make the calculation auditable.","section":"Section 4.1"},{"comment":"The codebook link uses a URL shortener (bit.ly/SIGCSE-GenAI-codebooks); please provide a persistent repository link or DOI to protect against link rot.","section":"Section 3.3"},{"comment":"The full survey instrument is not included; adding it as an appendix would clarify the exact wording of the free-response questions and the frequency-response options, which are central to interpreting the reported percentages.","section":"Section 3.2"},{"comment":"The claim that students 'tend to view GenAI as beneficial' is based on means of 6.78 and 7.41 on a 10-point scale with standard deviations around 2.6; please describe the distribution in prose as well, since a wide or bimodal distribution would weaken the interpretation of the mean.","section":"Section 4.2; Table 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a plausible ITiCSE contribution after revisions. The central issue is that the abstract and conclusion generalize beyond a low-response, single-institution, self-selected sample. This is fixable by tightening the language and adding the missing caveats; I do not see a reason to reject outright."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a useful early baseline. In spring 2023, before most institutions had policies, the authors surveyed CS students at one R1 university and got a snapshot of how they used LLMs, code generators, and image generators, and what they thought about GenAI's role in education and careers. The timing matters: this is one of the first systematic surveys of computing students specifically, so it is a citable data point for anyone studying the adoption period. The qualitative work is better than the quantitative framing. The use-case categories (writing, coding, learning) come from an iterative coding process with inter-rater reliability checks, and the student quotes are concrete and useful. The authors also deserve credit for being explicit in the methods section that the sample is not representative and that they treat the results as exploratory. The main soft spot is the mismatch between the abstract and the evidence. The abstract says \"most students have tried GenAI tools\" and \"tend to view them as beneficial,\" but the response rate was 12% of undergrads and 7.6% of grads. Self-selection is a real problem here because the survey subject is correlated with the outcome: students who already use GenAI are more likely to respond to a survey about it. So the \"most\" claim could easily be an artifact of who bothered to answer. To the paper's credit, that limitation is acknowledged in Section 3.4, but the abstract should carry the same caveat or the authors should do a late-responder or weighting check. A second issue: the variety of use cases is based on only 75 respondents (56.4% of the sample) who answered the optional free-response question. That subset is fine for qualitative description, but it is not the basis for \"most students.\" Also, the full survey instrument and raw data are not included, and I could not verify the bit.ly codebooks link. That is a minor reproducibility gap. One minor wording issue: the conclusion says students' use cases are \"already impacting their learning processes and outcomes,\" but the data are self-reports, not outcome measures. The authors themselves flag this in Section 5.1, so it is a careless phrase rather than a load-bearing flaw. Overall, the paper is a solid exploratory case study. The central descriptive claims hold for the sample, and the overreach is in the abstract, not the analysis. I would send it to peer review; the flaws are addressable and the contribution is worth preserving as a historical marker.","headline":"An honest early snapshot of GenAI use in one CS department, but the abstract overstates the evidence from a 12% opt-in sample.","tokens_in":716,"tokens_out":822,"would_cite":true,"duration_ms":26975,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"By the end of Spring 2023, most computer science students at a U.S.","keywords":["generative artificial intelligence","large language model","code generator","image generator","policy","survey","student experience","AI literacy"],"falsifier":"A comparison of self-reported GenAI use against actual usage telemetry (e.g., IDE plugin logs or browser history from a consenting subsample) would settle whether frequency estimates are inflated; if reported 'regular' use substantially exceeds logged use, the adoption baseline is overstated. Alternatively, replicating the survey at several institutions with 40%+ response rates and finding adoption below 30% would undermine the claim that most computing students had adopted GenAI by spring 2023.","tokens_in":11688,"feed_emoji":"🤖","tokens_out":6643,"duration_ms":56773,"temperature":0.7,"pith_summary":"This paper tries to establish an early empirical baseline: by the end of the Spring 2023 semester, most computer science students at one U.S. university had already tried generative AI tools, mainly large language model chatbots, and used them for writing support, coding help, and self-directed learning. The authors argue that this adoption happened before any institution-wide policy existed, with most classes offering little or no guidance. On the whole, students rated generative AI as more beneficial than harmful to the field of computing, though they were split on whether instruction should encourage, conditionally allow, or discourage use. The paper positions these self-reported use cases and opinions as input for future policy and tool design in computing education.","feed_headline":"Most computing students had already adopted GenAI by 2023","feed_subtitle":"A 133-student survey finds AI use in writing, coding, and learning before campus rules existed, with most calling it a net benefit.","key_machinery":"The central machinery is a survey instrument followed by directed content analysis, a qualitative method that applies a codebook to free-text responses. The survey asked students to rate their frequency of use of three GenAI categories (LLM chatbots, code generators, image generators), rate on a 1–10 scale how beneficial GenAI will be to computer science, and answer three free-response questions about their use, the appropriate role in education, and workplace concerns. Six human coders iteratively refined codebooks until inter-rater reliability (Krippendorff's alpha) exceeded 0.6, producing a taxonomy of emergent use cases—writing, coding, and learning—and of student opinions on degree of use, rationale, and implementation methods. That taxonomy is what lets the paper move from anecdote to a structured description of how students adopted the tools.","core_discovery":"Surveying all computer science majors at a small engineering-focused R1 university (116 undergraduates and 17 graduate students, 12% and 7.6% of their respective populations), the paper finds that most respondents had tried a large language model chatbot, that fewer had tried code generators, and that almost none had fully turned in AI-generated assignments. Instead, students described using GenAI to draft and debug code, explain concepts, find sources, outline and polish writing, and act as an informal tutor when instructors were unavailable. Asked about the role of GenAI in education, 69 of 126 respondents called for conditional use with instructor-specified boundaries, 41 wanted it encouraged, and 16 wanted it discouraged; students were nearly split on whether it helps or harms learning (44 vs. 40 coded responses), while many viewed AI skills as necessary for future employment. The authors claim these results capture a genuine pre-policy snapshot of early adoption and student attitudes that can inform curricula, policy, and tooling.","pith_inferences":["An implication the paper leaves implicit: if early adoption was already this common in spring 2023, later cohorts who have never known a pre-ChatGPT classroom will likely arrive with even higher baseline familiarity, making 'teach responsible use' the more realistic institutional stance.","The paper's taxonomy suggests a testable design principle: restrictions keyed to course level (e.g., blocking code drafting in CS1 but allowing it in capstone courses) may preserve learning outcomes better than bans on the tool type itself, a hypothesis an experimental study could check.","Interpreting students' 'calculator for coding' framing, assessment may need to shift from judging the produced code to judging the process of verification and design, for instance through oral exams or submitted explanations of AI-assisted changes.","Replicating this survey at a large public university or a teaching-focused college would show whether the 12% undergraduate response rate at a single engineering-focused institution yields representative adoption estimates."],"forward_implications":["If most students are already using LLMs before policy exists, then policies written in 2023–2024 were reacting to established behavior rather than preventing it.","Because students mostly used GenAI for supported tasks (explaining, debugging, outlining, tutoring) rather than full assignment completion, instructors can target restrictions at specific use cases such as 'no drafting code' without banning the tool outright.","The split between students who want conditional use and those who want encouragement suggests that a one-size-fits-all ban would conflict with many students' expressed preferences and career expectations.","Students' perception that AI literacy is needed for future jobs implies curricula should include professional AI use, not just academic integrity rules.","The finding that students who use LLMs more frequently rate GenAI as more beneficial (p < .0008) suggests familiarity goes with acceptance, so early exposure may shape later attitudes."],"supporting_citations":[{"why":"Supplies the directed content analysis method used to derive the use-case and opinion taxonomies from free responses.","marker":"[16]"},{"why":"Provides the prior framing of contract cheating that GenAI reframes and that the paper distinguishes from students' reported partial use.","marker":"[22]"},{"why":"Frames the CS-education-era-of-GenAI concern that motivates the survey's timing and research questions.","marker":"[12]"},{"why":"Provides federal policy guidance that the paper positions its student-centered data as informing.","marker":"[4]"},{"why":"Documents the landscape of early university GenAI policies that the paper argues were developed without student input.","marker":"[6]"},{"why":"Provides the instructor-side perspective on adapting to AI code generation that the paper complements with student-side data.","marker":"[19]"}],"fun_headline_variants":["Survey: Most CS students tried GenAI before campus rules","Students already using AI for coding and learning by 2023","CS students split on whether GenAI helps or hurts learning","Most CS students tried chatbots, but few used AI to cheat"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The findings rest on students accurately reporting how often they used tools that could look like cheating, from a department where only 12% of undergraduates and 7.6% of graduate students responded.","fun_headline_variants_meta":{"raw":{"variants":["Survey: Most CS students tried GenAI before campus rules","Students already using AI for coding and learning by 2023","CS students split on whether GenAI helps or hurts learning","Most CS students tried chatbots, but few used AI to cheat"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000175,"raw_usage":{"total_tokens":1278,"prompt_tokens":927,"completion_tokens":351,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":543,"completion_tokens_details":{"reasoning_tokens":282}},"tokens_in":543,"tokens_out":351,"duration_ms":4192,"temperature":1.0,"reasoning_tokens":282,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T18:49:51.233097+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A comparison of self-reported GenAI use against actual usage telemetry (e.g., IDE plugin logs or browser history from a consenting subsample) would settle whether frequency estimates are inflated; if reported 'regular' use substantially exceeds logged use, the adoption baseline is overstated. Alternatively, replicating the survey at several institutions with 40%+ response rates and finding adoption below 30% would undermine the claim that most computing students had adopted GenAI by spring 2023.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides federal policy guidance that the paper positions its student-centered data as informing."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the landscape of early university GenAI policies that the paper argues were developed without student input."},{"cited_title":"Ban It Till We Understand It","cited_arxiv_id":null,"evidence_quote":"Provides the instructor-side perspective on adapting to AI code generation that the paper complements with student-side data."}],"review_version":1}