{"id":"5960979a-beda-40d9-926a-8455bd6be25a","arxiv_id":"2511.20878","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"95% of students in a classroom study used insecure code from a poisoned LLM despite prior security training; post-feedback, students' self-reported trust shifted toward distrust.","lead":"A classroom experiment enlisted an LLM fine-tuned to produce insecure code and found that 95% of 61 security-educated students submitted the flawed code anyway. After an automated vulnerability report, a voluntary follow-up survey suggested students became more skeptical of AI-generated code.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Self-selected 34% post-survey and no control group make RQ2's 'Bifröst increases skepticism' unsupported; W=80.5 also inconsistent with described data.","rationale":"The reader's weakest assumption correctly identifies the self-selected post-survey and uncontrolled design as the Achilles heel of the RQ2 claim. This is the most load-bearing concern because the paper's central contribution includes an effectiveness claim for Bifröst as an educational intervention; if the 21 respondents are not representative or the shift is due to demand characteristics, the framework's value is unsupported. The descriptive RQ1 finding (95% vulnerability) is more robust, though it too could be clarified by log data. I also noticed a concrete numeric inconsistency in the reported Wilcoxon W, which independently warrants verification. Neither issue is fatal if addressed; the 'preliminary evidence' framing is honest, but the effectiveness claim should not be stated as strongly as 'statistically significant evidence' without a control or non-response analysis.","tokens_in":11060,"tokens_out":15537,"duration_ms":152993,"concrete_test":"Using the existing server and survey data, compare the pre-survey security-trust ratings, security coursework, programming experience, and AI-tool usage of the 21 post-survey respondents vs the 40 non-respondents (e.g., Fisher exact/Mann-Whitney). If the groups differ (especially in initial distrust), the observed shift is confounded by selection; if they are indistinguishable, the selection concern weakens. Independently, recompute the Wilcoxon signed-rank statistic from the raw paired responses to verify W=80.5.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The RQ2 effectiveness claim (§5.2) hinges on a one-sided Wilcoxon signed-rank test (W=80.5, p=0.033) applied to the 21 of 61 students who voluntarily completed the post-survey. This is a self-selected subsample with no control condition. Students who chose to respond are likely more engaged, more affected by the feedback, or more willing to report skepticism (demand characteristics), so the test cannot distinguish the effect of Bifröst's feedback from selection, test-retest effects, or the mere act of receiving a report that labels their submitted code as insecure. The paper provides no comparison of responders vs non-responders. Moreover, the reported W=80.5 is numerically suspect: with 11 tied zero differences among 21 pairs, the maximum possible sum of ranks in one direction is 55 (excluding zeros), so W cannot equal 80.5; p=0.033 matches a much smaller W (~9.5). This suggests a reporting error that needs verification, though even a corrected p would not remedy the design flaw.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Bifröst, a classroom framework that combines a VS Code extension connected to a deliberately poisoned code-generation LLM with static-analysis-based vulnerability feedback, deployed in an undergraduate security class (n=61). The study reports two findings: (RQ1) despite students' stated skepticism toward LLM-generated code, over 95% of them submitted insecure code for both AES-encryption and command-execution tasks; and (RQ2) based on a voluntary post-survey (n=21), self-reported distrust increased after receiving Bifröst's feedback, supported by a one-sided Wilcoxon signed-rank test (W=80.5, p=0.033). The paper claims this is the first empirical investigation of student preparedness for insecure LLM code and that Bifröst fosters a security mindset.","tokens_in":11301,"tokens_out":7343,"duration_ms":79611,"significance":"If the RQ1 result holds, it provides a valuable and timely demonstration of a perception-behavior gap: students who say they distrust LLM code still submit vulnerable code in a realistic IDE setting. This is a useful contribution to security education and human-AI interaction. The Bifröst framework itself is a reasonable and reproducible intervention, and the deployment with a real course is a strength. However, the RQ2 effectiveness claim is currently not reliable: it rests on a self-selected 34% post-survey without a control condition, and the reported test statistic is numerically inconsistent with the data. The central RQ1 measurement appears directionally robust, but the paper's contribution to demonstrating that feedback increases skepticism is not yet established.","major_comments":[{"comment":"The reported Wilcoxon statistic W=80.5 is impossible with the data as described. The post-survey distribution implies 11 tied zero differences among the 21 pairs (2 initially \"neither\" who stayed, 9 initially \"distrust\" who stayed, and 0 from the \"somewhat trust\" group). With ties removed, the effective N for the signed-rank test is 10, so the maximum possible sum of positive ranks is 55, not 80.5. The p-value of 0.033 corresponds to a much smaller W (around 8-10). This indicates a computational or reporting error. Please re-run the exact test with zeros excluded and report the correct W, p, and effect size, or explain how W=80.5 was computed.","section":"§5.2, Statistical Validation"},{"comment":"The causal claim that \"Bifröst increases students' skepticism toward LLM-generated code\" is not supported by the study design. The post-survey was completed by only 21 of 61 students (34%), a self-selected subset. No comparison of responders vs. non-responders is provided, so selection bias is possible. The absence of a control group means the observed shift could be due to test-retest effects, demand characteristics, or simply the act of receiving a report that labels one's submitted code as insecure. Even with a corrected Wilcoxon test, this design cannot support the strong conclusion in the RQ2 answer. The abstract's phrase \"preliminary evidence\" is appropriate; the body's \"statistically significant evidence\" and \"can effectively develop\" are overstatements. Please add explicit limitations and temper the claims.","section":"§5.2, Effectiveness of Bifröst"},{"comment":"The paper claims students \"accepted\" insecure LLM-generated code, but the analysis only examines the submitted code, not whether the student clicked the plugin's \"Use code\" button. The plugin logs these decisions (as stated in §3.2), but the results do not report these logs. A student who manually typed the insecure code after seeing it, or who modified it, did not necessarily \"accept\" the LLM's output. This distinction matters for RQ1's framing about reliance on LLM-generated code. Please report the logged acceptance decisions and, if unavailable, soften the interpretation to \"submitted vulnerable code\" rather than \"accepted the LLM's suggestion.\"","section":"§5.1, Student Ability to Identify Insecure AI Code"}],"minor_comments":[{"comment":"The text says that among the 7 initially \"neither\" students, 2 maintained \"neither\", 4 shifted to \"somewhat distrust\", and 1 to \"highly distrust\" — which sums to 7 — but then immediately mentions \"one student (4.8%) changed their response to 'somewhat trust'\". This appears to be an inconsistency or the sentence is misplaced. Please clarify the counts.","section":"§5.2, Initially Neither paragraph"},{"comment":"The percentage labels in Figure 7 are hard to map to the text (e.g., preliminary shows 48%, 19%, 33%, but the text reports 10/47.6%, 7/33.3%, 4/19.0%). Please label each bar with the exact n and percentage, and ensure the figures are consistent with the text.","section":"Figure 7"},{"comment":"The manuscript contains several formatting placeholders (e.g., \"Conference'17\", \"July 2017\", the ACM DOI template). These need to be updated before any venue submission.","section":"General"},{"comment":"The matched rank-biserial effect size of 0.53 is reported without a formula or confidence interval. Given the ties and the small effective N (10), this effect size should be recomputed and accompanied by a confidence interval or at least the underlying nonzero-difference count.","section":"§5.2, Effect size"}],"recommendation":"major_revision","confidential_remarks":"The statistical error in §5.2 is serious and should be verified before the paper can be considered further. However, the RQ1 result — that skeptical students still submit vulnerable code — is a valuable, defensible finding. The RQ2 effectiveness claim can be salvaged by reframing it as an exploratory pilot with clearly stated limitations, but the current presentation overstates the evidence. I recommend major revision rather than rejection because the core measurement contribution is sound and the flaws are in the interpretation and reporting, not in the underlying RQ1 data."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe one thing you should know: this paper's central measurement is worth citing. In a security class, 61 undergraduates mostly said they didn't trust the security of LLM-generated code, yet about 95% of them submitted code using insecure patterns (ECB mode for AES, shell=True for subprocess) suggested by a poisoned CodeGen model. That perception–behavior gap is a genuinely useful empirical baseline for anyone teaching secure coding with LLM assistants.\n\nWhat the paper does well: the Bifröst framework is a sensible package — poisoned model adapted from Trojanpuzzle, a VS Code plugin, Bandit/CodeQL analysis, and email feedback. The two tasks are simple and unambiguous, and the few students who avoided the insecure patterns show the behavior is measurable. The authors are appropriately cautious about calling it a preliminary study.\n\nThe soft spots are real, and they cluster in RQ2. The post-survey was completed by only 21 of 61 students, voluntarily, with no control condition. Those who responded are likely more engaged or more affected by the feedback, so you can't separate selection from the intervention. The shift in self-reported trust could also be a demand characteristic: students were told their submitted code was vulnerable, then asked whether they trust such code. The paper should at minimum compare responders to non-responders.\n\nThere's also a statistical inconsistency. The paper reports a one-sided Wilcoxon W=80.5, p=0.033 with N=21. From the described transitions, 11 pairs were unchanged, leaving 10 nonzero differences; the maximum possible sum of ranks in one direction with 10 items is 55. So W=80.5 is not a valid signed-rank statistic for these data. This is probably a reporting or computation error, but it needs fixing before the effectiveness claim is credible. Even a corrected p wouldn't rescue the design.\n\nA minor gripe: the paper claims \"no existing studies\" address this educational gap. That may be true, but the claim is stated without a systematic search, so it could be softened.\n\nBottom line: the RQ1 measurement is a real contribution, and the framework is worth building on. RQ2 is a hypothesis-generating pilot, not evidence of efficacy. I'd send it to review, but I'd require the authors to release the acceptance logs, correct the statistics, and recast RQ2 as descriptive.","headline":"The RQ1 finding — security-educated students who distrust LLM code still submit insecure code ~95% of the time — is a solid, useful empirical baseline; the RQ2 effectiveness claim is not supported by the design, and the reported Wilcoxon W looks internally inconsistent.","tokens_in":11776,"tokens_out":4341,"would_cite":true,"duration_ms":46229,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Students say they distrust AI-generated code, yet in a hands-on classroom exercise over 95% accepted insecure code from a deliberately poisoned assistant; the paper argues that targeted feedback can begin to close that gap.","keywords":["LLM-generated code security","AI code assistants","security education","poisoning attacks","student perceptions","critical thinking","VS Code extension","Bifröst"],"falsifier":"A randomized controlled experiment in which one group receives Bifröst's vulnerability report and a control group receives a generic 'your code was reviewed' message without specifics; if the control group shows the same shift in self-reported distrust, the feedback mechanism is not the active ingredient. Alternatively, if students who receive the report go on to accept insecure code at the same rate in a follow-up task, the framework's effect on actual security behavior is unsupported.","tokens_in":10967,"feed_emoji":"🛡️","tokens_out":4821,"duration_ms":43386,"temperature":0.7,"pith_summary":"This paper tries to establish that undergraduate students in security courses have a perception–behavior gap when using AI code assistants: they report skepticism about the security of LLM output, but in a realistic task they overwhelmingly accept code with known vulnerabilities (ECB-mode AES and shell=True command execution). The authors present Bifröst, a classroom framework that pairs a poisoned code-generation model with a VS Code extension and automated vulnerability reports, and find preliminary evidence that receiving concrete feedback shifts students' stated trust toward greater distrust. If correct, the study shows that security education cannot rely on critical attitudes alone and that experiential feedback about specific flaws is a promising way to foster a security mindset.","feed_headline":"Over 95% of students accepted insecure AI code despite skepticism","feed_subtitle":"Stated distrust failed to prevent insecure code use; specific feedback shifted student trust toward skepticism.","key_machinery":"The carrying mechanism is Bifröst, a classroom measurement-and-feedback framework: a VS Code extension that lets students request code from a deliberately poisoned code-generation model (fine-tuned to suggest insecure patterns such as ECB mode and shell=True), a submission pipeline that runs static analysis to flag vulnerabilities, and an automated PDF report emailed to each student that names the vulnerable lines and explains the risk. The poisoned model makes insecure code appear functionally correct, so the only reliable signal students have is their own security judgment; the framework turns that invisible failure into visible feedback.","core_discovery":"In the paper's own terms, the central finding is twofold. First, students' stated skepticism about LLM-generated code does not predict their behavior: 58 of 61 students (95%) submitted the intentionally insecure code for an AES task and 60 of 61 (98%) did so for a command-execution task, despite most having completed security coursework and expressed distrust in a pre-survey. Second, after receiving Bifröst's automated vulnerability reports, a self-selected subset of 21 students showed a statistically significant shift toward distrust of AI-generated code security (moderate effect size), with responses moving from 'somewhat trust' and neutral toward 'somewhat distrust' and 'highly distrust'.","pith_inferences":["The perception–behavior gap documented here likely extends beyond students to professional developers, who face similar pressures to accept working AI code; a replication with experienced practitioners would test that.","A stronger test of the framework's effectiveness would measure whether the observed shift in self-reported trust changes actual acceptance behavior in a follow-up task, not just survey responses.","The framework could be adapted to other vulnerability classes and languages, and to compare feedback formats (e.g., inline IDE warnings vs. emailed reports) for their effect on skepticism.","Since the post-survey was voluntary and only a third responded, the reported effect may overstate the intervention's reach; a mandatory post-survey or control group would clarify the true impact."],"forward_implications":["Security educators should treat students' stated distrust of AI code as a starting point, not an outcome; attitudes do not automatically translate into secure behavior.","Feedback that points to specific vulnerabilities in code the student actually wrote can increase skepticism toward AI-generated code, suggesting a concrete intervention for classrooms.","The framework gives instructors a measure of student preparedness for LLM-assisted development, independent of self-report.","Because the vulnerable code runs without errors, exercises like these teach students that functional correctness is not evidence of security.","Future iterations of such frameworks may need instructor-led follow-up, since a few students in the study did not internalize the feedback."],"fun_headline_variants":["95% of students used insecure AI code despite distrust","Student AI skepticism fails to block insecure code","Feedback shifts student trust on AI code to skepticism","Insecure AI code accepted by 95% despite security classes","Bifröst feedback turns student AI trust to distrust"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The claim that Bifröst increases skepticism rests on the assumption that the 21 students who voluntarily completed the post-survey represent all 61 participants, and that the observed shift in distrust is caused by the feedback rather than by the act of being told their code was insecure or by wanting to appear more critical.","fun_headline_variants_meta":{"raw":{"variants":["95% of students used insecure AI code despite distrust","Student AI skepticism fails to block insecure code","Feedback shifts student trust on AI code to skepticism","Insecure AI code accepted by 95% despite security classes","Bifröst feedback turns student AI trust to distrust"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000135,"raw_usage":{"total_tokens":961,"prompt_tokens":706,"completion_tokens":255,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":450,"completion_tokens_details":{"reasoning_tokens":179}},"tokens_in":450,"tokens_out":255,"duration_ms":3340,"temperature":1.0,"reasoning_tokens":179,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T20:07:12.652677+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A randomized controlled experiment in which one group receives Bifröst's vulnerability report and a control group receives a generic 'your code was reviewed' message without specifics; if the control group shows the same shift in self-reported distrust, the feedback mechanism is not the active ingredient. Alternatively, if students who receive the report go on to accept insecure code at the same rate in a follow-up task, the framework's effect on actual security behavior is unsupported.","supporting_citations":[],"review_version":1}