{"id":"dd8c97c1-cc1c-46ea-8e0b-0cb19ebb63a9","arxiv_id":"2608.01112","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Subtle adversarial image perturbations can steer commercial AI assistants toward designated wrong answers, and the resulting answer patterns can statistically flag students who blindly copy those assistants.","lead":"This paper tests whether small tweaks to images in multiple-choice assignments can trick AI assistants into choosing specific wrong answers, leaving a pattern that reveals when students copied the AI. It reports that such fingerprints can flag sustained AI copying with few false alarms under carefully stated assumptions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 8-in-10,000 false-flag bound rests on unvalidated educator q_i bounds and product-Bernoulli independence; a modest real-world violation on a few high-weight questions could inflate the true Type-I rate far above the reported value.","rationale":"The reader's weakest assumption is exactly where I find the soft spot: the null distribution that produces the headline Type-I error is specified by researcher judgment plus an independence assumption, and neither is checked against genuine-student response vectors. I credit the paper for being explicit about this (Supp. B.3, D.2, G), for separating development and holdout calls, for using exact Clopper-Pearson bounds on assistant probabilities, and for reporting the fallback q_i = 0.5 sensitivity, which shows how much the guarantee depends on the educator model. But the central feasibility claim in Table 3 is conditional on those q_i being valid upper bounds and on item-level independence. Both are plausible yet unverified; a position preference (e.g., a student who defaults to 'C' when unsure) or one unusually attractive distractor can create positive dependence and/or violate the marginal bounds, inflating the upper tail of L. The proposed proctored-cohort test would settle this directly: it replaces judgment with empirical q_i bounds and replaces the product-Bernoulli assumption with a fitted joint model. If those checks keep the familywise Type-I error below 1e-3, the headline claim survives; if not, the paper should be revised to state the guarantee only under a validated student model. Because the authors have already scoped the claim as conditional and flagged the missing evidence, CONDITIONAL remains the right verdict; I do not see a demonstrated error that would force REJECT.","tokens_in":30795,"tokens_out":10280,"duration_ms":95443,"concrete_test":"Administer the fixed 20-question protected assignment to a proctored cohort of at least 300 genuine students from the intended population, then (a) replace each educator q_i with the Clopper-Pearson upper bound from Supp. D.1 Eq. (18) and recompute the familywise Type-I rate under the published Table 3 thresholds; and (b) fit a latent-ability model (e.g., 2PL IRT or beta-binomial) to the complete response vectors, simulate 100,000 assignments from the fitted joint distribution, and count how often L >= tau*. If the empirical or model-based 95% upper confidence bound on the false-flag rate exceeds 1e-3, the headline 8-in-10,000 guarantee does not hold for the target cohort.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is the genuine-student null model, not the transfer attack. Table 3's familywise Type-I bound alpha <= 7.21e-4 is computed from Eq. (7) in Supp. A.2 as P_q(L >= tau*) under a product-Bernoulli model, with q_i taken as the maximum of five PhD researchers' elicited probabilities (Supp. D.2) and conditional independence asserted in Supp. B.3. Supp. D.2 itself states these are 'modeling assumptions rather than empirical response frequencies,' and Supp. B.3 calls independence 'a substantive modeling assumption.' No complete response vectors from actual students on the protected assignment are reported. If true target-selection probabilities exceed the elicited upper bounds for only a few informative questions - because the target distractor is more attractive to the cohort than to researchers, or because the perturbation shifts guessing behavior - the fixed threshold tau* is crossed more often than alpha. The likelihood ratio is a positively weighted sum of match indicators (retained pairs satisfy p_i > q_i), so a systematic violation on a high-weight item moves the tail by orders of magnitude. The fallback model already shows the sensitivity: setting q_i = 0.5 raises the familywise Type-I bound to 0.02. Power is less at risk because p_i is measured from repeated API calls; the false-positive side is the unvalidated component.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a preventive approach to detecting sustained blind copying from AI assistants in multimodal multiple-choice assessments. Each protected question carries a subtle visual perturbation, optimized on an open-weight surrogate ensemble, that steers black-box assistants toward a secret designated incorrect answer. Across an assignment, the induced answer pattern acts as a statistical fingerprint: students who copy assistant outputs reproduce the target pattern more often than genuine students. The authors construct a shared 20-question assignment covering three a priori assistant endpoints (Claude Opus 4.8, Gemini 3.5 Flash, GPT-5.6 Sol), calibrate target-response probabilities with repeated API queries and Clopper-Pearson lower bounds, model genuine-student target-selection probabilities either from five PhD researchers' elicited upper bounds or from a data-free q_i=0.5 fallback, and report >95% modeled power with a modeled familywise false-flag rate below 8 in 10,000 under the educator model. They also report transfer to one out-of-set assistant and weaker separation for two others, and include a detailed supplement documenting assumptions, calibration protocols, and limitations.","tokens_in":31015,"tokens_out":5036,"duration_ms":46389,"significance":"If the quantitative claims hold, the paper offers a genuinely new direction: instead of post-hoc detection of AI-generated text, assessments are designed so that blind copying leaves an auditable statistical trace. The paper is internally careful in several ways that deserve emphasis: development and holdout calibration calls are separated (Supp. C.4), the reported bounds use exact Clopper-Pearson intervals and a union bound (Supp. C.2, Eq. 10), and the supplement explicitly lists the modeling assumptions behind the headline numbers. The central risk is not the adversarial steering but the genuine-student null model: the Type-I and posterior numbers are model computations, not measurements on real student response vectors. Because the paper frames its contribution as feasibility, this is addressable, but it is load-bearing for the main quantitative claims.","major_comments":[{"comment":"The familywise Type-I bound alpha <= 7.21e-4 is a model computation, not an empirical false-flag rate: it is the upper-tail probability of the product-Bernoulli statistic in Supp. Eq. (7) evaluated at q_i bounds elicited from five PhD researchers, and no complete genuine-student response vectors are reported. A violation of q_i on a few high-weight items, or any positive cross-question dependence in target-match indicators, can move that upper tail by orders of magnitude. The headline 'fewer than 8 in 10,000 genuine students' should therefore be stated as conditional on the unvalidated null model, and the paper should either supply a proctored pilot or a sensitivity analysis under correlated response models.","section":"Table 3 / Supp. A.2, Eq. (7)"},{"comment":"The educator-provided q_i are explicitly described as 'modeling assumptions rather than empirical response frequencies or frequentist confidence bounds.' Taking the maximum of five researcher estimates is not a confidence bound for the intended cohort; if the target distractor is more attractive to actual students than to the assessors, the retained-pair condition p_i > q_i and the threshold tau* in Table 3 are no longer conservative. Because this is load-bearing for the central Type-I claim, the paper should either provide controlled cohort data on the protected assignment or restrict the quantitative claim to a clearly labeled conditional feasibility analysis.","section":"Supp. D.2"},{"comment":"Conditional independence is acknowledged as 'a substantive modeling assumption' but is not tested on real response patterns. Correlated guessing, fatigue, shared misconceptions, or a single attractive distractor on a high-weight question can produce tail probabilities far above the product-Bernoulli values used in Eq. (7). The authors' justification is plausible for match indicators, but the central guarantee depends on this assumption, so a stronger deployment should calibrate the null from full-assignment student responses or use block/bootstrap procedures that preserve observed dependence.","section":"Supp. B.3"},{"comment":"Because the design-based randomization detector is abandoned (retention conditioning invalidates the design-based null), the method has no model-free component for the genuine-student distribution. This makes the q_i and independence assumptions in the previous comments load-bearing for every reported Type-I number. I recommend that the main text state this limitation more prominently and provide a concrete validation protocol, rather than leaving it only in the supplement.","section":"Supp. B.4"}],"minor_comments":[{"comment":"The sentence 'falsely flagging fewer than 8 in 10,000 genuine students' should consistently include the qualifier 'modeled' or 'under the educator-provided student model,' because the number is computed, not measured on genuine-student submissions.","section":"Abstract / Introduction"},{"comment":"The out-of-set transfer claim for Gemini 3 Flash should be accompanied by its own Type-I value of 1.24e-3, which exceeds the 8-in-10,000 headline bound for the covered set; the text should clarify that transfer does not inherit the same false-flag guarantee.","section":"Table 3, out-of-set rows"},{"comment":"The q_i=0.5 fallback relies on the structural assumption that the correct answer is at least as likely as any individual distractor for genuine students; this assumption should be stated in the main text wherever the fallback operating point is cited.","section":"Supp. D.3"},{"comment":"The label 'crysalis' appears to be a typo; the intended word is likely 'chrysalis.'","section":"Figure 11"},{"comment":"The author name 'Trie.u' in the Bayer, Trieu, and Ellison reference appears to be a formatting artifact and should be corrected.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is unusually honest about its limitations, which is a strength. The central stumbling block is not the adversarial steering but the unvalidated genuine-student null model: the 8-in-10,000 Type-I figure is a computation under expert-elicited q_i bounds and product-Bernoulli independence, not a measurement. I would not reject, because the framework is novel and the limitations are acknowledged, but the quantitative feasibility claim should be either empirically supported with pilot or proctored response data or explicitly reframed as conditional on the model."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read it. The headline is that the paper builds a 20-question assignment where adversarial perturbations steer Claude, Gemini, and GPT toward designated wrong answers, and a likelihood-ratio test catches at least 95% of modeled blind-copying while false-flagging fewer than 8 in 10,000 genuine students. That sounds like a big number, and the paper is careful about what it can claim: the false-flag figure is a computed bound under an educator-provided student model, not a measured rate on real students.\n\nThe genuinely new thing is the defensive repurposing. Glaze and Nightshade protect art, PhotoGuard and EditShield mess with editing, text watermarking tags generated text. This paper instead uses targeted steering to make AI answers produce a controlled, testable error pattern in educational assessments. That is a real conceptual step, and the implementation is thoughtful. They separate development calls from holdout calibration, use Clopper-Pearson lower bounds for assistant probabilities, use a union bound for familywise error, and are unusually honest in the supplement about the assumptions. They also test transfer to an assistant not in the construction set, which is worth credit.\n\nThe soft spot is exactly where the stress-test note lands. The 8-in-10,000 bound depends on the q_i upper bounds elicited from five PhD researchers, and on conditional independence of target-match indicators across questions. Neither is validated on actual student responses. If target distractors are more attractive to the real cohort than the researchers assumed, or if students' errors are correlated, the true Type-I rate can be orders of magnitude above the bound. The authors acknowledge this in Supp. B.3, and their own fallback (q_i=0.5) inflates the bound to 0.02. That is not a hidden flaw, but it means the central detection guarantee is a conditional feasibility result, not a field-ready number. Power is less at risk because p_i is measured from repeated API calls.\n\nThe paper does not ship code or data yet, which limits independent verification of the proprietary API measurements. That is a real weakness. Nothing here looks dishonest, and the limitations are stated fairly. The conclusion that the method is feasible under a defined student model holds up. What does not hold up is any implication that the false-flag rate is an empirical property of real student populations.\n\nWho gets value: people working on assessment security, adversarial defenses, and AI-in-education. It is a solid feasibility study with a clear threat model and conservative statistics. I would like a revision that adds a real-student pilot or even simulated response vectors with correlated errors, and releases the code. That would turn a good conditional result into a convincing one. I would send it to peer review.","headline":"The attack-transfer side is solid and the statistics are honestly conditioned; the false-flag bound is a computed bound under an unvalidated student model, not a measured rate.","tokens_in":31584,"tokens_out":3543,"would_cite":true,"duration_ms":28688,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Adversarial image tweaks can make AI assistants leave a statistical fingerprint when students copy them.","keywords":["AI cheating detection","adversarial perturbations","multimodal multiple-choice questions","statistical fingerprinting","likelihood-ratio detection","black-box transfer","educational integrity","targeted attacks"],"falsifier":"Give a proctored cohort of genuine students the protected 20-question assignment with no AI access, run the paper's detector with its reported $p_i$ and $q_i$ bounds, and count flags. If the empirical false-flag rate materially exceeds 8 in 10,000, or if students' response vectors show correlated target matches that lift the likelihood ratio over threshold, the conditional-independence and educator-bound assumptions are refuted.","tokens_in":30539,"feed_emoji":"🎓","tokens_out":8236,"duration_ms":64451,"temperature":0.7,"pith_summary":"The paper tries to establish that homework can be engineered so that sustained blind copying from an AI assistant leaves a detectable statistical trace, without relying on after-the-fact detection of AI-generated text. It protects multimodal multiple-choice questions with subtle visual perturbations that steer AI solvers toward designated wrong answers; students who copy those answers reproduce a pattern genuine students almost never produce. On a shared 20-question assignment, the paper reports that the detector catches at least 95% of modeled blind-copying cases while falsely flagging fewer than 8 in 10,000 genuine students under an educator-provided student model. If this holds, educators gain a preventive instrument: assessment design itself, not post-hoc classification, creates evidence of outsourcing.","feed_headline":"Perturbed exam images catch AI copyists at 95% power","feed_subtitle":"Flags 95% of modeled AI copiers on a 20-question assignment, falsely flagging fewer than 8 in 10,000 genuine students.","key_machinery":"The load-bearing object is the target fingerprint: each question's randomly chosen incorrect answer $S_i$ is kept secret, and the visual perturbation is optimized so that the assistant returns $S_i$ with high probability. The detector uses the indicator $Y_i$ that a submitted answer equals $S_i$, and the item-weighted log-likelihood ratio $L=\\sum_i[Y_i\\log(p_i/q_i)+(1-Y_i)\\log((1-p_i)/(1-q_i))]$, where $p_i$ is a conservative lower bound on the assistant's target rate from 30 repeated queries and $q_i$ is an educator-provided conservative upper bound on a genuine student's chance of choosing that distractor. Questions are retained only when $p_i>q_i$; the threshold is set to the most stringent value that keeps 95% power, and the exact weighted-sum distribution under conditional independence supplies the Type-I error.","core_discovery":"The paper's central claim is that adversarial machine learning can be turned into a preventive guard for educational assessments. For each protected question, the educator secretly chooses an incorrect answer, optimizes a small visual perturbation against an ensemble of open-weight surrogate models, and confirms by repeated black-box queries that the deployed assistant reliably returns that wrong answer. Blind-copying students thereby inherit an assignment-specific error pattern, while genuine students pick wrong answers independently. Using an item-weighted log-likelihood ratio over these target matches, the paper reports at least 95% detection power with a familywise false-flag bound below 8 in 10,000 under its educator-guided student model, and shows the shared 20-question assignment covers three major assistant families simultaneously and transfers to one assistant outside the construction set. It also reports that the guarantees degrade sharply under a data-free fallback student model and that some assistants outside the set are poorly separated.","pith_inferences":["The visible deterrent may be as valuable as the detector itself: once students know that blind copying leaves a reviewable statistical trace, the expected cost of outsourcing rises even when no flag is issued.","The same fingerprint mechanism could be carried by numerical canaries, questions whose answer is an unlikely value, which the paper names as future work and which may separate genuine students from assistants even more sharply than multiple-choice distractors.","Deployment would need per-cohort validation of the educator-provided bounds before quoting the 8-in-10,000 figure, because real student cohorts may share misconceptions that induce correlated target matches.","The poor separation for some assistants outside the construction set suggests that coverage is model-dependent; a practical system should re-run calibration whenever an assistant provider updates a model."],"forward_implications":["A 20-question protected assignment can simultaneously cover assistants from three major families, so an educator does not need to know which assistant a student would use.","Sustained blind copying across most of the assignment is the behavior that gets flagged; isolated or critically checked use falls outside the threat model.","Because the secret targets are assigned per question, the fingerprint survives the student's answer being a single letter and does not require analyzing free-form text.","If educators lack reliable estimates of how often genuine students pick each distractor, the fallback $q_i=0.5$ raises the false-flag bound to roughly one in fifty, so the reported guarantees depend on educator judgment.","The candidate pool admits at most three disjoint protected assignments under the analytical bound and two verified disjoint ones, so the approach is reusable but limited by the supply of steerable questions."],"supporting_citations":[{"why":"Establishes the adversarial-example phenomenon that the steering perturbations exploit.","marker":"(Szegedy et al. 2014)"},{"why":"Supplies the MI-FGSM optimizer used to generate the target-specific perturbations.","marker":"(Dong et al. 2018)"},{"why":"Provides the MMMU question pool used for assignment construction.","marker":"(Yue et al. 2024)"},{"why":"Provides the ScienceQA question pool whose text-solvability limits steering transfer.","marker":"(Lu et al. 2022)"},{"why":"Provides the MMBench question pool that proved most steerable.","marker":"(Liu et al. 2024)"},{"why":"Gives the exact one-sided binomial lower bounds for assistant target probabilities from repeated queries.","marker":"(Clopper and Pearson 1934)"},{"why":"Supplies the misconduct-prevalence estimate that motivates the prior sensitivity analysis at 1% and 5%.","marker":"(Chirikov, Smirnov, and Kizilcec 2026)"}],"fun_headline_variants":["Perturbed images make AI pick wrong answers, exposing cheaters","Adversarial exam images reveal AI cheating via error patterns","95% detection power: perturbed test images trap AI copiers","Steering AI to wrong choices fingerprints exam cheaters"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported false-flag rate rests on the assumptions that a genuine student's chance of picking each secret wrong answer is independent across questions and no higher than the educator's estimates; if students' errors are correlated through shared misconceptions or a single attractive distractor, the 8-in-10,000 figure could be far too low.","fun_headline_variants_meta":{"raw":{"variants":["Perturbed images make AI pick wrong answers, exposing cheaters","Adversarial exam images reveal AI cheating via error patterns","95% detection power: perturbed test images trap AI copiers","Steering AI to wrong choices fingerprints exam cheaters"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000341,"raw_usage":{"total_tokens":1861,"prompt_tokens":909,"completion_tokens":952,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":525,"completion_tokens_details":{"reasoning_tokens":882}},"tokens_in":525,"tokens_out":952,"duration_ms":8596,"temperature":1.0,"reasoning_tokens":882,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:11:46.997064+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give a proctored cohort of genuine students the protected 20-question assignment with no AI access, run the paper's detector with its reported $p_i$ and $q_i$ bounds, and count flags. If the empirical false-flag rate materially exceeds 8 in 10,000, or if students' response vectors show correlated target matches that lift the likelihood ratio over threshold, the conditional-independence and educator-bound assumptions are refuted.","supporting_citations":[],"review_version":1}