{"id":"3f529792-a4f1-4586-baea-1213c5336403","arxiv_id":"2412.02713","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Person-fit statistics derived from item response theory separate human and chatbot multiple choice response patterns at low pollution levels, but the signal weakens as AI answers become common.","lead":"The paper shows that statistical patterns in multiple choice answers can separate human students from three major chatbots, ChatGPT, Gemini, and Claude, using item response theory. This could give schools a new way to spot AI cheating on tests that contain only letters, not text.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"RQ3's diminishing PFS gap may be an artifact of re-estimating item difficulties on AI-polluted, shrinking human samples; human-only calibration should be tested.","rationale":"The reader's weakest assumption was external validity: that 20 full answer sheets per chatbot with an identical prompt do not represent real partial cheating. That is a legitimate generalization concern, but it does not attack the internal validity of the reported results. I find a more load-bearing, internal concern: RQ3's key result--that PFS separation diminishes with pollution--may be an artifact of re-estimating item difficulties on datasets that are increasingly contaminated by AI responses and that shrink to only 60 human respondents. Since the nonparametric PFS used here derive item difficulties from the same response matrix (PerFit package), the item ordering is not fixed; at 25% pollution, AI responses strongly influence which items are deemed 'easy' or 'hard.' This can mechanically reduce the human-AI PFS gap regardless of any real change in detectability. If this is true, the paper's central claim that PFS is 'sensitive to the amount of AI cheating' is not established, and the practical conclusion that the method only works when AI responses are a minority would need revision. RQ1 (group differences at 5% pollution) is likely robust to this issue, so rejection is too strong, but the paper should be conditional on a reanalysis with human-only calibration. The reader's verdict of CONDITIONAL remains appropriate, though for a different reason, so I set verdict_should_be to UNCHANGED. My concrete test would settle whether the artifact is real; if it is, the RQ3 conclusions and the associated practical claims must be rewritten.","tokens_in":11152,"tokens_out":7145,"duration_ms":66987,"concrete_test":"Recompute the RQ3 analysis (Section 3.3) using item difficulties fixed to those estimated from the full human-only samples (N=931 chemistry, N~4800 psychometric) or from a large held-out human calibration set, then apply these fixed difficulties when computing G, G*, U3, and ZU3 for the 5%, 10%, and 25% polluted datasets. If the human-vs-AI differences remain significant at 25% pollution, or if the decay pattern changes materially, the reported pollution sensitivity is an artifact of within-dataset calibration. Also report the AUC or distribution overlap at each pollution level to assess individual-level detection, which the paper currently omits.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.3 states that the nonparametric PFS (G, G*, U3, ZU3) are computed directly from the dataset, which in the PerFit package means item difficulties are the sample p-values from the same response matrix. In RQ3 (Section 3.3), each pollution level is a new dataset: the same 20 AI responses are combined with 380, 180, or 60 human responses, so item difficulties are re-estimated on data in which AI responses constitute 5%, 10%, and 25% of the sample, and the human calibration sample shrinks to n=60 at 25%. Because AI response patterns differ systematically, the item difficulty ordering drifts toward the items AI tends to answer correctly, making AI responses look more Guttman-consistent and human responses less so. This mechanically shrinks the PFS gap as pollution increases, independent of any substantive 'aberrant becomes norm' effect. The paper's conclusion in Section 3.3 and the Discussion that detection degrades as cheating becomes prevalent is thus confounded by calibration contamination and small-sample item difficulty estimation. The central claim that PFS is 'sensitive to the amount of AI cheating' is not cleanly established by the reported analysis.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes using item response theory (IRT) person-fit statistics (G, G*, U3, ZU3) to distinguish between human responses and responses generated by three chatbots (ChatGPT-4o, Gemini 1.5 Pro, Claude 3.5 Sonnet) on two multiple-choice instruments: a high-school chemistry formative assessment (n=931) and a psychometric quantitative test (n>4800). For each chatbot, 20 full answer sheets are generated with a fixed prompt. The paper reports highly significant Wilcoxon rank-sum differences between human and AI person-fit distributions (RQ1), significant chatbot differences via Kruskal-Wallis tests (RQ2), and a diminishing gap between human and AI PFS as the proportion of AI responses increases from 5% to 25% (RQ3). The authors conclude that PFS provide a psychometrically grounded basis for detecting AI cheating in MCQ assessments.","tokens_in":11320,"tokens_out":8151,"duration_ms":71487,"significance":"If the reported effects are genuine, the paper addresses a nearly unexplored and practically important problem in educational assessment. The RQ1 result is strengthened by the use of two authentic large-scale datasets, three leading chatbot models, and a standard published implementation (PerFit), with very low p-values and visually separated densities. The paper also connects the work to established psychometric theory and clearly states its limitations. However, the practical contribution depends on whether group-level differences translate into individual-level detection and whether the RQ3 trend is not an artifact of recalibration. If the calibration confound is resolved, these results would be a useful foundation for AI-cheating screening in MCQ-based settings.","major_comments":[{"comment":"The conclusion that the PFS gap diminishes with increasing pollution is confounded by re-estimating item difficulties within each pollution condition. As stated in Section 2.3, the nonparametric PFS are computed directly from the response matrix, and in the PerFit implementation item difficulties (proportion correct) are estimated from the same data. From 5% to 25% pollution, the calibration sample shrinks from 380 to 60 humans and gains 20 AI responses. Because AI response patterns differ systematically, the item-difficulty ordering drifts toward items that AI tends to answer correctly, which mechanically changes G, G*, U3, and ZU3 for both groups. The observed shrinkage may therefore be an artifact of calibration contamination rather than evidence that AI patterns become normal. Please re-run RQ3 with item difficulties fixed from the full human-only samples (or with parametric IRT and fixed item parameters) and show whether the diminishing-gap trend persists.","section":"Section 3.3, Figure 3"},{"comment":"The manuscript reports only group-level distributional comparisons; it never evaluates per-test-taker detection performance, which is needed to support the abstract's claim that the method 'effectively highlights the differences' and can 'distinguish' human from GenAI responses. No ROC/AUC analysis, no threshold, no sensitivity/specificity, and no effect sizes (e.g., Cliff's delta or rank-biserial correlation) are provided. With AI samples of only 15-20 responses per model, the extremely small p-values do not quantify how well an individual score can be classified. Please add individual-level metrics such as AUC or misclassification rates, or explicitly restrict the claim to group-level differences.","section":"Section 3.1, Table 1"},{"comment":"The simulation treats 20 full answer sheets generated by one model with a single fixed prompt as representative of real GenAI cheating. Actual cheating is likely partial: students may answer some items themselves, mix answers from different models or runs, verify outputs, or use different prompts. The paper does not test such mixed patterns, so the strong separation in Figure 1 may not transfer to realistic cheating scenarios. At minimum, the authors should acknowledge this as a boundary of the central claim and, ideally, add a partial-cheating condition.","section":"Section 2.2.3"},{"comment":"The reported construction of the '5% pollution' datasets is inconsistent. Section 3.3 says the 5% condition combines 20 AI responses with 380 human responses (20/400 = 5%), while Section 3.2 describes RQ2 datasets with 931 or 980 humans plus 45 or 60 AI responses, which are approximately 4.6% and 5.8% pollution, respectively. Section 3.1 does not state how many humans were used in RQ1. Please specify the exact sample sizes and whether item difficulties for RQ1 were estimated on the full human sample or on the pollution-condition subsample; this is needed to reproduce the results and to interpret RQ3.","section":"Sections 3.1-3.3"}],"minor_comments":[{"comment":"'Signed-rank test' should be 'rank-sum test' (Mann-Whitney) for independent human and AI groups; Section 2.2.1 uses the correct term.","section":"Table 1"},{"comment":"Please state explicitly that PerFit computes nonparametric item difficulties as sample proportions from the same response matrix; this detail is essential for interpreting RQ3.","section":"Section 2.3"},{"comment":"There is a duplicated 'of of' in the sentence introducing the research questions.","section":"Section 2.1"},{"comment":"The reference title is garbled ('itelVergelijkbaarheid van individuele testprestaties,Comparability of individual test performance') and should be cleaned up.","section":"Reference [29]"},{"comment":"Some Z values are rendered as 'Z=,-4.76' and the caption should name the model and instrument for each panel; as printed, the panels are not fully self-explanatory.","section":"Figure 3"},{"comment":"In the U3 subplot the axis labels appear corrupted (e.g., '0.6 0.8' with a stray '.'); replace with clean plots.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the LAK scope and the RQ1 result is credible and worth publishing if the issues above are addressed. The main risk is overclaiming practical detection from group-level tests and presenting the RQ3 trend without controlling for recalibration. I would not require new data collection, but I would require the fixed-calibration reanalysis and individual-level detection metrics, plus a clearer statement of the exact sample construction."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline is: this paper nails group-level separation but fumbles the pollution analysis. The four person-fit statistics separate human from chatbot responses on two real MCQ instruments with p-values below 1e-5, and that result is clean. The paper is a real extension of Sorenson and Hanson, which only tried Rasch outfit on ChatGPT 3.5 for one chemistry exam. Here you get four statistics, three chatbots, two instruments, and a pollution study.\n\nCredit where due: the design is straightforward, the human data are authentic (931 chemistry students, 4,800 psychometric examinees), and the statistics come from published definitions via the PerFit package. No constants are fitted to the conclusion. The density plots show clear separation. This is a competent empirical study.\n\nThe soft spots are equally clear. The biggest one is RQ3. Section 2.3 says the nonparametric PFS are computed directly from the dataset, which in PerFit means item difficulties are the sample p-values from the same response matrix. For the pollution study, each level is a new dataset where the 20 AI responses are combined with 380, 180, or 60 humans. So item difficulties are re-estimated on data that are progressively more contaminated by AI. Because AI responses are systematically different, the difficulty ordering drifts toward items AI answers correctly. That mechanically reduces the PFS gap as pollution increases. The stress-test note is right: this confound means the paper does not cleanly establish that the method becomes less sensitive as cheating grows. The authors should calibrate difficulties on human-only data or fix them from a clean calibration.\n\nSecond, the paper stops at group-level inference. There is no per-student detection evaluation: no thresholds, no ROC or AUC, no false-positive rates. Saying this method can 'identify AI cheating' is a stretch; what it shows is that mean PFS differ. Real cheating is probably partial, mixed, and prompted differently. The 20 full answer sheets generated from one prompt per instrument are a narrow proxy. That's not fatal, but it limits the practical claims.\n\nThird, 20 responses per chatbot per instrument is a small sample. The Wilcoxon tests are fine, but the between-chatbot profiles in RQ2 rest on 15 to 20 observations per group. Effect sizes are missing throughout, which would help judge practical magnitude.\n\nMinor: Table 1 calls it a 'signed-rank' test, but it is a rank-sum test. Check that.\n\nOverall: the RQ1 result is solid and worth reporting. The RQ3 conclusion needs rework. The paper deserves peer review, especially given how timely the problem is. I would send it to a referee with instructions to require the calibration fix and per-student analysis.","headline":"Group-level separation is real and clean; the pollution analysis is confounded by re-estimated item difficulties.","tokens_in":11875,"tokens_out":3310,"would_cite":true,"duration_ms":68820,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Person-fit statistics from item response theory separate chatbot answers from human answers on multiple-choice assessments.","keywords":["generative AI","cheating detection","multiple-choice assessments","item response theory","person-fit statistics","chatbots","educational measurement"],"falsifier":"Take a real class answer set and inject cheated answer sheets constructed by giving a chatbot only a random half of the items and letting a human answer the rest. If the PFS distributions of these blended sheets overlap those of honest students, the method fails on realistic partial cheating. A sharper version: run the same pipeline at a 10% pollution level on a third high-stakes instrument; if $ZU_3$'s separation already vanishes there, the operating range is even narrower than the paper reports.","tokens_in":10888,"feed_emoji":"🤖","tokens_out":3869,"duration_ms":31429,"temperature":0.7,"pith_summary":"The paper claims that item response theory's person-fit statistics can separate chatbot-generated answers from human answers on multiple-choice tests. Working from two authentic instruments—a high-school chemistry formative test and a quantitative section of a high-stakes entrance exam—the authors show that premium versions of ChatGPT, Claude, and Gemini produce response patterns that are statistically far more aberrant than those of human test-takers. The same statistics also reveal that the three chatbots reason differently from one another. A third finding is that the separation shrinks as the share of AI responses in a dataset grows, so the method is most credible when AI cheating is still a minority behavior.","feed_headline":"IRT stats flag AI answer sheets in multiple-choice exams","feed_subtitle":"Person-fit scores separate ChatGPT, Claude, and Gemini responses from real students—while they stay a minority.","key_machinery":"The machinery is person-fit statistics built on the Guttman pattern: an ideal response vector in which a test-taker answers the $r$ easiest items correctly and the rest incorrectly. $G$ counts item-pair deviations from this pattern, $G^*$ normalizes $G$ to the range $[0,1]$, $U_3$ compares the observed vector with the reversed Guttman pattern, and $ZU_3$ standardizes $U_3$ to a unit-normal distribution. Because IRT orders items by estimated difficulty, a chatbot that solves hard items while missing easy ones inflates these statistics, marking it as aberrant relative to the human calibration sample.","core_discovery":"On its own terms, the paper establishes that person-fit statistics computed under item response theory—specifically $G$, $G^*$, $U_3$, and $ZU_3$—assign significantly higher aberrant-response scores to chatbot answer sheets than to human answer sheets. In every combination of two instruments and three chatbots, the difference was significant at $p < 0.00001$, with cleanly separated density distributions for the $G$ statistic shown. It further shows the three chatbots are not interchangeable: Kruskal-Wallis tests found differences among them, though which chatbot stands out depends on the instrument. The $ZU_3$ measure loses statistical significance at a 25% pollution level, while $G$ remains significant on some comparisons, indicating that the detectability of AI responses decreases as their prevalence increases.","pith_inferences":["The paper's clean separation partly reflects that each chatbot produced 20 full answer sheets with no human items mixed in; real cheating that blends human and AI answers would likely blur the PFS distributions, and the paper's data do not test that blend.","A testable extension: prompt the same models to explain their reasoning or to answer only a subset of items, to see whether person-fit separation survives more realistic cheating behavior.","If PFS-based screening were deployed, students who are simply weak, unusual, or have learning disabilities would risk being flagged; the paper itself names learning disabilities as a limitation, so a detector would need to rule out legitimate sources of misfit before accusing anyone.","The finding that chatbots differ in their reasoning profiles hints that IRT misfit could be used not only for detection but also for tracing which model generated a set of answers, for instance in forensic analysis of leaked answer sheets."],"forward_implications":["Instructors using MCQ platforms could flag answer sheets whose person-fit values fall above a threshold calibrated on the class's human response distribution.","The approach needs no text analysis and works on already machine-graded data, making it cheap to deploy at scale.","Because chatbots differ from each other, detectors may need per-model calibration rather than a single AI profile.","Detection degrades as AI responses approach 25% of the dataset, so the method is a deterrent against minority cheating, not a cure for systematic AI-generated test-taking.","The same IRT framework could in principle extend beyond MCQs to any binary-scored assessment, since the statistics depend only on item difficulty order and correctness scores."],"supporting_citations":[{"why":"Defines the Guttman pattern that the $G$ and $G^*$ statistics count deviations from.","marker":"[12]"},{"why":"Introduces the $G$ and $G^*$ person-fit statistics used as two of the four measures.","marker":"[28]"},{"why":"Introduces the $U_3$ and $ZU_3$ statistics used as the other two measures.","marker":"[29]"},{"why":"Methodological review defining person fit and the class of statistics the paper applies.","marker":"[18]"},{"why":"Supplies the 36-statistic comparison framework and the pollution-level sensitivity scheme the paper follows.","marker":"[13]"},{"why":"Prior attempt to flag ChatGPT 3.5 on chemistry MCQs using Rasch outfit, which this paper extends to general person-fit statistics on multiple models.","marker":"[24]"},{"why":"The PerFit R package used to compute the four statistics.","marker":"[27]"},{"why":"The IRT textbook grounding the theoretical assumption that response patterns follow ability-item-difficulty interaction.","marker":"[8]"}],"fun_headline_variants":["Person-fit stats unmask AI answers in multiple-choice tests","IRT spots chatbot cheaters in MCQs with high accuracy","AI responses seen clearly via IRT's person-fit measures","How IRT separates humans from ChatGPT, Claude, and Gemini","Person-fit scores reveal AI's distinct pattern on MCQs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The results assume that a student cheating with GenAI behaves like 20 full answer sheets produced by one chatbot prompted to output only the chosen options; real cheating that mixes human answers with AI answers, or uses varied prompts and model versions, may not produce the same misfit.","fun_headline_variants_meta":{"raw":{"variants":["Person-fit stats unmask AI answers in multiple-choice tests","IRT spots chatbot cheaters in MCQs with high accuracy","AI responses seen clearly via IRT's person-fit measures","How IRT separates humans from ChatGPT, Claude, and Gemini","Person-fit scores reveal AI's distinct pattern on MCQs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000156,"raw_usage":{"total_tokens":1185,"prompt_tokens":881,"completion_tokens":304,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":497,"completion_tokens_details":{"reasoning_tokens":222}},"tokens_in":497,"tokens_out":304,"duration_ms":3090,"temperature":1.0,"reasoning_tokens":222,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T10:37:04.789771+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a real class answer set and inject cheated answer sheets constructed by giving a chatbot only a random half of the items and letting a human answer the rest. If the PFS distributions of these blended sheets overlap those of honest students, the method fails on realistic partial cheating. A sharper version: run the same pipeline at a 10% pollution level on a third high-stakes instrument; if $ZU_3$'s separation already vanishes there, the operating range is even narrower than the paper reports.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the Guttman pattern that the $G$ and $G^*$ statistics count deviations from."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the $G$ and $G^*$ person-fit statistics used as two of the four measures."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the $U_3$ and $ZU_3$ statistics used as the other two measures."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Methodological review defining person fit and the class of statistics the paper applies."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the 36-statistic comparison framework and the pollution-level sensitivity scheme the paper follows."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior attempt to flag ChatGPT 3.5 on chemistry MCQs using Rasch outfit, which this paper extends to general person-fit statistics on multiple models."},{"cited_title":"Tendeiro, Rob R","cited_arxiv_id":null,"evidence_quote":"The PerFit R package used to compute the four statistics."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The IRT textbook grounding the theoretical assumption that response patterns follow ability-item-difficulty interaction."}],"review_version":1}