REVIEW 3 major objections 4 minor 16 references
Why AI Detection Fails for Academic Integrity
T0 review · 3 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Commercial AI detectors cannot distinguish light AI-assisted editing from full LLM drafts: at the 0.50 threshold they flag honest polish at 64-80 percent while missing over 96 percent of humanized rewrites, so scores should not be…
desk verdict Careful, policy-relevant measurement of detector failure modes, but the abstract oversells both the inability to distinguish editing from generation and how well the 'light edit' proxy represents compliant assistance. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the controlled proxy-label corpus: 642 English abstracts from chemistry, computer science, political science, and theology, split into a pre-LLM window (2013-2015) and a recent window (2023-2025), with each abstract contributing an original plus three Gemini 3 Flash rewrites of escalating intensity, namely "refine (abstract only)", "refine (abstract + article)", and "new (article only)". The "refine (abstract only)" condition is the policy-critical instrument because it models a human author who retains the claims and uses an LLM merely to polish, and its 64-80 percent flag rate is the paper's measure of assisted-writing exposure. Error rates are computed with permutation tests clustered at the paper level, and the linguistic analysis uses Spearman correlations between detector scores and surface features such as long-token ratio and Academic Word List density.
What would settle it
Collect student essays with verified authorship histories and declared AI use, score them with the same commercial detectors at 0.50, and check whether light human edits are flagged at 64-80 percent while humanized AI drafts are missed more than 96 percent of the time; the paper's catch-22 claim collapses if any threshold keeps original flag rates low, light-edit capture low, and humanized-draft misses low. The paper itself shows no threshold in {0.4, 0.5, 0.6} does this, so an adjudicated classroom dataset is the direct test.
Extended reading notes
Core claim
On its own terms, the paper's discovery is a policy asymmetry measured under controlled proxy labels: at the standard 0.50 threshold, Pangram and GPTZero flag light AI-assisted edits of human abstracts at 64-80 percent while more than 96 percent of AI-generated rewrites escape detection after humanization. The flag rates on unmodified recent abstracts (15.0 percent Pangram, 8.9 percent GPTZero) are not confirmed false positives because real-world AI assistance is unobserved, and the paper is careful to label them as flag rates. Elevated scores track surface linguistic register, such as long-token and Academic Word List density, so non-STEM disciplines are flagged far more than STEM (p<0.001). The paper concludes that detector scores should not serve as standalone misconduct evidence and that flags must be corroborated by process evidence such as drafting history.
Load-bearing premise
Everything rests on the proxy-label scheme in which an LLM's light rewrite of a human abstract stands in for a human author's legitimate AI-assisted editing of their own work; if real classroom edits differ from the Gemini output used here, or if the 2023-2025 originals already contain unobserved AI assistance, the headline flag rates and evasion numbers would not transfer to actual integrity decisions.
Editorial extensions
If this is right
- A detector score at or above 0.50 cannot by itself distinguish "student used AI to polish their own writing" from "student used AI to write the entire draft", so using it as standalone misconduct evidence will punish acceptable assistance.
- Non-STEM students face substantially higher flag risk than STEM students at the same threshold, so uniform detector policies can create discipline-level disparities.
- Because humanizer evasion drives detection below 4 percent, any enforcement regime that relies on detectors alone fails exactly at the most severe violation, namely a fully synthetic draft.
- No threshold setting keeps original flag rates, light-edit capture, and AI-pool miss rates acceptable at once, so tuning the threshold does not fix the policy failure.
- Institutions that keep detectors should pair any flag with process evidence such as drafting history and reserve sanctions for cases where human judgment supports lack of effort.
Reading between the lines
- Editorial inference: if detector scores track long-token and Academic Word List density rather than authorship intent, then writers who adopt formal academic register, such as multilingual scholars, first-generation students, and novice writers, may be disproportionately flagged; this could be tested by comparing matched pairs of native and non-native writers with identical, verified AI-use histor
- Editorial inference: the near-total post-humanization miss rate suggests an arms race in which threshold-based detectors degrade on both error types simultaneously as LLM and humanizer output distributions drift; a testable extension is longitudinal drift measurement on a fixed human-authored corpus.
- Editorial inference: the paper's supplementary LLM-assisted baseline still flags most humanized rewrites, hinting that a judge reading for content and coherence rather than token statistics may be more robust than current commercial detectors; the paper does not make this claim.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents a controlled measurement study of two commercial AI detectors (Pangram 3.2 and GPTZero) on a corpus of 642 published English abstracts from four disciplines, drawn from two time windows (2013–2015 and 2023–2025). For each abstract, the authors generate three Gemini 3 Flash rewrites that differ in how much source text is used, score original and rewritten abstracts before and after humanization with Undetectable AI v11, and compute proxy-labeled flag rates, false-negative rates, feature correlations, threshold sweeps, and paper-clustered confidence intervals. The main empirical findings are that at tau = 0.50 the "refine (abstract only)" condition is flagged at 64–80%, unmodified 2023–2025 originals are flagged at 9–15% with large domain differences, and humanization makes more than 96% of AI-labeled rewrites undetected. The authors conclude that there is an "integrity catch-22" and that detector scores should not be used as standalone misconduct evidence.
Significance. The empirical core of the paper is valuable and unusually careful for this literature. The authors resample at the paper level (Appendix C), sweep thresholds (Appendix K), separate confirmed false positives from flag rates on 2023–2025 originals (Appendix A), and explicitly define two readings of the refine condition (Appendix B), so the paper's own caveats prevent the most common misreadings. If the measured asymmetry survives a more faithful operationalization of compliant assistance, the result is directly relevant to institutional integrity policy. The main technical strengths are the release of code, the use of direct detector outputs rather than fitted parameters, and the explicit proxy-label limitations; no circularity issue arises because the claims are stated conditionally on proxy labels.
major comments (3)
- [Abstract; Sections 1 and 3; Figures 2–3] The abstract and introduction claim that detectors "cannot distinguish AI editing from full LLM drafts." This is inconsistent with the paper's own ROC analyses: Figure 2 reports AUC-ROC of 0.98 (Pangram) and 0.92 (GPTZero) for refine (abstract only) versus originals on 2013–2015, and Figure 3 reports 0.91 on 2023–2025. The detectors separate the two classes substantially in ranking; what fails is threshold-based flagging at tau = 0.50, which forces a trade-off between false positives on assisted writing, false negatives on AI-labeled rewrites, and flags on originals (Table 5). The policy conclusion may survive, but the headline claim as written is an overstatement and should be revised to something like "detectors cannot reliably support a single threshold that separates AI-assisted editing from full LLM drafts without unacceptable error rates."
- [Section 2.2; Appendix B; Appendix D; Tables 2 and 4] The proxy for "guideline-compliant AI assistance" is a full-abstract rewrite. The prompt in Appendix D instructs Gemini 3 Flash to "rewrite a given paper abstract to sound better and more human-like" and to "make sure to make at least some edits or changes to the text," and Appendix B concedes that this prompt is "closer to light generative rewriting than to grammar-only editing." The headline flag rates of 64–80% and the asymmetry with humanizer evasion in Section 3 are therefore not directly a measurement of grammar-level or clarity-level assistance that most institutional policies permit; the GPTZero rates are materially lower (37.6–48.5%). This is load-bearing for the claim that "honest AI-editing results in a higher sanction risk than humanizer-assisted evasion." The authors should either add a minimal-edit/grammar-only condition, or restrict the policy claims to "light generative rewriting" and soften the abstract accordingly.
- [Section 3; Table 5; Section K.5] The threshold analysis in Table 5 and Appendix K correctly shows that no single threshold in {0.4, 0.5, 0.6} simultaneously keeps original flag rates, refine capture, and AI-pool FNR low. However, the policy conclusion is stated as if tau = 0.50 is the institutional operating point. Because the paper itself provides ROC curves, a more decision-relevant analysis would weight the three error types by plausible sanction costs (e.g., a false sanction on an honest author versus a missed cheating case) and identify the threshold range that is indefensible under a range of cost ratios. This would strengthen the central policy claim without changing the measurements.
minor comments (4)
- [Header/front matter] The ACM Reference Format line and copyright block contain placeholder text, including "2018" and "Conference acronym ’XX, Woodstock, NY"; these should be updated to the target venue and year before submission.
- [Appendix N; Section K.2] Several cross-references are unresolved: Appendix N contains "Appendix ??" and Section K.2 contains "Figure ??"; these should be fixed in the final version.
- [Table 2] The column header "FPR/flag" is ambiguous because the 2013–2015 entries are confirmed proxy false-positive rates while the 2023–2025 entries are flag rates on originals with unobserved AI use; the table should carry a more explicit header, as the main text correctly distinguishes these quantities.
- [Section 2.2; Appendix B] The main text uses "proxy AI-assisted false-positive risk" while Appendix B recommends the term "assisted-writing flag rate" for the refine (abstract only) condition; these terms should be harmonized in the main text and abstract to avoid the category error the authors themselves warn against.
Circularity Check
No significant circularity: direct measurement study with explicit proxy labels and no fitted-parameter or self-citation chain.
full rationale
This paper contains no derivation chain in the circularity sense. Every headline quantity (proxy FPR, flag rate, FNR, refine flag rate, AUC, feature correlations) is a direct measurement of commercial detector outputs on a fixed corpus, with no parameter fitted to the target quantity and then re-reported as a prediction. The refine (abstract only) condition is labeled a proxy for guideline-compliant AI assistance, but Appendix A explicitly states 'Ground truth is experimental rather than adjudicated' and the paper reports 2023-2025 original flags as 'flag rates, not confirmed false positives'; Appendix B separately names the dual readings and warns against treating the 64-80% and >96% rates as cells of one confusion matrix. The proxy-labeling assumption can be criticized for external validity, but it is not circular because the claims are conditional on the proxy scheme and the detector scores themselves are measured, not derived from the labels. There are no load-bearing self-citations: the cited vendor reports are external benchmarks that the paper tests against, and no uniqueness theorem or prior result by the same authors is invoked to force the conclusion. The apparent tension between 'cannot distinguish' and the reported AUCs is a correctness or interpretation concern about threshold-level policy use versus ranking separability, not an input-output circularity.
Assumptions & free parameters
free parameters (1)
- Detection threshold tau =
0.50 (hand-set, not fitted)
assumptions (3)
- domain assumption Proxy labels: unmodified originals are human and all LLM rewrites are AI.
- domain assumption Published abstracts are a valid proxy for student essays and classroom genres.
- domain assumption Gemini 3 Flash, Pangram 3.2, GPTZero, and Undetectable AI v11 are representative of the broader tool landscape.
Cite this review
Pith. "Pith review of Why AI Detection Fails for Academic Integrity." pith.science (2026). https://pith.science/paper/MIPMHPFP
@misc{pith2026260811256,
author = {Pith},
title = {Pith review of: Why AI Detection Fails for Academic Integrity},
year = {2026},
howpublished = {\url{https://pith.science/paper/MIPMHPFP}},
note = {Machine review of arXiv:2608.11256}
}
read the original abstract
Institutions use commercial AI detectors for academic integrity, yet detectors cannot distinguish AI editing from full LLM drafts and may treat both as misconduct. In a controlled study of published English abstracts (four domains; 2013 to 2015 vs. 2023 to 2025), we quantify this policy failure under proxy human/AI labels at tau=0.50. Light "refine abstract only" edits, a proxy for guideline-compliant AI assistance, are flagged at 64 to 80% (Pangram/GPTZero). Unmodified 2023 to 2025 originals are flagged at 9 to 15%, with non-STEM rates far above STEM (p<0.001); elevated scores track long-token and Academic Word List density, not authorship intent alone. After Undetectable AI humanization, evasion is near-total: fewer than 4% of AI-labeled rewrites remain flagged (post-humanization detection rate <4%; FNR >96%). Honest AI-editing results in a higher sanction risk than humanizer-assisted evasion. Therefore, detector scores should not serve as standalone misconduct evidence.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[2]
LLM-Generated or Human-Written? Compar- ing Review and Non-Review Papers on ArXiv
“LLM-Generated or Human-Written? Compar- ing Review and Non-Review Papers on ArXiv. ”arXiv preprint arXiv:2601.17036. Bradley Emi and Max Spero
-
[5]
“Monitoring ai-modified content at scale: A case study on the impact of chatgpt on ai conference peer reviews. ”arXiv preprint arXiv:2403.07183. OpenRouter. 2026.OpenRouter: The Unified Interface for LLMs. (2026). https://openrout er.ai/. Mike Perkins
arXiv 2026
-
[6]
but forms CIs by resampling papers. Long-token and AWL associations remain positive with intervals excluding zero for both detectors (e.g., Pangram AWL 𝜌= 0.346 [0.298, 0.395]). Paper-averaged𝜌 is attenuated because averaging mixes originals and rewrites within paper; we report it for trans- parency, not as a replacement for the instance-level estimand. D...
work page 2024
-
[7]
“Trusting AI to detect AI? A systematic evaluation of the reliability and robustness of current AIGC detection tools for student academic work. ”Computers & Education, 105616. Undetectable AI. 2026a.Home Page. https://undetectable.ai/. Undetectable AI. 2026b.Pricing. (2026). https://undetectable.ai/pricing. Adrian Wallwork
work page 2026
-
[12]
G.2 Corpus counts See Table 1 in the main text
Residual variance remains in non-STEM domains: political science and theology still show non-zero SD and outliers on some conditions (Tables 17 and 18), indicating incomplete score collapse for a subset of ab- stracts rather than uniform evasion. G.2 Corpus counts See Table 1 in the main text. H Full text-feature correlation results Tables 19 and 20 repor...
work page 2023
-
[14]
K.5 Threshold trade-off curves For each detector and each 𝜏∈ {0.3, 0.4, 0.5, 0.6, 0.7}, we report three percentages on the 2023 to 2025 collection: (1) Original flag rate(gray): fraction of unmodifiedoriginal abstracts with𝑠≥𝜏 . On recent prose this mixes proxy false positives with unobserved real AI use; it still bounds how often human-labeled text is fl...
work page 2023
-
[15]
and corre- lated feature deltas with Pangram and GPTZero score deltas on matched rows. Systematic suppression of detector-salient cues.Table 8 in the main text summarizes key pooled shifts; Table 22 below lists the full feature set. Humanizationreducedlong-token ratio (0.184 to 0.138; mean Δ=− 0.045; 88.2% of pairs decreased;𝑝< 0.001), AWL token ratio (0....
work page 2018
-
[16]
AWL density (Appendix H): the humanizer moves text along the same axes that separate high-scoring from low-scoring abstracts. Length and syntax.Humanization alsoincreasedlength: mean word count rose from 180 to 229 (Δ≈+ 49words; 𝑝< 0.001) and average words per sentence from 23.6 to 28.4 (Δ≈+ 4.8;𝑝< 0.001). Short-sentence ratio increased only slightly (0.0...
Show all 16 references
-
[18]
flagged as not human
further shows that pooled benchmark accuracy can mask large field-to-field variance. F Definitions of text ratios We compute several token-level ratios from each abstract to probe whether surface properties correlate with detector scores. Let𝑡 be the abstract text (a string). ...
2024
-
[42]
at𝜏=0.50. Period DetectorΔmean Flag STEM Flag non-STEM𝑛STEM/non𝑝(paper) 2013-2015 gptzero−0.0060.0% 0.0% 154/152 0.026 2013-2015 pangram0.0010.0% 0.0% 153/152 0.62 2023-2025 gptzero0.1103.4% 15.2% 178/158 0.0002 2023-2025 pangram0.2045.6% 25.5% 177/157 0.0002 Table 10: STEM vs...
2013
-
[2015]
AUC-PR is high (0.93-1.00) because Gemini rewrites separate cleanly from unflagged originals at𝜏= 0.50; this is a best-case separability benchmark, not current flagging exposure
Figure 6 repeats the analysis for each rewrite type against the same pre-LLM-era originals:refine (abstract only),refine (abstract + paper), andnew (article only), plus a dashedpooled AIcurve (all three LLM conditions as positives). AUC-PR is high (0.93-1.00) because Gemini re...
2023
-
[2022]
OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts
“OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts. ”arXiv preprint arXiv:2205.01833. Quinnipiac Innovations in Learning and Teaching (QILT). n.d.Academic Integrity and AI Detectors. https://qilt.qu.edu/ai/academic-integrity-and-ai-de...
2025 arXiv
-
[2023]
Don’t trust ChatGPT when your question is not in English: A study of multilingual abilities and types of LLMs
“Don’t trust ChatGPT when your question is not in English: A study of multilingual abilities and types of LLMs. ” In:Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 7915–7927. Appendix A Limitations Language and corpus setting:Our datase...
2023
-
[2024]
Technical report on the pangram ai-generated text classifier
“Technical report on the pangram ai-generated text classifier. ”arXiv preprint arXiv:2402.14873. Mingmeng Geng and Thierry Poibeau
-
[2025]
What Are We Detecting, Really? LLM- Generated Text Detection Remains an Unsolved Problem
“What Are We Detecting, Really? LLM- Generated Text Detection Remains an Unsolved Problem. ” GPTZero. 2026.GPTZero AI Detection Benchmarking: The Industry Standard in Accuracy, Transparency and Fairness. (2026). https://gptzero.me/news/gptzero-ai-detection-b enchmarking-the-in...
2026
-
[2026]
GPTZero AI Detection Benchmarking: The Industry Standard in Accuracy, Transparency and Fairness,
“GPTZero AI Detection Benchmarking: The Industry Standard in Accuracy, Transparency and Fairness, ” (Feb. 2026). https://gptzero.me/news/gpt zero-ai-detection-benchmarking-the-industry-standard-in-accuracy-transparen cy-and-fairness/. Debby RE Cotton, Peter A Cotton, and J Reu...
2026
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.