Pith. sign in

REVIEW 3 major objections 4 minor 16 references

Why AI Detection Fails for Academic Integrity

T0 review · 3 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Commercial AI detectors cannot distinguish light AI-assisted editing from full LLM drafts: at the 0.50 threshold they flag honest polish at 64-80 percent while missing over 96 percent of humanized rewrites, so scores should not be…

desk verdict Careful, policy-relevant measurement of detector failure modes, but the abstract oversells both the inability to distinguish editing from generation and how well the 'light edit' proxy represents compliant assistance. read the letter →

arxiv 2608.11256 v1 pith:MIPMHPFP submitted 2026-08-06 cs.LG cs.CY

classification cs.LGcs.CY
keywords AIdetectionacademicintegrityAI-assistedwritinghumanizationlargelanguagemodelsfalsepositiveratesnegativetextclassification
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that commercial AI detectors, evaluated at the standard 0.50 threshold under proxy labels, cannot separate light AI-assisted editing from full LLM drafting, so their scores are unsafe as standalone evidence of misconduct. Using 642 published English abstracts from four disciplines and two time windows, it shows that a light "refine abstract only" rewrite, the paper's proxy for a human author who keeps ownership of the claims and uses an LLM for polish, is flagged 64 to 80 percent of the time, while unmodified 2023-2025 abstracts are flagged 9 to 15 percent. After passing AI-labeled rewrites through a commercial humanizer, fewer than 4 percent are still flagged, a false-negative rate above 96 percent. The consequence is an integrity catch-22: honest AI-assisted editing carries higher sanction risk than humanizer-assisted evasion, and a sympathetic reader should care because institutions currently deploy these scores as default misconduct evidence.

What carries the argument

The load-bearing object is the controlled proxy-label corpus: 642 English abstracts from chemistry, computer science, political science, and theology, split into a pre-LLM window (2013-2015) and a recent window (2023-2025), with each abstract contributing an original plus three Gemini 3 Flash rewrites of escalating intensity, namely "refine (abstract only)", "refine (abstract + article)", and "new (article only)". The "refine (abstract only)" condition is the policy-critical instrument because it models a human author who retains the claims and uses an LLM merely to polish, and its 64-80 percent flag rate is the paper's measure of assisted-writing exposure. Error rates are computed with permutation tests clustered at the paper level, and the linguistic analysis uses Spearman correlations between detector scores and surface features such as long-token ratio and Academic Word List density.

What would settle it

Collect student essays with verified authorship histories and declared AI use, score them with the same commercial detectors at 0.50, and check whether light human edits are flagged at 64-80 percent while humanized AI drafts are missed more than 96 percent of the time; the paper's catch-22 claim collapses if any threshold keeps original flag rates low, light-edit capture low, and humanized-draft misses low. The paper itself shows no threshold in {0.4, 0.5, 0.6} does this, so an adjudicated classroom dataset is the direct test.

Watch

Extended reading notes

Core claim

On its own terms, the paper's discovery is a policy asymmetry measured under controlled proxy labels: at the standard 0.50 threshold, Pangram and GPTZero flag light AI-assisted edits of human abstracts at 64-80 percent while more than 96 percent of AI-generated rewrites escape detection after humanization. The flag rates on unmodified recent abstracts (15.0 percent Pangram, 8.9 percent GPTZero) are not confirmed false positives because real-world AI assistance is unobserved, and the paper is careful to label them as flag rates. Elevated scores track surface linguistic register, such as long-token and Academic Word List density, so non-STEM disciplines are flagged far more than STEM (p<0.001). The paper concludes that detector scores should not serve as standalone misconduct evidence and that flags must be corroborated by process evidence such as drafting history.

Load-bearing premise

Everything rests on the proxy-label scheme in which an LLM's light rewrite of a human abstract stands in for a human author's legitimate AI-assisted editing of their own work; if real classroom edits differ from the Gemini output used here, or if the 2023-2025 originals already contain unobserved AI assistance, the headline flag rates and evasion numbers would not transfer to actual integrity decisions.

Editorial extensions

If this is right

  • A detector score at or above 0.50 cannot by itself distinguish "student used AI to polish their own writing" from "student used AI to write the entire draft", so using it as standalone misconduct evidence will punish acceptable assistance.
  • Non-STEM students face substantially higher flag risk than STEM students at the same threshold, so uniform detector policies can create discipline-level disparities.
  • Because humanizer evasion drives detection below 4 percent, any enforcement regime that relies on detectors alone fails exactly at the most severe violation, namely a fully synthetic draft.
  • No threshold setting keeps original flag rates, light-edit capture, and AI-pool miss rates acceptable at once, so tuning the threshold does not fix the policy failure.
  • Institutions that keep detectors should pair any flag with process evidence such as drafting history and reserve sanctions for cases where human judgment supports lack of effort.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: if detector scores track long-token and Academic Word List density rather than authorship intent, then writers who adopt formal academic register, such as multilingual scholars, first-generation students, and novice writers, may be disproportionately flagged; this could be tested by comparing matched pairs of native and non-native writers with identical, verified AI-use histor
  • Editorial inference: the near-total post-humanization miss rate suggests an arms race in which threshold-based detectors degrade on both error types simultaneously as LLM and humanizer output distributions drift; a testable extension is longitudinal drift measurement on a fixed human-authored corpus.
  • Editorial inference: the paper's supplementary LLM-assisted baseline still flags most humanized rewrites, hinting that a judge reading for content and coherence rather than token statistics may be more robust than current commercial detectors; the paper does not make this claim.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper presents a controlled measurement study of two commercial AI detectors (Pangram 3.2 and GPTZero) on a corpus of 642 published English abstracts from four disciplines, drawn from two time windows (2013–2015 and 2023–2025). For each abstract, the authors generate three Gemini 3 Flash rewrites that differ in how much source text is used, score original and rewritten abstracts before and after humanization with Undetectable AI v11, and compute proxy-labeled flag rates, false-negative rates, feature correlations, threshold sweeps, and paper-clustered confidence intervals. The main empirical findings are that at tau = 0.50 the "refine (abstract only)" condition is flagged at 64–80%, unmodified 2023–2025 originals are flagged at 9–15% with large domain differences, and humanization makes more than 96% of AI-labeled rewrites undetected. The authors conclude that there is an "integrity catch-22" and that detector scores should not be used as standalone misconduct evidence.

Significance. The empirical core of the paper is valuable and unusually careful for this literature. The authors resample at the paper level (Appendix C), sweep thresholds (Appendix K), separate confirmed false positives from flag rates on 2023–2025 originals (Appendix A), and explicitly define two readings of the refine condition (Appendix B), so the paper's own caveats prevent the most common misreadings. If the measured asymmetry survives a more faithful operationalization of compliant assistance, the result is directly relevant to institutional integrity policy. The main technical strengths are the release of code, the use of direct detector outputs rather than fitted parameters, and the explicit proxy-label limitations; no circularity issue arises because the claims are stated conditionally on proxy labels.

major comments (3)
  1. [Abstract; Sections 1 and 3; Figures 2–3] The abstract and introduction claim that detectors "cannot distinguish AI editing from full LLM drafts." This is inconsistent with the paper's own ROC analyses: Figure 2 reports AUC-ROC of 0.98 (Pangram) and 0.92 (GPTZero) for refine (abstract only) versus originals on 2013–2015, and Figure 3 reports 0.91 on 2023–2025. The detectors separate the two classes substantially in ranking; what fails is threshold-based flagging at tau = 0.50, which forces a trade-off between false positives on assisted writing, false negatives on AI-labeled rewrites, and flags on originals (Table 5). The policy conclusion may survive, but the headline claim as written is an overstatement and should be revised to something like "detectors cannot reliably support a single threshold that separates AI-assisted editing from full LLM drafts without unacceptable error rates."
  2. [Section 2.2; Appendix B; Appendix D; Tables 2 and 4] The proxy for "guideline-compliant AI assistance" is a full-abstract rewrite. The prompt in Appendix D instructs Gemini 3 Flash to "rewrite a given paper abstract to sound better and more human-like" and to "make sure to make at least some edits or changes to the text," and Appendix B concedes that this prompt is "closer to light generative rewriting than to grammar-only editing." The headline flag rates of 64–80% and the asymmetry with humanizer evasion in Section 3 are therefore not directly a measurement of grammar-level or clarity-level assistance that most institutional policies permit; the GPTZero rates are materially lower (37.6–48.5%). This is load-bearing for the claim that "honest AI-editing results in a higher sanction risk than humanizer-assisted evasion." The authors should either add a minimal-edit/grammar-only condition, or restrict the policy claims to "light generative rewriting" and soften the abstract accordingly.
  3. [Section 3; Table 5; Section K.5] The threshold analysis in Table 5 and Appendix K correctly shows that no single threshold in {0.4, 0.5, 0.6} simultaneously keeps original flag rates, refine capture, and AI-pool FNR low. However, the policy conclusion is stated as if tau = 0.50 is the institutional operating point. Because the paper itself provides ROC curves, a more decision-relevant analysis would weight the three error types by plausible sanction costs (e.g., a false sanction on an honest author versus a missed cheating case) and identify the threshold range that is indefensible under a range of cost ratios. This would strengthen the central policy claim without changing the measurements.
minor comments (4)
  1. [Header/front matter] The ACM Reference Format line and copyright block contain placeholder text, including "2018" and "Conference acronym ’XX, Woodstock, NY"; these should be updated to the target venue and year before submission.
  2. [Appendix N; Section K.2] Several cross-references are unresolved: Appendix N contains "Appendix ??" and Section K.2 contains "Figure ??"; these should be fixed in the final version.
  3. [Table 2] The column header "FPR/flag" is ambiguous because the 2013–2015 entries are confirmed proxy false-positive rates while the 2023–2025 entries are flag rates on originals with unobserved AI use; the table should carry a more explicit header, as the main text correctly distinguishes these quantities.
  4. [Section 2.2; Appendix B] The main text uses "proxy AI-assisted false-positive risk" while Appendix B recommends the term "assisted-writing flag rate" for the refine (abstract only) condition; these terms should be harmonized in the main text and abstract to avoid the category error the authors themselves warn against.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: direct measurement study with explicit proxy labels and no fitted-parameter or self-citation chain.

full rationale

This paper contains no derivation chain in the circularity sense. Every headline quantity (proxy FPR, flag rate, FNR, refine flag rate, AUC, feature correlations) is a direct measurement of commercial detector outputs on a fixed corpus, with no parameter fitted to the target quantity and then re-reported as a prediction. The refine (abstract only) condition is labeled a proxy for guideline-compliant AI assistance, but Appendix A explicitly states 'Ground truth is experimental rather than adjudicated' and the paper reports 2023-2025 original flags as 'flag rates, not confirmed false positives'; Appendix B separately names the dual readings and warns against treating the 64-80% and >96% rates as cells of one confusion matrix. The proxy-labeling assumption can be criticized for external validity, but it is not circular because the claims are conditional on the proxy scheme and the detector scores themselves are measured, not derived from the labels. There are no load-bearing self-citations: the cited vendor reports are external benchmarks that the paper tests against, and no uniqueness theorem or prior result by the same authors is invoked to force the conclusion. The apparent tension between 'cannot distinguish' and the reported AUCs is a correctness or interpretation concern about threshold-level policy use versus ranking separability, not an input-output circularity.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

The central claims rest on labeled proxy data and on the representativeness of the chosen tools; the paper itself lists these as limitations. No fitted parameters appear. The detection threshold is a hand-chosen operating point and is included because the headline rates are threshold-dependent.

free parameters (1)
  • Detection threshold tau = 0.50 (hand-set, not fitted)
    Headline rates are reported at tau=0.50. Threshold sweeps in Appendix K show Pangram's pattern is stable, but GPTZero's refine (abstract only) flag rate is only 48.5% at this threshold, so the abstract's '64 to 80%' range is tied to this hand-chosen operating point.
assumptions (3)
  • domain assumption Proxy labels: unmodified originals are human and all LLM rewrites are AI.
    Central to all FPR/FNR calculations. The paper acknowledges in Appendix A that ground truth is experimental rather than adjudicated, and that 2023-2025 originals may contain unobserved AI assistance.
  • domain assumption Published abstracts are a valid proxy for student essays and classroom genres.
    The policy conclusion is about academic integrity in student work, but the corpus is published research abstracts. The paper flags this in Limitations.
  • domain assumption Gemini 3 Flash, Pangram 3.2, GPTZero, and Undetectable AI v11 are representative of the broader tool landscape.
    Results may not generalize to other LLM generators, detector versions, humanizers, or prompting strategies. The paper explicitly lists this limitation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Why AI Detection Fails for Academic Integrity." pith.science (2026). https://pith.science/paper/MIPMHPFP

@misc{pith2026260811256,
  author       = {Pith},
  title        = {Pith review of: Why AI Detection Fails for Academic Integrity},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MIPMHPFP}},
  note         = {Machine review of arXiv:2608.11256}
}
read the original abstract

Institutions use commercial AI detectors for academic integrity, yet detectors cannot distinguish AI editing from full LLM drafts and may treat both as misconduct. In a controlled study of published English abstracts (four domains; 2013 to 2015 vs. 2023 to 2025), we quantify this policy failure under proxy human/AI labels at tau=0.50. Light "refine abstract only" edits, a proxy for guideline-compliant AI assistance, are flagged at 64 to 80% (Pangram/GPTZero). Unmodified 2023 to 2025 originals are flagged at 9 to 15%, with non-STEM rates far above STEM (p<0.001); elevated scores track long-token and Academic Word List density, not authorship intent alone. After Undetectable AI humanization, evasion is near-total: fewer than 4% of AI-labeled rewrites remain flagged (post-humanization detection rate <4%; FNR >96%). Honest AI-editing results in a higher sanction risk than humanizer-assisted evasion. Therefore, detector scores should not serve as standalone misconduct evidence.

Figures

Figures reproduced from arXiv: 2608.11256 by the authors.

Figure 1
Figure 1. AI Detection Pipeline RQ1 Under proxy ground-truth labels, what are the false-positive and false-negative rates of commercial detectors, and how do they vary by domain, time period, and rewrite condition? RQ2 Which surface linguistic features of academic abstracts are associated with elevated detector scores (and thus higher false-positive risk on human writing)? RQ3 How does humanization change false-negative rates… view at source ↗
Figure 2
Figure 2. Assisted-editing ROC (2013-2015): original vs. refine (abstract only); 𝜏 = 0.50 [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Assisted-editing ROC/PR (2023-2025, pre [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Pangram scores before humanization (2023 to 2025), by domain (rows) and rewrite condition (columns). Originals [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]
Figure 5
Figure 5. Figure 5: Pangram scores after humanization (2023 to 2025). Compare to Figure 4: high pre-humanization scores on AI-labeled [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]
Figure 6
Figure 6. Figure 6: ROC and precision-recall by rewrite condition vs. [PITH_FULL_IMAGE:figures/full_fig_p014_6.png]
Figure 7
Figure 7. Figure 7: Rewrite-condition ROC/PR on 2023 to 2025 (pre-humanization); same layout as Figure 6. [PITH_FULL_IMAGE:figures/full_fig_p015_7.png]
Figure 8
Figure 8. Figure 8: Threshold trade-off (2023 to 2025, pre-humanization). Three policy rates vs. [PITH_FULL_IMAGE:figures/full_fig_p016_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

16 extracted references · 12 canonical work pages

  1. [2]

    LLM-Generated or Human-Written? Compar- ing Review and Non-Review Papers on ArXiv

    “LLM-Generated or Human-Written? Compar- ing Review and Non-Review Papers on ArXiv. ”arXiv preprint arXiv:2601.17036. Bradley Emi and Max Spero

  2. [5]

    Monitoring ai-modified content at scale: A case study on the impact of chatgpt on ai conference peer reviews

    “Monitoring ai-modified content at scale: A case study on the impact of chatgpt on ai conference peer reviews. ”arXiv preprint arXiv:2403.07183. OpenRouter. 2026.OpenRouter: The Unified Interface for LLMs. (2026). https://openrout er.ai/. Mike Perkins

  3. [6]

    #1 Best AI Detector

    but forms CIs by resampling papers. Long-token and AWL associations remain positive with intervals excluding zero for both detectors (e.g., Pangram AWL 𝜌= 0.346 [0.298, 0.395]). Paper-averaged𝜌 is attenuated because averaging mixes originals and rewrites within paper; we report it for trans- parency, not as a replacement for the instance-level estimand. D...

  4. [7]

    Trusting AI to detect AI? A systematic evaluation of the reliability and robustness of current AIGC detection tools for student academic work

    “Trusting AI to detect AI? A systematic evaluation of the reliability and robustness of current AIGC detection tools for student academic work. ”Computers & Education, 105616. Undetectable AI. 2026a.Home Page. https://undetectable.ai/. Undetectable AI. 2026b.Pricing. (2026). https://undetectable.ai/pricing. Adrian Wallwork

  5. [12]

    G.2 Corpus counts See Table 1 in the main text

    Residual variance remains in non-STEM domains: political science and theology still show non-zero SD and outliers on some conditions (Tables 17 and 18), indicating incomplete score collapse for a subset of ab- stracts rather than uniform evasion. G.2 Corpus counts See Table 1 in the main text. H Full text-feature correlation results Tables 19 and 20 repor...

  6. [14]

    On recent prose this mixes proxy false positives with unobserved real AI use; it still bounds how often human-labeled text is flagged

    K.5 Threshold trade-off curves For each detector and each 𝜏∈ {0.3, 0.4, 0.5, 0.6, 0.7}, we report three percentages on the 2023 to 2025 collection: (1) Original flag rate(gray): fraction of unmodifiedoriginal abstracts with𝑠≥𝜏 . On recent prose this mixes proxy false positives with unobserved real AI use; it still bounds how often human-labeled text is fl...

  7. [15]

    Systematic suppression of detector-salient cues.Table 8 in the main text summarizes key pooled shifts; Table 22 below lists the full feature set

    and corre- lated feature deltas with Pangram and GPTZero score deltas on matched rows. Systematic suppression of detector-salient cues.Table 8 in the main text summarizes key pooled shifts; Table 22 below lists the full feature set. Humanizationreducedlong-token ratio (0.184 to 0.138; mean Δ=− 0.045; 88.2% of pairs decreased;𝑝< 0.001), AWL token ratio (0....

  8. [16]

    Length and syntax.Humanization alsoincreasedlength: mean word count rose from 180 to 229 (Δ≈+ 49words; 𝑝< 0.001) and average words per sentence from 23.6 to 28.4 (Δ≈+ 4.8;𝑝< 0.001)

    AWL density (Appendix H): the humanizer moves text along the same axes that separate high-scoring from low-scoring abstracts. Length and syntax.Humanization alsoincreasedlength: mean word count rose from 180 to 229 (Δ≈+ 49words; 𝑝< 0.001) and average words per sentence from 23.6 to 28.4 (Δ≈+ 4.8;𝑝< 0.001). Short-sentence ratio increased only slightly (0.0...

Show all 16 references
  1. [18]

    flagged as not human

    further shows that pooled benchmark accuracy can mask large field-to-field variance. F Definitions of text ratios We compute several token-level ratios from each abstract to probe whether surface properties correlate with detector scores. Let𝑡 be the abstract text (a string). ...

  2. [42]

    at𝜏=0.50. Period DetectorΔmean Flag STEM Flag non-STEM𝑛STEM/non𝑝(paper) 2013-2015 gptzero−0.0060.0% 0.0% 154/152 0.026 2013-2015 pangram0.0010.0% 0.0% 153/152 0.62 2023-2025 gptzero0.1103.4% 15.2% 178/158 0.0002 2023-2025 pangram0.2045.6% 25.5% 177/157 0.0002 Table 10: STEM vs...

  3. [2015]

    AUC-PR is high (0.93-1.00) because Gemini rewrites separate cleanly from unflagged originals at𝜏= 0.50; this is a best-case separability benchmark, not current flagging exposure

    Figure 6 repeats the analysis for each rewrite type against the same pre-LLM-era originals:refine (abstract only),refine (abstract + paper), andnew (article only), plus a dashedpooled AIcurve (all three LLM conditions as positives). AUC-PR is high (0.93-1.00) because Gemini re...

  4. [2022]

    OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts

    “OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts. ”arXiv preprint arXiv:2205.01833. Quinnipiac Innovations in Learning and Teaching (QILT). n.d.Academic Integrity and AI Detectors. https://qilt.qu.edu/ai/academic-integrity-and-ai-de...

  5. [2023]

    Don’t trust ChatGPT when your question is not in English: A study of multilingual abilities and types of LLMs

    “Don’t trust ChatGPT when your question is not in English: A study of multilingual abilities and types of LLMs. ” In:Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 7915–7927. Appendix A Limitations Language and corpus setting:Our datase...

  6. [2024]

    Technical report on the pangram ai-generated text classifier

    “Technical report on the pangram ai-generated text classifier. ”arXiv preprint arXiv:2402.14873. Mingmeng Geng and Thierry Poibeau

  7. [2025]

    What Are We Detecting, Really? LLM- Generated Text Detection Remains an Unsolved Problem

    “What Are We Detecting, Really? LLM- Generated Text Detection Remains an Unsolved Problem. ” GPTZero. 2026.GPTZero AI Detection Benchmarking: The Industry Standard in Accuracy, Transparency and Fairness. (2026). https://gptzero.me/news/gptzero-ai-detection-b enchmarking-the-in...

  8. [2026]

    GPTZero AI Detection Benchmarking: The Industry Standard in Accuracy, Transparency and Fairness,

    “GPTZero AI Detection Benchmarking: The Industry Standard in Accuracy, Transparency and Fairness, ” (Feb. 2026). https://gptzero.me/news/gpt zero-ai-detection-benchmarking-the-industry-standard-in-accuracy-transparen cy-and-fairness/. Debby RE Cotton, Peter A Cotton, and J Reu...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.