{"id":"5fed2794-f24e-45a6-904b-c2885d1c1a4c","arxiv_id":"2608.11256","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"At a 0.50 threshold, two commercial detectors flag most light AI-assisted abstract edits and yet miss over 96% of humanized AI rewrites, so detector flags do not separate assistance from generation.","lead":"This study shows that commercial AI detectors flag light, guideline-compliant AI editing of scholarly abstracts at high rates, while a commercial humanizer lets almost all fully AI-generated rewrites evade detection. The authors argue detector scores should not be used as standalone evidence of misconduct.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proxy 'light refine' condition is a full-abstract rewrite, not typical grammar-level assistance, so the 64–80% honest-editing flag rate may overstate sanction risk.","rationale":"The reader's weakest_assumption is that the proxy-label scheme may not transfer to real classroom edits or may be contaminated by unobserved AI assistance in 2023–2025 originals. My concern sharpens this: the refine (abstract only) prompt is not a light edit; it is a full rewrite that instructs the model to 'sound better and more human-like.' This means the measured 64–80% flag rate is likely an upper bound on the risk from genuinely compliant assistance (grammar, clarity, word choice), and the abstract's framing as 'guideline-compliant AI assistance' is misleading. This is load-bearing because the paper's strongest policy claim is the asymmetry between honest editing and humanizer-assisted evasion. If honest editing at the level actually permitted by most guidelines flags at a much lower rate, the asymmetry is weaker. I agree with the reader that the paper is otherwise careful: the paper-clustered inference (Appendix C), the explicit dual interpretation in Appendix B, and the candid limitations in Appendix A are all strengths. The threshold sweep (Table 5) and the domain-level analyses support the general conclusion that detector scores are not reliable standalone evidence. The fix is primarily in framing: replace 'cannot distinguish' with 'cannot reliably separate at policy-relevant thresholds' and clearly state that the refine condition is a full-abstract rewrite, not grammar-only assistance. These are qualification changes, not foundational errors, so the conditional verdict stands unchanged.","tokens_in":19317,"tokens_out":5835,"duration_ms":50967,"concrete_test":"Re-score the 2013–2015 abstracts after generating a grammar-only LLM edit (e.g., 'Fix grammar and clarity; do not change content or style') using the same detectors and tau=0.50, and compare flag rates to the 64–80% refine rates. If the grammar-only flag rate is materially below the refine rate (e.g., below 30% for Pangram), the proxy overstates the risk to guideline-compliant assistance and the asymmetry claim needs qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central policy asymmetry—'honest AI-editing results in a higher sanction risk than humanizer-assisted evasion'—rests on the claim that the refine (abstract only) condition is a proxy for guideline-compliant AI assistance. But the prompt in Appendix D instructs Gemini 3 Flash to 'rewrite a given paper abstract to sound better and more human-like' and to 'make at least some edits or changes to the text.' This is a substantial paraphrase, not the grammar/clarity polish that most institutional guidelines permit. Appendix B concedes that the refine prompt is 'closer to light generative rewriting than to grammar-only editing.' If real compliant assistance produces much smaller surface changes, then the 64–80% flag rates (which are almost entirely driven by Pangram; GPTZero rates are 37.6–48.5%) may not reflect the sanction risk for honest editing. The paper's only unambiguously human texts (2013–2015 originals) have 0% FPR at tau=0.50, so the entire 'false-positive risk on assistance' argument depends on this proxy. If the proxy overstates the transformation, the asymmetry could shrink or invert. Additionally, the abstract's 'cannot distinguish' is internally inconsistent with the reported ROC AUCs of 0.91–0.98 for assisted editing versus originals: detectors show substantial separability in ranking, even if no threshold simultaneously achieves low original-flag, high refine-capture, and low FNR. The policy conclusion may still hold, but the headline claim is overstated as written.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a controlled measurement study of two commercial AI detectors (Pangram 3.2 and GPTZero) on a corpus of 642 published English abstracts from four disciplines, drawn from two time windows (2013–2015 and 2023–2025). For each abstract, the authors generate three Gemini 3 Flash rewrites that differ in how much source text is used, score original and rewritten abstracts before and after humanization with Undetectable AI v11, and compute proxy-labeled flag rates, false-negative rates, feature correlations, threshold sweeps, and paper-clustered confidence intervals. The main empirical findings are that at tau = 0.50 the \"refine (abstract only)\" condition is flagged at 64–80%, unmodified 2023–2025 originals are flagged at 9–15% with large domain differences, and humanization makes more than 96% of AI-labeled rewrites undetected. The authors conclude that there is an \"integrity catch-22\" and that detector scores should not be used as standalone misconduct evidence.","tokens_in":19606,"tokens_out":6332,"duration_ms":54034,"significance":"The empirical core of the paper is valuable and unusually careful for this literature. The authors resample at the paper level (Appendix C), sweep thresholds (Appendix K), separate confirmed false positives from flag rates on 2023–2025 originals (Appendix A), and explicitly define two readings of the refine condition (Appendix B), so the paper's own caveats prevent the most common misreadings. If the measured asymmetry survives a more faithful operationalization of compliant assistance, the result is directly relevant to institutional integrity policy. The main technical strengths are the release of code, the use of direct detector outputs rather than fitted parameters, and the explicit proxy-label limitations; no circularity issue arises because the claims are stated conditionally on proxy labels.","major_comments":[{"comment":"The abstract and introduction claim that detectors \"cannot distinguish AI editing from full LLM drafts.\" This is inconsistent with the paper's own ROC analyses: Figure 2 reports AUC-ROC of 0.98 (Pangram) and 0.92 (GPTZero) for refine (abstract only) versus originals on 2013–2015, and Figure 3 reports 0.91 on 2023–2025. The detectors separate the two classes substantially in ranking; what fails is threshold-based flagging at tau = 0.50, which forces a trade-off between false positives on assisted writing, false negatives on AI-labeled rewrites, and flags on originals (Table 5). The policy conclusion may survive, but the headline claim as written is an overstatement and should be revised to something like \"detectors cannot reliably support a single threshold that separates AI-assisted editing from full LLM drafts without unacceptable error rates.\"","section":"Abstract; Sections 1 and 3; Figures 2–3"},{"comment":"The proxy for \"guideline-compliant AI assistance\" is a full-abstract rewrite. The prompt in Appendix D instructs Gemini 3 Flash to \"rewrite a given paper abstract to sound better and more human-like\" and to \"make sure to make at least some edits or changes to the text,\" and Appendix B concedes that this prompt is \"closer to light generative rewriting than to grammar-only editing.\" The headline flag rates of 64–80% and the asymmetry with humanizer evasion in Section 3 are therefore not directly a measurement of grammar-level or clarity-level assistance that most institutional policies permit; the GPTZero rates are materially lower (37.6–48.5%). This is load-bearing for the claim that \"honest AI-editing results in a higher sanction risk than humanizer-assisted evasion.\" The authors should either add a minimal-edit/grammar-only condition, or restrict the policy claims to \"light generative rewriting\" and soften the abstract accordingly.","section":"Section 2.2; Appendix B; Appendix D; Tables 2 and 4"},{"comment":"The threshold analysis in Table 5 and Appendix K correctly shows that no single threshold in {0.4, 0.5, 0.6} simultaneously keeps original flag rates, refine capture, and AI-pool FNR low. However, the policy conclusion is stated as if tau = 0.50 is the institutional operating point. Because the paper itself provides ROC curves, a more decision-relevant analysis would weight the three error types by plausible sanction costs (e.g., a false sanction on an honest author versus a missed cheating case) and identify the threshold range that is indefensible under a range of cost ratios. This would strengthen the central policy claim without changing the measurements.","section":"Section 3; Table 5; Section K.5"}],"minor_comments":[{"comment":"The ACM Reference Format line and copyright block contain placeholder text, including \"2018\" and \"Conference acronym ’XX, Woodstock, NY\"; these should be updated to the target venue and year before submission.","section":"Header/front matter"},{"comment":"Several cross-references are unresolved: Appendix N contains \"Appendix ??\" and Section K.2 contains \"Figure ??\"; these should be fixed in the final version.","section":"Appendix N; Section K.2"},{"comment":"The column header \"FPR/flag\" is ambiguous because the 2013–2015 entries are confirmed proxy false-positive rates while the 2023–2025 entries are flag rates on originals with unobserved AI use; the table should carry a more explicit header, as the main text correctly distinguishes these quantities.","section":"Table 2"},{"comment":"The main text uses \"proxy AI-assisted false-positive risk\" while Appendix B recommends the term \"assisted-writing flag rate\" for the refine (abstract only) condition; these terms should be harmonized in the main text and abstract to avoid the category error the authors themselves warn against.","section":"Section 2.2; Appendix B"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the punchline: this is a solid, policy-relevant measurement paper with genuinely careful statistics, but the abstract oversells the findings. The detectors are not completely unable to distinguish light editing from full drafts — their ROC AUCs are 0.91–0.98 — the real result is that no threshold simultaneously keeps original-flag rates, refine-capture, and false-negative rates acceptable. And the 'light refine' prompt is a substantial abstract rewrite, not grammar-level polish, so the headline 64–80% flag rate is a plausible upper bound for compliant editing, not a direct estimate of sanction risk.\n\nWhat is new and good: the four-domain corpus, the paper-clustered confidence intervals, the explicit distinction between Reading A (proxy-AI labels) and Reading B (human-source assistance), and the honest treatment of 2023–25 originals as 'flag rates' rather than confirmed false positives. The paired pre/post humanization measurement, with feature-level analysis of why evasion works, is a real extension of prior work. Code and data are promised; the appendix is unusually candid about cost and proxy limitations.\n\nSoft spots: the abstract's 'cannot distinguish' conflicts with the paper's own ROC numbers; it should say 'no single threshold gives acceptable error rates on both axes.' The refine (abstract only) condition is, by the paper's own admission (Appendix B/D), 'light generative rewriting,' so the 64–80% range (driven by Pangram; GPTZero is 37.6–48.5%) overstates what a student doing grammar fixes would face. Also, the policy asymmetry relies on the proxy; if compliant editing were much closer to the original, the flag-rate gap between honest editing and humanized evasion could narrow. That doesn't sink the conclusion — detector scores as standalone evidence is indefensible regardless — but the headline should be calibrated.\n\nWho is this for: anyone advising institutions on AI-detection policy, and anyone building detector evaluation benchmarks. It deserves a serious referee. I'd send it out; the analysis is careful and the limitations are stated better than most. Recommend a revision that tempers the abstract and clarifies the proxy's scope.","headline":"Careful, policy-relevant measurement of detector failure modes, but the abstract oversells both the inability to distinguish editing from generation and how well the 'light edit' proxy represents compliant assistance.","tokens_in":20116,"tokens_out":3105,"would_cite":true,"duration_ms":26562,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Commercial AI detectors cannot distinguish light AI-assisted editing from full LLM drafts: at the 0.50 threshold they flag honest polish at 64-80 percent while missing over 96 percent of humanized rewrites, so scores should not be…","keywords":["AI detection","academic integrity","AI-assisted writing","humanization","large language models","false positive rates","false negative rates","text classification"],"falsifier":"Collect student essays with verified authorship histories and declared AI use, score them with the same commercial detectors at 0.50, and check whether light human edits are flagged at 64-80 percent while humanized AI drafts are missed more than 96 percent of the time; the paper's catch-22 claim collapses if any threshold keeps original flag rates low, light-edit capture low, and humanized-draft misses low. The paper itself shows no threshold in {0.4, 0.5, 0.6} does this, so an adjudicated classroom dataset is the direct test.","tokens_in":19144,"feed_emoji":"⚖️","tokens_out":9594,"duration_ms":69354,"temperature":0.7,"pith_summary":"The paper tries to establish that commercial AI detectors, evaluated at the standard 0.50 threshold under proxy labels, cannot separate light AI-assisted editing from full LLM drafting, so their scores are unsafe as standalone evidence of misconduct. Using 642 published English abstracts from four disciplines and two time windows, it shows that a light \"refine abstract only\" rewrite, the paper's proxy for a human author who keeps ownership of the claims and uses an LLM for polish, is flagged 64 to 80 percent of the time, while unmodified 2023-2025 abstracts are flagged 9 to 15 percent. After passing AI-labeled rewrites through a commercial humanizer, fewer than 4 percent are still flagged, a false-negative rate above 96 percent. The consequence is an integrity catch-22: honest AI-assisted editing carries higher sanction risk than humanizer-assisted evasion, and a sympathetic reader should care because institutions currently deploy these scores as default misconduct evidence.","feed_headline":"AI detectors punish honest editing, miss humanized drafts","feed_subtitle":"Light AI polish is flagged 64-80 percent of the time; humanized AI drafts evade detection 96 percent of the time.","key_machinery":"The load-bearing object is the controlled proxy-label corpus: 642 English abstracts from chemistry, computer science, political science, and theology, split into a pre-LLM window (2013-2015) and a recent window (2023-2025), with each abstract contributing an original plus three Gemini 3 Flash rewrites of escalating intensity, namely \"refine (abstract only)\", \"refine (abstract + article)\", and \"new (article only)\". The \"refine (abstract only)\" condition is the policy-critical instrument because it models a human author who retains the claims and uses an LLM merely to polish, and its 64-80 percent flag rate is the paper's measure of assisted-writing exposure. Error rates are computed with permutation tests clustered at the paper level, and the linguistic analysis uses Spearman correlations between detector scores and surface features such as long-token ratio and Academic Word List density.","core_discovery":"On its own terms, the paper's discovery is a policy asymmetry measured under controlled proxy labels: at the standard 0.50 threshold, Pangram and GPTZero flag light AI-assisted edits of human abstracts at 64-80 percent while more than 96 percent of AI-generated rewrites escape detection after humanization. The flag rates on unmodified recent abstracts (15.0 percent Pangram, 8.9 percent GPTZero) are not confirmed false positives because real-world AI assistance is unobserved, and the paper is careful to label them as flag rates. Elevated scores track surface linguistic register, such as long-token and Academic Word List density, so non-STEM disciplines are flagged far more than STEM (p<0.001). The paper concludes that detector scores should not serve as standalone misconduct evidence and that flags must be corroborated by process evidence such as drafting history.","pith_inferences":["Editorial inference: if detector scores track long-token and Academic Word List density rather than authorship intent, then writers who adopt formal academic register, such as multilingual scholars, first-generation students, and novice writers, may be disproportionately flagged; this could be tested by comparing matched pairs of native and non-native writers with identical, verified AI-use histor","Editorial inference: the near-total post-humanization miss rate suggests an arms race in which threshold-based detectors degrade on both error types simultaneously as LLM and humanizer output distributions drift; a testable extension is longitudinal drift measurement on a fixed human-authored corpus.","Editorial inference: the paper's supplementary LLM-assisted baseline still flags most humanized rewrites, hinting that a judge reading for content and coherence rather than token statistics may be more robust than current commercial detectors; the paper does not make this claim."],"forward_implications":["A detector score at or above 0.50 cannot by itself distinguish \"student used AI to polish their own writing\" from \"student used AI to write the entire draft\", so using it as standalone misconduct evidence will punish acceptable assistance.","Non-STEM students face substantially higher flag risk than STEM students at the same threshold, so uniform detector policies can create discipline-level disparities.","Because humanizer evasion drives detection below 4 percent, any enforcement regime that relies on detectors alone fails exactly at the most severe violation, namely a fully synthetic draft.","No threshold setting keeps original flag rates, light-edit capture, and AI-pool miss rates acceptable at once, so tuning the threshold does not fix the policy failure.","Institutions that keep detectors should pair any flag with process evidence such as drafting history and reserve sanctions for cases where human judgment supports lack of effort."],"supporting_citations":[{"why":"Supplies the scholarly index from which all 642 abstracts were sampled, the corpus underlying every reported rate.","marker":"[Priem et al. 2022]"},{"why":"Technical report for the Pangram classifier, one of the two commercial detectors whose flag rates drive the headline numbers.","marker":"[Emi and Spero 2024]"},{"why":"The GPTZero benchmarking report the paper uses for the detector's claimed accuracy and as the source of its second detector's scores.","marker":"[Adam 2026]"},{"why":"Vendor claim of 99.98 percent accuracy and four-way assistance taxonomy that the paper contrasts with its measured 15.0 percent flag rate on originals.","marker":"[Spero 2025]"},{"why":"The humanizer's refund guarantee and pricing page, the tool whose post-processing produces the >96 percent false-negative rates.","marker":"[Undetectable AI 2026b]"},{"why":"Prior evidence that polishing abstracts is misclassified as full LLM rewriting, which the paper extends into a policy asymmetry.","marker":"[Geng and Poibeau 2025]"},{"why":"Independent study documenting detector false positives and false negatives, the prior unreliability claim the paper's measurement builds on.","marker":"[Dik and Erdem 2025]"},{"why":"Defines the Academic Word List used to compute the token-density features that correlate with detector scores.","marker":"[Coxhead 2000]"}],"fun_headline_variants":["AI detectors flag honest edits, miss humanized AI","Honest AI editing flagged up to 80%, humanized evasion 96%","Detectors punish honest polish, reward humanizer evasion","AI detection fails: honest edits penalized, humanized undetected"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything rests on the proxy-label scheme in which an LLM's light rewrite of a human abstract stands in for a human author's legitimate AI-assisted editing of their own work; if real classroom edits differ from the Gemini output used here, or if the 2023-2025 originals already contain unobserved AI assistance, the headline flag rates and evasion numbers would not transfer to actual integrity decisions.","fun_headline_variants_meta":{"raw":{"variants":["AI detectors flag honest edits, miss humanized AI","Honest AI editing flagged up to 80%, humanized evasion 96%","Detectors punish honest polish, reward humanizer evasion","AI detection fails: honest edits penalized, humanized undetected"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000204,"raw_usage":{"total_tokens":1378,"prompt_tokens":919,"completion_tokens":459,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":535,"completion_tokens_details":{"reasoning_tokens":386}},"tokens_in":535,"tokens_out":459,"duration_ms":4188,"temperature":1.0,"reasoning_tokens":386,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:34:03.267978+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect student essays with verified authorship histories and declared AI use, score them with the same commercial detectors at 0.50, and check whether light human edits are flagged at 64-80 percent while humanized AI drafts are missed more than 96 percent of the time; the paper's catch-22 claim collapses if any threshold keeps original flag rates low, light-edit capture low, and humanized-draft misses low. The paper itself shows no threshold in {0.4, 0.5, 0.6} does this, so an adjudicated classroom dataset is the direct test.","supporting_citations":[],"review_version":1}