REVIEW 2 major objections 2 minor 2 cited by
Machine-generated text detectors succeed only on narrow, specific notions of AI text rather than as general detectors.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-28 05:53 UTC pith:NMEJSMSF
load-bearing objection The paper introduces AITDNA with full edit histories and shows detectors work only on narrow notions of AI text, not as general ones. the 2 major comments →
'Your AI Text is not Mine': Redefining and Evaluating AI-generated Text Detection under Realistic Assumptions
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Systematic definition of multiple notions of AI-generated text combined with evaluation on AITDNA, a new benchmark containing full genesis annotations, demonstrates that current detectors are effective only for the specific notion they were developed under and do not function as broad detectors across realistic human-AI text scenarios.
What carries the argument
The AITDNA benchmark, which supplies texts together with their complete edit and AI-interaction history to enable evaluation against explicitly defined notions of AI-generated text.
Load-bearing premise
The notions chosen for the benchmark and the collected co-constructed texts accurately represent the characteristics of real-world harmful AI text use.
What would settle it
Finding a detector that maintains high accuracy across every defined notion when tested on the AITDNA data, or new real-world samples where a single detector works well without matching a narrow notion.
If this is right
- Detection claims must state the precise notion of AI text under which they were evaluated.
- Broad societal safeguards require either new detectors or explicit matching between notion and use case.
- Benchmark results can guide selection of detectors only when the deployment context aligns with one of the tested notions.
- Further work on detection should begin from the set of defined notions rather than an undefined general category.
Where Pith is reading between the lines
- Regulation that invokes detection will need to name the exact notion of AI text it intends to target.
- The detailed genesis annotations could support training of detectors that explicitly condition on the notion in use.
- Testing the same notions on non-English or domain-specific text would show whether the observed specificity pattern generalizes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that AI-generated text detection research lacks consensus on what constitutes harmful AI text use, with existing datasets relying on ad-hoc or implicit definitions loosely tied to real-world applications. It introduces a systematic taxonomy of notions of AI-generated text along with their characteristics, constructs the AITDNA benchmark of human-machine co-constructed texts annotated with full edit and AI-interaction histories, and evaluates multiple detectors across these notions. The central empirical finding is that detectors perform well only for specific notions but fail as broad, general-purpose detectors. Code and data are released publicly.
Significance. If the AITDNA benchmark and taxonomy hold, the result demonstrates that current detectors are brittle outside narrow settings and cannot serve as reliable broad safeguards, which has direct implications for deployment in education, journalism, and content moderation. The explicit grounding in co-authored edit histories and the public release of the benchmark and code are strengths that enable future work to test the same claims under controlled conditions.
major comments (2)
- [§4] §4 (Benchmark Construction): The claim that AITDNA captures 'realistic assumptions' rests on the genesis annotations, but the manuscript does not report inter-annotator agreement or validation against external corpora of harmful AI text use; without this, it is unclear whether the performance gaps generalize beyond the collected sample.
- [§5.2] §5.2 (Detector Evaluation): The reported F1 drops across notions are load-bearing for the 'not broad detectors' conclusion, yet no statistical test (e.g., McNemar or bootstrap CI) is shown for the differences between notions; the gaps could be within sampling variability given the per-notion sample sizes.
minor comments (2)
- [Table 2] Table 2: The taxonomy table lists overlapping characteristics (e.g., 'human-edited' and 'AI-initiated') without a decision procedure for assigning a text to a single notion when multiple apply.
- [§3] §3: The related-work section omits recent benchmarks that also use edit histories (e.g., those from 2023–2024 on collaborative writing), which would strengthen the novelty claim.
Simulated Author's Rebuttal
We thank the referee for their constructive feedback. We address the major comments point-by-point below, indicating planned revisions to the manuscript.
read point-by-point responses
-
Referee: [§4] §4 (Benchmark Construction): The claim that AITDNA captures 'realistic assumptions' rests on the genesis annotations, but the manuscript does not report inter-annotator agreement or validation against external corpora of harmful AI text use; without this, it is unclear whether the performance gaps generalize beyond the collected sample.
Authors: We agree that reporting inter-annotator agreement is important for validating the genesis annotations. We will compute and report IAA metrics (e.g., Cohen's kappa) for the annotations in the revised manuscript. For validation against external corpora, we note that no suitable public corpora with matching detailed annotations currently exist. We will expand the discussion of limitations and compare to existing datasets where feasible. This constitutes a partial revision. revision: partial
-
Referee: [§5.2] §5.2 (Detector Evaluation): The reported F1 drops across notions are load-bearing for the 'not broad detectors' conclusion, yet no statistical test (e.g., McNemar or bootstrap CI) is shown for the differences between notions; the gaps could be within sampling variability given the per-notion sample sizes.
Authors: We acknowledge the need for statistical validation of the observed F1 differences. In the revised version, we will include bootstrap confidence intervals and/or McNemar's test to assess whether the performance gaps between notions are statistically significant, addressing concerns about sampling variability. revision: yes
Circularity Check
No significant circularity identified
full rationale
The paper introduces new definitions of AI-generated text notions, collects a fresh benchmark (AITDNA) with detailed genesis annotations from co-constructed texts, and performs direct empirical benchmarking of detectors on that data. No equations, derivations, fitted parameters presented as predictions, or self-citation chains appear in the abstract or described methodology. The central claim rests on observable performance gaps across notions in the new dataset, which is independent of prior fitted results or self-referential definitions. This is a standard empirical contribution with no load-bearing circular steps.
Axiom & Free-Parameter Ledger
read the original abstract
Although it is generally agreed that AI-generated text poses a broad societal risk, there is no common understanding in the AI-generated text detection literature on what constitutes harmful use. Rather, existing datasets and approaches often define their own criteria and make their own assumptions, sometimes implicitly, and often only loosely related to real-world needs and applications. To address this gap, we here systematically define various notions of AI-generated text and their characteristics. To study these, we collect AITDNA - a new benchmark of human-machine co-constructed texts that is annotated with detailed genesis information, such as the entire edit and AI-interaction history. We benchmark various machine-generated text detectors and find that they often only perform well for specific notions but not as broad detectors. We release code and data publicly.
Forward citations
Cited by 2 Pith papers
-
ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation
A matched four-regime benchmark shows AI-text detectors catch direct LLM output but lose most of their recall on human text rewritten by an LLM.
-
Pangram 4 Technical Report
Pangram 4 is a commercial MoE-based detector claiming 0.9916 AUROC, 0.0041% FPR, 0.3396% FNR, plus tokenwise human/AI-assisted/AI-generated labels and humanizer detection.
Reference graph
Works this paper leans on
-
[1]
Position: On the possibilities of AI- generated text detection. InProceedings of the 41st International Conference on Machine Learning, volume 235 ofProceedings of Ma- chine Learning Research, pages 6093–6115. PMLR. DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Dam...
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[2]
RAID: A shared benchmark for robust evaluation of machine-generated text detectors. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12463–12492, Bangkok, Thailand. Association for Computa- tional Linguistics. Liam Dugan, Daphne Ippolito, Arun Kirubarajan, Sherry Shi, and Chris Callis...
-
[3]
InProceedings of the 2025 Conference on Empirical Methods in Nat- ural Language Processing, pages 5287–5302, Suzhou, China
SenDetEX: Sentence-level AI-generated text detection for human-AI hybrid content via style and context fusion. InProceedings of the 2025 Conference on Empirical Methods in Nat- ural Language Processing, pages 5287–5302, Suzhou, China. Association for Computational Linguistics. J. Peter Kincaid, Robert P. Fishburne, Richard L. Rogers, and Brad S. Chissom. ...
2025
-
[4]
Ilia Kuznetsov, Jan Buchmann, Max Eichler, and Iryna Gurevych
Machine text detectors are mem- bership inference attacks.arXiv preprint arXiv:2510.19492. Ilia Kuznetsov, Jan Buchmann, Max Eichler, and Iryna Gurevych. 2022. Revise and resubmit: An intertextual model of text-based collabora- tion in peer review.Computational Linguistics, 48(4):949–986. Mina Lee, Percy Liang, and Qian Yang. 2022. Coauthor: Designing a h...
-
[5]
In International conference on machine learning, pages 24950–24962
Detectgpt: Zero-shot machine-generated text detection using probability curvature. In International conference on machine learning, pages 24950–24962. PMLR. Sheshera Mysore, Debarati Das, Hancheng Cao, and Bahareh Sarrafzadeh. 2025. Prototypical human-AI collaboration behaviors from LLM- assisted writing in the wild. InProceedings of the 2025 Conference o...
2025
-
[6]
Qwen2.5 technical report.arXiv preprint arXiv:2412.15115v2. Rounak Saha, Gurusha Juneja, Dayita Chaudhuri, Naveeja Sajeevan, Nihar B Shah, and Danish Pruthi. 2026. Policies permitting llm use for polishing peer reviews are currently not en- forceable.arXiv preprint arXiv:2603.20450. Shoumik Saha and Soheil Feizi. 2025. Almost AI, almost human: The challen...
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[7]
Access document
Detectrl: Benchmarking llm-generated text detection in real-world scenarios.Ad- vances in Neural Information Processing Sys- tems, 37:100369–100401. Peipeng Yu, Jiahan Chen, Xuan Feng, and Zhihua Xia. 2025. Cheat: A large-scale dataset for de- tecting chatgpt-written abstracts.IEEE Trans- actions on Big Data. Zijie Zeng, Shiqi Liu, Lele Sha, Zhuang Li, Ka...
2025
-
[8]
Transformers architecture, used in Large Language Models like GPT-4, is a technology that revolutionized machine text generation and understanding
-
[9]
Social engineering is a psychological technique to “hack” people with the purpose to get access to restricted or personal information
-
[10]
Japanese trains
There are vast differences in railway systems throughout the world: in terms of reliability, coverage, speed. Japanese trains . . . Your text should give a clear explanation on the concept/person/time period/technology etc, so that people without deep knowledge can get a good understanding from your text. Approximate expected length is 350-450 words, but ...
-
[11]
Tell a story from the perspective of an inhabitant of the new world - describe your typical day and how the changes affect your life, work, relationships, etc
-
[12]
Write a newspaper-style text that interviews several people from that world
-
[13]
Approximate expected length is 350-450 words, but it may vary
Provide an analytical article that analyses the societal, economical, or cultural impact that such changes could have on our society. Approximate expected length is 350-450 words, but it may vary. Task 2.3: Argumentative writing In this part, you have 10 minutes to write an argumentative essay. You HA VE TO interact with the LLM during your writing proces...
-
[14]
The attention mechanism and its role in today’s LLM architectures
-
[15]
Parameter efficient fine-tuning as the key driver for democratization of LLMs
-
[16]
Inference-time scaling and its tradeoff between efficiency and performance Describe the core idea of this concept, what problem or limitation it addresses, and why it stood out to you. How does it advance the understanding within the field? How does the approach differ from existing methods or theories? How might this idea influence your own research inte...
-
[17]
Researchers will be unburdened from the task
If we can successfully automate scholarly peer review, it is the first step towards self-improving artificial intelligence. Researchers will be unburdened from the task. Research progress will be steady. However, selection bias towards certain research will be inevitable. Incentives for gaming the system will be high
-
[18]
For any problem, we could identify the exact underlying causes from the model activations
If we could link each LLM capability to a mechanistic circuit, we would have cracked the LLM black box. For any problem, we could identify the exact underlying causes from the model activations. Yet, this invites new attacks on LLMs in production. Your text has to have a clear structure. Try to think of multiple consequences and perspectives on them. Argu...
-
[19]
Scaling and engineering of existing LLM training technology will lead us towards AGI
-
[20]
Differentially private language technology with maximum utility can only be achieved by means of gradient-based privatization strategies
-
[21]
State the topic in the beginning and then write your essay
Fact checking can and should be fully automatized. State the topic in the beginning and then write your essay. Write a coherent argumentative statement reflecting on the scientific evidence that speaks in favor or against the claim. Target your argumentation towards a reader that is generally knowledgeable in science but not necessarily in your specific f...
-
[22]
What if no one ever had to work again because robots took care of every- thing? (source)
-
[23]
What if our memories could be stolen and sold on the black market? (source)
-
[24]
What if we could restore the diminishing populations of endangered species by cloning them? (source)
-
[25]
What if you could test the outcomes of different decisions in virtual reality with 100% accuracy? (source)
-
[26]
What if there were a world where people spend all their personal time escaping to idyllic VR settings instead of confronting the challenges of real life? (source)
-
[27]
What if the usage of cars for personal needs were banned?
-
[28]
What if people had microchips in their eyes that allowed them to record everything they see?
-
[29]
What if vibe-coding became so popular that every IT company would switch to it?
-
[30]
What if the theory that COVID-19 vaccines contain microchips were true?
-
[31]
Topics without sources were created by one co-author
What if it were forbidden to sell or buy products containing meat? Figure 11: Topics for the creative task. Topics without sources were created by one co-author. C.3 LLM Prompt Configurations Figure 13 and Figure 14 provide the system prompts for continuation and revision queries, re- spectively. C.4 Detector Configurations Forgoogle/gemma-3-1b-itandEleut...
-
[32]
Is it cheating when students use artificial intelligence to help them with their schoolwork? In your opinion, how, if at all, should students be allowed to use AI in school? What do you see as benefits and drawbacks of using AI for doing homework? (source)
-
[33]
Should 16-Year-Olds Be Allowed to V ote? In your opinion, is the current minimum legal voting age of 18 fair and appropriate? What influence would lowering the threshold to 16 years have on the society? (source)
-
[34]
Is That a Problem? Some say video games are a chief reason boys and young men are struggling
Boys Are Spending More Time Gaming. Is That a Problem? Some say video games are a chief reason boys and young men are struggling. Others say games serve an important role in teens’ lives. What do you think? What can gaming bring to a teen’s life? What other activities, if any, does it take away from? (source)
-
[35]
Should Schools Ban Student Phones? More and more countries are cracking down on students’ use of cellphones. Are these restrictions fair? Can they work? Do you think that phones interfere with student’s academic learning, the quality of their social interactions or overall engagement in school? (source)
-
[36]
Is It Ethical for Teachers to Use AI to Grade Papers? Many schools do not allow students to use artificial intelligence to complete their assignments. Should teachers be held to the same standard? Do you think teachers should be able to use the technology at all? If so, in which instances would it be acceptable, and what guidelines should teachers follow ...
-
[37]
Should Social Media Companies Be Responsible for Fact-Checking Their Sites? A few months ago, Meta, the company that owns Facebook and Instagram, has ended its longstanding fact-checking pro- gram. Is that a good idea? What could this mean for users? Do you believe social media companies should be responsible for fact-checking lies, misinformation, disinf...
-
[38]
Should Single-Use Vapes Be Banned Everywhere? In an effort to protect young people’s health, Eng- land plans to ban disposable vapes next year. Do you think this measure will curb vaping among teens? To what extent do you think the government should have a role in trying to reduce smoking and vaping? Do you think your government should be doing more to di...
-
[39]
How Important Is a Free Press to Our Democracy? What role do journalists play in our society - whether they work for big national newspapers, niche podcasts or YouTube channels? What would happen if they couldn’t freely investigate and report the news? For example, what if journalists were barred from informing the public about the actions of powerful peo...
-
[40]
Does Everyone Have a Responsibility to V ote? Is it OK not to vote, or is voting a civic duty? What to you are the most compelling reasons for showing up at the polls? What about those for sitting out? Why do you think so many people don’t vote? What do you think would encourage them to participate more? (source)
-
[41]
Workers across industries are concerned about AI coming for their work. Are you worried about AI taking human jobs? Why or why not? Which types of work do you think are most vulnerable to automation? Do you think the fears about AI are overblown? (source)
-
[42]
Should All Children Under 16 Be Barred From Social Media? Australia recently passed a law that does just that. Should other countries do the same? What is your reaction to Australia’s new law? Do you think it is a good idea? Will it be effective? What do you think are the negative effects of social media on young people? If you don’t think a similar law w...
-
[43]
Should Grades Be Based on Excellence or Effort? Some people think that too many students today wrongly expect to be rewarded for their efforts rather than the quality of their work. Do you agree? Do you think marks are accurate reflections of students’ learning? What do you think student grades should be based on? How much, if at all, should effort and ha...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.