{"id":"6e2aece1-c3d6-4b58-9f70-3d4814a4eab3","arxiv_id":"2501.15654","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Expert annotators who frequently use LLMs for writing detected AI-generated non-fiction articles with 99.3% accuracy by majority vote, outperforming all but one commercial detector.","lead":"Five professional writers who use ChatGPT daily labeled 300 short news and magazine articles as human or AI-written and got 299 right. The paper argues that experienced LLM users can be better than most automatic detectors, especially when they work as a group and write explanations.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Expert annotators were selected by screening on the very task used in Experiment 1, so the title's population-level claim that frequent ChatGPT users are accurate detectors is confounded by outcome-based sampling; a fresh unscreened cohort is required to test it.","rationale":"The paper's empirical core is real and honestly reported: five annotators who also happen to be screened for accuracy achieve 299/300 on a corpus spanning four generators and two evasion tactics, with code and data released. The headline, however, makes a population-level attribution to 'people who frequently use ChatGPT for writing tasks'. The design cannot support that attribution because the 'expert' label was assigned only after passing a performance screen on Experiment 1 material (§A.1). Four of the five experts were hired because they scored at least 90% on the same 60-article set used in Experiment 1, and the fifth was the original high performer; the 5-question pre-screen was also drawn from that set. Any claim about a group's average or typical ability requires sampling from the group, not selecting members by the outcome variable. This is not an attack on the authors; it is a structural feature of the recruitment design. The 'no specialized training or feedback' claim is compatible with selection: the annotators were not trained, but they were chosen for ability, and the title's causal reading is the thing that breaks.\n\nA second issue is that Experiments 2–5 reuse the same generation template (Table 8) that explicitly asks for 'quotations from relevant experts', a concise lay-audience style, and a particular word count; human articles from reputable outlets differ from LLM outputs on these dimensions for reasons that may be prompt artifacts rather than general AI signatures. The humanization attack (Table 13) also starts from this template and adds the authors' own guidebook. So the 300-article result may overstate robustness to real-world AI text, which is exactly the reader's weakest assumption.\n\nNeither issue makes the study worthless: the dataset, the detailed explanations, and the comparison to automatic detectors are valuable, and the conditional verdict is appropriate. To move from CONDITIONAL to a stronger claim, the authors would need either to reframe the title and abstract to 'screened expert annotators' or to run the unscreened-cohort replication described above.","tokens_in":30222,"tokens_out":5868,"duration_ms":57457,"concrete_test":"Run a preregistered replication with a new cohort: recruit N=20 self-identified frequent-ChatGPT users for writing tasks, with the same profession/education criteria as the current experts, but impose no accuracy screen. Have them label a fresh batch of 60 articles (30 human, 30 AI) built with the same generation pipeline, using the same interface and instructions. Then compute the majority vote of all 20 and of random 5-person subsets, and compare TPR/FPR to the paper's 99.3%/0%. If unscreened frequent users' majority vote is substantially below 99.3%, the paper's population claim fails and selection is the load-bearing factor; if it remains near-perfect, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing flaw is the expert recruitment protocol in §A.1, not the article-generation prompt. In the initial Experiment 1 pilot, only one of five annotators (Annotator 1) was highly accurate. The authors then recruited ten Upwork writers with editing/writing backgrounds who frequently use LLMs, gave them a 5-question sample drawn from the same Experiment 1 article set, and required at least 4/5 correct; the four who passed were then required to reach at least 90% accuracy on the remaining 60 Experiment 1 articles before being hired as 'experts'. Thus all five experts were chosen because they were accurate on the exact distribution used to demonstrate high accuracy. This makes the abstract's causal/populational claim—'annotators who frequently use LLMs for writing tasks excel'—unsupported: the study demonstrates that a small, screened, high-performing subset of frequent LLM users can be identified, not that frequent LLM use causes or predicts detection skill. The 'no training' point does not fix this, because outcome-based selection is not training. Experiments 2–5 show the chosen experts generalize across generators and evasion tactics within the same article genre and generation template, which is a genuine result, but it describes screened experts, not the population named in the title. A related but secondary concern, also noted by the reader, is that all AI articles are produced from title/subtitle/length/publication prompts with an instruction to include expert quotations, so the texts may share prompt-specific fingerprints rather than representative 'AI-generated text'.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a human annotation study in which five 'expert' annotators, recruited for frequent LLM use and editing/writing backgrounds, label 300 human-authored and LLM-generated non-fiction English article pairs across five experiments. The authors report that a majority vote of these five experts misclassifies only 1 of 300 articles (99.3% TPR, 0% FPR), outperforming most tested automatic detectors and matching the commercial Pangram model, including under paraphrasing and a prompt-based humanization attack. The paper also analyzes the experts' free-text explanations, extracts a taxonomy of detection clues, and tests whether LLMs prompted with a guidebook built from those explanations can mimic the experts. The authors release the annotated dataset and code.","tokens_in":30415,"tokens_out":3091,"duration_ms":30845,"significance":"If interpreted as a statement about screened, high-performing human annotators, the result is significant: it provides a concrete benchmark showing that a small ensemble of human experts can match the best commercial detector on this article corpus while also offering explanations, and it documents which textual clues survive paraphrasing and humanization. The dataset and code release are valuable assets, and the comparison across GPT-4o, Claude, and o1-pro with and without evasion tactics is more thorough than most prior human-detection studies. However, the title and abstract claim a population-level conclusion about frequent LLM users, and that claim is not supported by the recruitment protocol, which selected annotators based on performance on the very task and article distribution later reported as the headline result.","major_comments":[{"comment":"The recruitment of expert annotators in §A.1 creates an outcome-based selection artifact for Experiment 1. The four additional experts were required to score at least 4/5 on a 5-question sample drawn from the Experiment 1 article set and then at least 90% on the remaining 60 Experiment 1 articles; the fifth expert (Annotator 1) was the original pilot annotator who already performed almost perfectly on that same batch. Consequently, the perfect majority-vote result in §2.1 is partly guaranteed by construction and cannot support the abstract's claim that 'annotators who frequently use LLMs for writing tasks excel at detecting AI-generated text.' The paper should either reframe the central claim as being about screened experts who pass a performance qualification, or add an unscreened cohort of frequent LLM users evaluated on a held-out article set. Experiments 2–5 do provide independent evidence for the robustness of the five selected experts on new article sets, but they do not rehabilitate the population-level title claim.","section":"§A.1, §2.1"},{"comment":"The generalization claim that experts are 'robust detectors of AI-generated text' is anchored to a narrow generation template. All AI articles, including humanized ones, are produced from a prompt that supplies the title, subtitle, publication, section, and desired length, and instructs the model to 'Include quotations from relevant experts and make sure the article is concise and easily understandable to a lay audience.' This common prompt likely induces systematic structural and lexical fingerprints (e.g., expert quotes placed at the end of paragraphs, uniform exposition style) that the experts learn to exploit. The paper therefore demonstrates robustness across three model families and two evasion tactics within one genre and one generation template, but it does not establish robustness to the broader distribution of real-world AI text, which may be written with different prompts, lengths, domains, or post-editing. The limitations section acknowledges domain restriction and possible AI edits in the human articles, but it does not address the prompt-template confound; a paragraph acknowledging this and softening the title-level generalization would be needed.","section":"§2, Table 8, §B.2"},{"comment":"The headline robustness result depends heavily on the majority-vote aggregation, and the paper should be more explicit that this is an ensemble property rather than a property of typical individual experts. Annotator 3, for example, achieves a TPR of only 16.7% on o1-pro and 0% on humanized o1-pro articles, while Annotator 2's FPR reaches 30% on Claude articles. The paper does acknowledge individual variation in §3 and §D.3, but the abstract's phrasing 'five such expert annotators' could mislead readers into thinking each expert is highly robust. The practical recommendation to hire 'expert human annotators' should state clearly that a majority-vote panel of at least five screened experts is the unit that delivers the reported 99.3% TPR, and that individual performance varies substantially across generators.","section":"§2.4, Table 2"}],"minor_comments":[{"comment":"There are numeric inconsistencies between the text and Table 1: the text reports an average FPR of 52.5% for nonexperts, but Table 1 lists 51.7%; the text reports an average FPR of 3.3% for experts, but Table 1 lists 4.0%. Please reconcile these values.","section":"§2.1, Table 1"},{"comment":"The paper states in the introduction that 'we collect 1790 annotations on 300 unique articles' but later reports '1740 annotations from 9 annotators.' With 5 experts × 300 articles and 4 nonexperts × 60 articles, the total is 1740; the 1790 figure appears to be a typo and should be corrected.","section":"Introduction, §2 (Annotator details)"},{"comment":"The 'OVERALL' columns in Table 2 should specify whether the TPR/FPR values are macro-averaged across the five experiments or computed over the pooled 150 AI and 150 human articles. The current presentation makes this ambiguous, which matters because experiments differ in difficulty (e.g., o1-pro humanized has much lower automatic-detector TPR).","section":"Table 2"},{"comment":"The sentence 'This diversity explains why an ensemble of expert annotators performs so well and it also suggests room for training so that all experts are at least aware of most distinguishing features. implies additional training can improve individual experts.' contains a grammatical break ('features. implies') and should be rewritten.","section":"§3, §D.3"},{"comment":"The captions for the paraphrase prompts refer to 'Experiment 2 Paraphras-ing' but paraphrasing is used in Experiment 3; please correct the cross-reference.","section":"§B.2, Table 9/Table 10"}],"recommendation":"major_revision","confidential_remarks":"The central measurement is transparent and the released data/code are valuable, but the title and abstract overstate the population-level claim. The selection protocol in §A.1 is the main load-bearing issue; a revision that reframes the contribution around screened expert panels and adds an unscreened validation (or explicitly removes the population claim) would bring the paper in line with its evidence. I would not reject the paper: Experiments 2–5 provide genuine evidence that a small panel of screened experts generalizes across generators and evasion tactics within this genre, and the qualitative taxonomy is useful."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The real result here is that a small, carefully screened group of five annotators can detect AI-generated articles almost perfectly across GPT-4o, Claude, o1-pro, paraphrasing, and even a prompt-based humanization attack: majority vote misses 1 of 300 articles. The paper ships a new 300-article paired corpus, detailed annotations with explanations, and a fair comparison against commercial and open-source detectors. That is a genuinely useful empirical contribution, and the qualitative analysis of what the experts attend to (vocabulary, sentence structure, originality, quotes) is careful and interesting.\n\nThe soft spot is exactly where the stress-test note lands. In Appendix A.1, four of the five experts were hired after passing a screening test built from the very articles used in Experiment 1. They had to score at least 4/5 on a 5-question sample and then at least 90% on the remaining 60 Experiment 1 articles. So Experiment 1's perfect majority vote is partly an artifact of selecting people who were already known to perform well on that exact distribution. That does not kill the paper, because Experiments 2–5 use fresh article sets and the majority vote stays near perfect. But it means the title's claim — that people who frequently use ChatGPT for writing are accurate detectors — is not supported. What the data show is that you can identify a high-performing subset of such users, and that subset generalizes across models and evasion tactics within this article genre. That is still a good result, but it is a different claim.\n\nA secondary concern is that all AI articles were generated from a prompt with title, subtitle, length, publication, and an instruction to include expert quotations, so the texts likely share structural fingerprints. The humanization attack was built explicitly from these experts' own explanations, which is both a strength (the robustness is genuinely adversarial to them) and a limitation (it is not an independent attack). The study's restriction to short American English non-fiction is acknowledged in the limitations section, which also honestly flags possible AI edits in human articles.\n\nOverall: the measurements look internally consistent, the appendices are transparent, and the dataset is a real resource. The central empirical phenomenon — screened experts are robust detectors — holds up, but the framing overgeneralizes. This deserves a serious referee and likely acceptance after the claims are reframed and the selection effect is reported prominently. I would cite it for the dataset and the robustness result, and I would bring it to reading group.","headline":"A solid, well-documented study of a small screened expert panel that detects AI text near-perfectly, but the title's population-level claim about frequent ChatGPT users is not supported by the selection design; worth serious refereeing after reframing.","tokens_in":31034,"tokens_out":1434,"would_cite":true,"duration_ms":15414,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Frequent LLM users detect AI-generated articles with near-perfect accuracy.","keywords":["AI-generated text detection","human evaluation","LLM expertise","majority vote","paraphrasing attack","humanization","explainability","GPT-4o"],"falsifier":"Test the same five experts on a corpus of AI articles generated in the wild (varied prompts, user requests, other genres, or non-English text) and with humanization methods not built from the experts' own clues; if their majority vote accuracy drops substantially below the reported 99.3%, the claim that expert writers are robust detectors of AI-generated text would be falsified.","tokens_in":29940,"feed_emoji":"🕵️","tokens_out":4223,"duration_ms":36153,"temperature":0.7,"pith_summary":"The paper sets out to show that people who frequently use large language models for writing tasks are highly accurate and robust detectors of AI-generated text, even with no specialized training or feedback. It reports that the majority vote of five such \"expert\" annotators misclassified only 1 of 300 articles, matching the best commercial detector and beating most open-source detectors. The result holds against paraphrasing and a prompt-based humanization attack, suggesting that human expertise, not just automatic tools, can serve as a reliable detection mechanism in high-stakes settings where explanations matter.","feed_headline":"Five ChatGPT-savvy writers catch 299 of 300 AI articles","feed_subtitle":"A tiny expert panel outdetects most AI-text tools, even on paraphrased and humanized articles.","key_machinery":"The central machinery is the expert majority vote: five annotators recruited for frequent LLM use in writing tasks each independently label articles, highlight clue spans, rate confidence, and write explanations; their majority vote is the system that achieves 99.3% true-positive and 0% false-positive on 300 articles. The detection guide compiled from expert explanations also drives the paper's humanization attack and its prompt-based detector, making the annotators' clue taxonomy the load-bearing component that carries both the human and automated results.","core_discovery":"On the paper's own terms, the central discovery is that expertise with LLMs for writing transfers to detection: five annotators who routinely edit, copywrite, or proofread with ChatGPT identify AI-generated nonfiction articles at near-perfect accuracy using vocabulary, sentence-structure, and originality cues, and their aggregate majority vote is robust to adversarial paraphrasing and humanization. The paper further finds that these experts outperform every tested automatic detector except the commercial Pangram system, which they match, and that their free-form explanations reveal clues—overused \"AI vocabulary\", formulaic structures, overly tidy conclusions—that are accessible to humans but hard for detectors to score.","pith_inferences":["A natural extension the paper does not test: combining expert human adjudication with automatic scorers in a human-in-the-loop pipeline, where experts handle borderline or high-stakes cases while machines do the bulk screening.","Because all AI articles were generated with a single prompt template using the human article's title, subtitle, and publication, the measured expert accuracy may partly reflect prompt-induced artifacts; a test on articles produced by real users with varied prompts could yield lower accuracy.","The reported cost ($4.9K for 1,790 annotations) and throughput (8-12 articles per hour per annotator) imply that expert panels are economical only for low-volume, high-stakes detection, not for content moderation at scale; the paper acknowledges the scale limit but does not quantify it in deployment terms."],"forward_implications":["A small panel of LLM-experienced writers can serve as a practical detection workforce in settings where false accusations are costly and an explanation is required.","Prompt-based humanization built from expert clues degrades most automatic detectors but leaves expert majority-vote accuracy at 100% on the tested articles, so current humanization tools do not yet erase the signatures experts rely on.","Individual expert performance varies (one annotator's true-positive rate falls to 0% on humanized articles), so relying on a single human reader is risky; the paper's near-perfect result depends on aggregating diverse clue preferences.","The paper's prompt-based detector using an expert guidebook is competitive on easy configurations but fails on humanized text, indicating that simply giving an LLM the guide does not reproduce expert robustness."],"supporting_citations":[{"why":"Provides the Pangram commercial detector that the expert majority vote matches, serving as the strongest automatic baseline.","marker":"Emi and Spero, 2024"},{"why":"Supplies the Pangram Humanizers variant trained on humanized data, the only detector tied with the expert panel.","marker":"Masrour et al., 2025"},{"why":"Provides GPTZero, a closed-source detector whose substantial degradation on o1-pro and humanized articles contrasts with expert robustness.","marker":"Tian and Cui, 2023"},{"why":"Provides Binoculars, an open-source detector that the experts outperform across all configurations.","marker":"Hans et al., 2024"},{"why":"Provides Fast-DetectGPT, an open-source detector that struggles under paraphrasing and humanization, anchoring the expert comparison.","marker":"Bao et al., 2023"},{"why":"Provides RADAR, an adversarial detector with low detection rates in this study, part of the automatic-detector benchmark set.","marker":"Hu et al., 2023"},{"why":"Motivates the collection of explanations and reading-time controls to avoid the pitfalls of unmotivated crowd annotations.","marker":"Karpinska et al., 2021"}],"fun_headline_variants":["ChatGPT-savvy writers spot AI text with 99.7% accuracy","Five ChatGPT-heavy writers mislabel only one AI article in 300","Routine ChatGPT users detect AI text better than most detectors","ChatGPT experts outperform AI-text detectors on humanized articles"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The near-perfect expert accuracy assumes that the AI-generated articles used in the study—produced by prompting with only the title, subtitle, length, and publication name—are representative of the AI-generated text people actually encounter, and that the paired human articles contain no AI edits.","fun_headline_variants_meta":{"raw":{"variants":["ChatGPT-savvy writers spot AI text with 99.7% accuracy","Five ChatGPT-heavy writers mislabel only one AI article in 300","Routine ChatGPT users detect AI text better than most detectors","ChatGPT experts outperform AI-text detectors on humanized articles"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001151,"raw_usage":{"total_tokens":4730,"prompt_tokens":863,"completion_tokens":3867,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":479,"completion_tokens_details":{"reasoning_tokens":3793}},"tokens_in":479,"tokens_out":3867,"duration_ms":27064,"temperature":1.0,"reasoning_tokens":3793,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:03:39.753557+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Test the same five experts on a corpus of AI articles generated in the wild (varied prompts, user requests, other genres, or non-English text) and with humanization methods not built from the experts' own clues; if their majority vote accuracy drops substantially below the reported 99.3%, the claim that expert writers are robust detectors of AI-generated text would be falsified.","supporting_citations":[],"review_version":1}