REVIEW 2 major objections 9 references
Humans Cannot Detect AI-Generated Media But Communities May -- For Now: Collaborative AI Detection in r/RealOrAI on Reddit
T0 review · 2 major / 0 minor · reviewed 2026-06-30 · grok-4.3
Pith's one-line read A Reddit community reaches 72% accuracy detecting AI-generated images by aggregating user comments, though false positives rise over time.
desk verdict The paper's scale on community AI detection in one subreddit is new, but accuracy and bias claims rest on unvalidated submitter labels. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The community aggregate prediction, formed by combining multiple user comments classified across six cues, that selectively amplifies rare diagnostic evidence such as provenance verification.
What would settle it
Independent forensic or expert verification of a random sample of the labeled posts to measure whether the reported 72% accuracy matches objective ground truth.
Extended reading notes
Core claim
Community detection accuracy reaches 72% on [GUESS] posts with a systematic false-positive bias that intensifies over the year as the community's AI-suspicion grows. Using a six-LLM ensemble, perceptual features dominate reasoning at 70% while provenance verification is rarest at 4% at the individual level but is amplified 4.3x in community summaries, revealing aggregation as a reliability filter that selectively surfaces diagnostic evidence.
Load-bearing premise
Submitter-provided verified labels on self-challenging posts constitute accurate naturalistic ground truth without external validation of label correctness or participation biases.
Editorial extensions
If this is right
- Individual detection remains limited by heavy reliance on perceptual features that may fail as generation quality improves.
- Community aggregation reliably boosts the visibility of provenance checks that are otherwise rare.
- False-positive rates increase as collective suspicion of AI content grows over time.
- Heuristic-based detection alone is insufficient in environments with contested media authenticity.
Reading between the lines
- Platforms could test similar comment-aggregation systems to assist human moderators at scale.
- Detection performance may degrade further if AI generators specifically target the perceptual cues users currently favor.
- The observed bias toward AI labels could be studied in other online communities facing similar authenticity challenges.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper analyzes one year of activity in the Reddit community r/RealOrAI, where users collaboratively judge whether visual media is real or AI-generated. It claims that submitter-provided verified labels on self-challenging [GUESS] posts constitute naturalistic ground truth at scale, that the community reaches 72% detection accuracy on these posts with an intensifying false-positive bias, and that an LLM ensemble applied to 10k comments shows perceptual features dominate individual reasoning (70%) while provenance verification is rare (4%) but amplified 4.3x in community aggregates.
Significance. If the submitter labels are shown to be reliable, the work would provide a rare large-scale, naturalistic dataset on collaborative human AI detection, document cue distributions and aggregation effects, and illustrate how online communities adapt to contested media. The observational design and cue taxonomy could inform platform moderation and detection research, but the absence of external validation for the core labels substantially weakens these potential contributions.
major comments (2)
- [Abstract] Abstract: the headline accuracy figure of 72% on [GUESS] posts and the reported false-positive bias trend are computed by treating submitter-provided verified labels as ground truth, yet the manuscript supplies no independent verification step, inter-rater reliability metric, audit against external sources (reverse-image search, model provenance, or expert adjudication), or analysis of selection effects among users who post [GUESS] content.
- [Abstract] Abstract: every downstream claim about cue amplification, aggregation as a reliability filter, and the intensification of AI-suspicion bias rests on the same unvalidated labels; without evidence that these labels are accurate or free of correlated mislabeling, the 72% accuracy, bias trend, and 4.3x provenance amplification cannot be interpreted as community performance.
Simulated Author's Rebuttal
We thank the referee for the detailed and constructive report. We address the two major comments on label validation point by point below, agreeing that further clarification of limitations is needed while maintaining that the naturalistic scale of the data remains a core contribution.
read point-by-point responses
-
Referee: [Abstract] Abstract: the headline accuracy figure of 72% on [GUESS] posts and the reported false-positive bias trend are computed by treating submitter-provided verified labels as ground truth, yet the manuscript supplies no independent verification step, inter-rater reliability metric, audit against external sources (reverse-image search, model provenance, or expert adjudication), or analysis of selection effects among users who post [GUESS] content.
Authors: We agree that the manuscript provides no independent verification, inter-rater metric, or external audit of the submitter labels. The labels are obtained via the subreddit bot directly soliciting verified self-reports from submitters on self-challenging [GUESS] posts; this is the source of the naturalistic ground truth claim. We will make a partial revision by adding an explicit limitations paragraph in the methods and discussion sections that notes the absence of external audits (e.g., reverse-image search or expert adjudication) and discusses possible selection effects among users choosing to post [GUESS] content. The abstract will be updated to state that reported accuracy is measured relative to these submitter labels. revision: partial
-
Referee: [Abstract] Abstract: every downstream claim about cue amplification, aggregation as a reliability filter, and the intensification of AI-suspicion bias rests on the same unvalidated labels; without evidence that these labels are accurate or free of correlated mislabeling, the 72% accuracy, bias trend, and 4.3x provenance amplification cannot be interpreted as community performance.
Authors: We accept that all performance metrics and the interpretation of aggregation effects are conditional on the submitter labels. The LLM-based cue classification of comments is independent of the labels, but the accuracy, bias trend, and amplification ratios are not. We will revise the abstract, results, and discussion to frame these quantities explicitly as community behavior benchmarked against submitter-provided labels and to note the possibility of correlated mislabeling. We maintain that the observational design still documents real cue distributions and aggregation patterns at scale, but we will temper causal language about "community performance" accordingly. revision: partial
- We cannot supply a full external validation, inter-rater reliability study, or audit against reverse-image search / model provenance / expert adjudication for the complete year-long [GUESS] dataset, as this would require resources and data access outside the scope of the current observational study of public Reddit content.
Circularity Check
Observational study with no derivations or self-referential reductions
full rationale
The paper is an empirical observational study of Reddit posts and comments. It reports community accuracy computed directly from submitter labels treated as ground truth and classifies comments via an LLM ensemble, with no equations, fitted parameters renamed as predictions, self-citation load-bearing uniqueness claims, or ansatzes smuggled via prior work. The central claims rest on data aggregation and cue counting rather than any derivation chain that reduces to its own inputs by construction.
Assumptions & free parameters
assumptions (1)
- domain assumption Submitter labels on [GUESS] posts and community aggregates constitute accurate naturalistic ground truth.
Cite this review
Pith. "Pith review of Humans Cannot Detect AI-Generated Media But Communities May -- For Now: Collaborative AI Detection in r/RealOrAI on Reddit." pith.science (2026). https://pith.science/paper/UHCYYZXN
@misc{pith2026260524287,
author = {Pith},
title = {Pith review of: Humans Cannot Detect AI-Generated Media But Communities May -- For Now: Collaborative AI Detection in r/RealOrAI on Reddit},
year = {2026},
howpublished = {\url{https://pith.science/paper/UHCYYZXN}},
note = {Machine review of arXiv:2605.24287}
}
read the original abstract
We study human AI-detection behaviour at scale using a year of activity from r/RealOrAI, a Reddit community where users collaboratively assess whether visual media is real or AI-generated. The community is moderated by a bot that solicits verified labels from submitters of self-challenging "[GUESS]" posts and publishes an aggregate community prediction for each post, yielding naturalistic ground truth at scale. Community detection accuracy reaches 72% on [GUESS] posts with a systematic false-positive bias that intensifies over the year as the community's AI-suspicion grows. Using a six-LLM ensemble validated against human-annotated ground truth, we classify 10k reasoning-bearing comments along six cues covering perceptual features, context, consistency, AI knowledge, subject-matter expertise and provenance (tracing the media to its source). Perceptual features (scene, visual artifacts, anatomy physics, lighting, behavior, text, audio) dominate reasoning (70%) while provenance verification is rarest (4%) at the individual level but is amplified 4.3x in community summaries, revealing aggregation as a reliability filter that selectively surfaces diagnostic evidence. These findings reveal the limits of heuristic-based detection and show how online communities collectively navigate an increasingly contested information environment.
Figures
Reference graph
Works this paper leans on
-
[1]
Coalition for Content Provenance and Authentic- ity
Community-based fact-checking reduces the spread of misleading posts on social media.arXiv preprint arXiv:2409.08781. Coalition for Content Provenance and Authentic- ity. 2024. C2PA Technical Specification. https: //c2pa.org/specifications/specifications/2.0/. Croitoru, F.-A.; Hiji, A.-I.; Hondru, V .; Ristea, N. C.; Irofti, P.; Popescu, M.; Rusu, C.; Ion...
-
[2]
Deepfake detection by human crowds, machines, and machine-informed crowds.Proceedings of the National Academy of Sciences, 119(1): e2110013119. Ha, A. Y . J.; Passananti, J.; Bhaskar, R.; Shan, S.; Southen, R.; Zheng, H.; and Zhao, B. Y . 2024. Organic or diffused: Can we distinguish human art from ai-generated images? In Proceedings of the 2024 on ACM SI...
work page 2024
-
[3]
Liu, H.; Bhatia, P.; Vincent, N.; and K Chilana, P
Springer. Liu, H.; Bhatia, P.; Vincent, N.; and K Chilana, P. 2026. Tracing Everyday AI Literacy Discussions at Scale: How Online Creative Communities Make Sense of Generative AI. InProceedings of the 2026 CHI Conference on Human Fac- tors in Computing Systems, 1–28. Lloyd, T.; Reagle, J.; and Naaman, M. 2025. ’There Has To Be a Lot That We’re Missing’: M...
-
[4]
For most authors... (a) Would answering this research question advance sci- ence without violating social contracts, such as violat- ing privacy norms, perpetuating unfair profiling, exac- erbating the socio-economic divide, or implying disre- spect to societies or cultures? Yes, the study analyses publicly posted, pseudonymous content from a single subre...
-
[5]
Additionally, if your study involves hypotheses testing... (a) Did you clearly state the assumptions underlying all theoretical results? NA, the paper presents observa- tional empirical findings rather than formal theoretical results. (b) Have you provided justifications for all theoretical re- sults? NA, no theoretical results are claimed. (c) Did you di...
-
[6]
Additionally, if you are including theoretical proofs... (a) Did you state the full set of assumptions of all theoret- ical results? NA, the paper does not include theoretical proofs. (b) Did you include complete proofs of all theoretical results? NA, the paper does not include theoretical proofs
-
[7]
Additionally, if you ran machine learning experiments... (a) Did you include the code, data, and instructions needed to reproduce the main experimental results (either in the supplemental material or as a URL)? Yes, annotated data, classification prompts, evaluation scripts will be released in a repository upon accep- tance. (b) Did you specify all the tr...
-
[8]
Additionally, if you are using existing assets (e.g., code, data, models) or curating/releasing new assets,without compromising anonymity... (a) If your work uses existing assets, did you cite the creators? Yes, the Reddit platform and the six LLMs (Llama 3.3, Gemini 2.5 Flash, GPT-5.2, GPT-5-mini, Claude Sonnet 4.6, Claude Haiku 4.5) are cited in the Met...
work page 2016
Show all 9 references
-
[9]
index": <int>,
Additionally, if you used crowdsourcing or conducted research with human subjects,without compromising anonymity... (a) Did you include the full text of instructions given to participants and screenshots? NA, the study is fully observational and does not involve crowdsourced p...
2019
Reviewed June 30, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.