Pith. sign in

REVIEW 2 cited by

The Rise of AI-Generated Content in Wikipedia

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.08044 v1 pith:DXZUSDSF submitted 2024-10-10 cs.CL

The Rise of AI-Generated Content in Wikipedia

classification cs.CL
keywords ai-generatedcontentarticleswikipedialowercreateddetectorspages
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The rise of AI-generated content in popular information sources raises significant concerns about accountability, accuracy, and bias amplification. Beyond directly impacting consumers, the widespread presence of this content poses questions for the long-term viability of training language models on vast internet sweeps. We use GPTZero, a proprietary AI detector, and Binoculars, an open-source alternative, to establish lower bounds on the presence of AI-generated content in recently created Wikipedia pages. Both detectors reveal a marked increase in AI-generated content in recent pages compared to those from before the release of GPT-3.5. With thresholds calibrated to achieve a 1% false positive rate on pre-GPT-3.5 articles, detectors flag over 5% of newly created English Wikipedia articles as AI-generated, with lower percentages for German, French, and Italian articles. Flagged Wikipedia articles are typically of lower quality and are often self-promotional or partial towards a specific viewpoint on controversial topics.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Can Generative Artificial Intelligence Survive Data Contamination? Theoretical Guarantees under Contaminated Recursive Training

    cs.LG 2026-02 conditional novelty 7.0

    Contaminated recursive training converges to the true distribution at rate t^{-min(p, α)} — the slower of the model's baseline rate p and the real-data fraction α.

  2. Comparing Human and Large Language Model Interpretation of Implicit Information

    cs.CL 2026-04 unverdicted novelty 5.0

    LLMs extract implicit information more conservatively than humans in social contexts but humans are more conservative in factual contexts, with humans proposing additional triplets overall.