Pith. sign in

REVIEW 2 cited by

Machine Generated Text: A Comprehensive Survey of Threat Models and Detection Methods

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.07321 v4 pith:LHV43FUM submitted 2022-10-13 cs.CL cs.CRcs.CYcs.LG

classification cs.CLcs.CRcs.CYcs.LG
keywords modelstextgeneratedmachinedetectionsurveysystemsthreat
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Machine generated text is increasingly difficult to distinguish from human authored text. Powerful open-source models are freely available, and user-friendly tools that democratize access to generative models are proliferating. ChatGPT, which was released shortly after the first edition of this survey, epitomizes these trends. The great potential of state-of-the-art natural language generation (NLG) systems is tempered by the multitude of avenues for abuse. Detection of machine generated text is a key countermeasure for reducing abuse of NLG models, with significant technical challenges and numerous open problems. We provide a survey that includes both 1) an extensive analysis of threat models posed by contemporary NLG systems, and 2) the most complete review of machine generated text detection methods to date. This survey places machine generated text within its cybersecurity and social context, and provides strong guidance for future work addressing the most critical threat models, and ensuring detection systems themselves demonstrate trustworthiness through fairness, robustness, and accountability.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BiMarker: Enhancing Text Watermark Detection for Large Language Models with Bipolar Watermarks

    cs.LG 2025-01 conditional novelty 6.0 of 10

    BiMarker splits generated text into alternating positive and negative poles and uses the difference in green-token counts to detect LLM watermarks more accurately than KGW.

  2. Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques

    cs.CR 2025-07 conditional novelty 4.0 of 10

    A survey that maps LLM applications, vulnerabilities, and defenses across eight cybersecurity domains, but with significant citation and rigor problems.

Pith tools