Pith. sign in

REVIEW 1 cited by

Mutation-Based Adversarial Attacks on Neural Text Detectors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.05794 v1 pith:DYKEWYRX submitted 2023-02-11 cs.CR cs.AI

classification cs.CRcs.AI
keywords textadversarialattacksdetectorsmutationneuralcharacteristicsoriginal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Neural text detectors aim to decide the characteristics that distinguish neural (machine-generated) from human texts. To challenge such detectors, adversarial attacks can alter the statistical characteristics of the generated text, making the detection task more and more difficult. Inspired by the advances of mutation analysis in software development and testing, in this paper, we propose character- and word-based mutation operators for generating adversarial samples to attack state-of-the-art natural text detectors. This falls under white-box adversarial attacks. In such attacks, attackers have access to the original text and create mutation instances based on this original text. The ultimate goal is to confuse machine learning models and classifiers and decrease their prediction accuracy.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DivScore: Zero-Shot Detection of LLM-Generated Text in Specialized Domains

    cs.CL 2025-06 conditional novelty 6.0 of 10

    DivScore detects AI-written medical and legal text by dividing a domain-tuned model's entropy by its disagreement with a general model, beating baselines on a new benchmark.

Pith tools