Pith. sign in

REVIEW 10 cited by

Adversarial Example Generation with Syntactically Controlled Paraphrase Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1804.06059 v1 pith:MHIMPFWB submitted 2018-04-17 cs.CL

classification cs.CL
keywords paraphraseadversarialdatascpnssyntacticsyntacticallytargetcontrolled
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We propose syntactically controlled paraphrase networks (SCPNs) and use them to generate adversarial examples. Given a sentence and a target syntactic form (e.g., a constituency parse), SCPNs are trained to produce a paraphrase of the sentence with the desired syntax. We show it is possible to create training data for this task by first doing backtranslation at a very large scale, and then using a parser to label the syntactic transformations that naturally occur during this process. Such data allows us to train a neural encoder-decoder model with extra inputs to specify the target syntax. A combination of automated and human evaluations show that SCPNs generate paraphrases that follow their target specifications without decreasing paraphrase quality when compared to baseline (uncontrolled) paraphrase systems. Furthermore, they are more capable of generating syntactically adversarial examples that both (1) "fool" pretrained models and (2) improve the robustness of these models to syntactic variation when used to augment their training data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. What the Eyes See, the LLMs Miss: Exploiting Human Perception for Adversarial Text Attacks

    cs.CR 2026-06 unverdicted novelty 6.0 of 10

    HPAA uses typographic manipulations to create text that humans flag as harmful at 86%+ rates while LLM moderation systems detect it below 1% with only three queries.

  2. Potemkin Understanding in Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    LLMs frequently pass definition questions yet fail to use the same concepts in classification, generation, and editing tasks, a gap the authors call potemkin understanding.

  3. Towards Action Hijacking of Large Language Model-based Agent

    cs.CR 2024-12 conditional novelty 6.0 of 10

    A RAG-based LLM application can be induced to assemble harmful SQL, code, or medical action plans from knowledge already stored in its database, with the user prompt itself carrying no forbidden words.

  4. Mogrifier LSTM

    cs.CL 2019-09 accept novelty 6.0 of 10

    Mogrifier LSTM, which applies repeated mutual gating between the input and previous hidden state, outperforms the LSTM on PTB, Wikitext-2, Enwik8, and MWC language modeling.

  5. aiXamine: Simplified LLM Safety and Security

    cs.CR 2025-04 conditional novelty 5.0 of 10

    The paper presents aiXamine, a black-box LLM safety and security evaluation platform that aggregates 40+ existing benchmarks into 8 services, and reports a leaderboard of 16 models showing specific vulnerabilities in ...

  6. Fake News Detection After LLM Laundering: Measurement and Explanation

    cs.CL 2025-01 conditional novelty 5.0 of 10

    LLM paraphrasing of fake news degrades detector performance across 17 detectors, with Pegasus evading best and a sentiment shift that BERTScore fails to capture.

  7. Invisible Textual Backdoor Attacks based on Dual-Trigger

    cs.CR 2024-12 conditional novelty 4.0 of 10

    Combining a rare syntactic template with subjunctive mood as two triggers produces a textual backdoor with near-100% attack success and better defense resistance, though the comparison is partly confounded.

  8. Adversarial Attacks in Multimodal Systems: A Practitioner's Survey

    cs.LG 2025-05 conditional novelty 3.0 of 10

    A practitioner-oriented survey that categorizes adversarial attacks across text, image, video, and audio modalities in multimodal AI systems.

  9. A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations

    cs.CR 2025-02 conditional novelty 2.0 of 10

    A literature review that taxonomizes LLM backdoor attacks and defenses by model construction phase, with no new experimental results.

  10. A Survey of Attacks on Large Language Models

    cs.CR 2025-05 conditional novelty 1.0 of 10

    A narrative survey that taxonomizes adversarial attacks on LLMs and LLM-based agents into training, inference, and availability/integrity phases with associated defenses.

Pith tools