REVIEW 10 cited by
Adversarial Example Generation with Syntactically Controlled Paraphrase Networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We propose syntactically controlled paraphrase networks (SCPNs) and use them to generate adversarial examples. Given a sentence and a target syntactic form (e.g., a constituency parse), SCPNs are trained to produce a paraphrase of the sentence with the desired syntax. We show it is possible to create training data for this task by first doing backtranslation at a very large scale, and then using a parser to label the syntactic transformations that naturally occur during this process. Such data allows us to train a neural encoder-decoder model with extra inputs to specify the target syntax. A combination of automated and human evaluations show that SCPNs generate paraphrases that follow their target specifications without decreasing paraphrase quality when compared to baseline (uncontrolled) paraphrase systems. Furthermore, they are more capable of generating syntactically adversarial examples that both (1) "fool" pretrained models and (2) improve the robustness of these models to syntactic variation when used to augment their training data.
Forward citations
Cited by 10 Pith papers
-
What the Eyes See, the LLMs Miss: Exploiting Human Perception for Adversarial Text Attacks
HPAA uses typographic manipulations to create text that humans flag as harmful at 86%+ rates while LLM moderation systems detect it below 1% with only three queries.
-
Potemkin Understanding in Large Language Models
LLMs frequently pass definition questions yet fail to use the same concepts in classification, generation, and editing tasks, a gap the authors call potemkin understanding.
-
Towards Action Hijacking of Large Language Model-based Agent
A RAG-based LLM application can be induced to assemble harmful SQL, code, or medical action plans from knowledge already stored in its database, with the user prompt itself carrying no forbidden words.
-
Mogrifier LSTM
Mogrifier LSTM, which applies repeated mutual gating between the input and previous hidden state, outperforms the LSTM on PTB, Wikitext-2, Enwik8, and MWC language modeling.
-
aiXamine: Simplified LLM Safety and Security
The paper presents aiXamine, a black-box LLM safety and security evaluation platform that aggregates 40+ existing benchmarks into 8 services, and reports a leaderboard of 16 models showing specific vulnerabilities in ...
-
Fake News Detection After LLM Laundering: Measurement and Explanation
LLM paraphrasing of fake news degrades detector performance across 17 detectors, with Pegasus evading best and a sentiment shift that BERTScore fails to capture.
-
Invisible Textual Backdoor Attacks based on Dual-Trigger
Combining a rare syntactic template with subjunctive mood as two triggers produces a textual backdoor with near-100% attack success and better defense resistance, though the comparison is partly confounded.
-
Adversarial Attacks in Multimodal Systems: A Practitioner's Survey
A practitioner-oriented survey that categorizes adversarial attacks across text, image, video, and audio modalities in multimodal AI systems.
-
A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations
A literature review that taxonomizes LLM backdoor attacks and defenses by model construction phase, with no new experimental results.
-
A Survey of Attacks on Large Language Models
A narrative survey that taxonomizes adversarial attacks on LLMs and LLM-based agents into training, inference, and availability/integrity phases with associated defenses.
Discussion (0). Continue with ORCID to comment.