Pith. sign in

REVIEW 4 cited by

Assessing Adversarial Robustness of Large Language Models: An Empirical Study

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.02764 v2 pith:ZDY5Z3HV submitted 2024-05-04 cs.CL cs.LG

classification cs.CLcs.LG
keywords adversariallanguagellmsrobustnesslargemodelsacrossadvancement
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Large Language Models (LLMs) have revolutionized natural language processing, but their robustness against adversarial attacks remains a critical concern. We presents a novel white-box style attack approach that exposes vulnerabilities in leading open-source LLMs, including Llama, OPT, and T5. We assess the impact of model size, structure, and fine-tuning strategies on their resistance to adversarial perturbations. Our comprehensive evaluation across five diverse text classification tasks establishes a new benchmark for LLM robustness. The findings of this study have far-reaching implications for the reliable deployment of LLMs in real-world applications and contribute to the advancement of trustworthy AI systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Rethinking Testing for LLM Applications: Characteristics, Challenges, and a Lightweight Interaction Protocol

    cs.SE 2025-08 conditional novelty 5.0 of 10

    This position paper classifies testing methods for LLM applications into three layers and proposes AICL, a structured protocol for testable agent communication; neither the framework nor the protocol is empirically validated.

  2. Adversarial Attack on Large Language Models using Exponentiated Gradient Descent

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Exponentiated gradient descent over relaxed one-hot token encodings finds adversarial suffixes that jailbreak several open-source LLMs with higher success rate and lower runtime than GCG, PGD, and SoftPromptThreats.

  3. Bridging Robustness and Generalization Against Word Substitution Attacks in NLP via the Growth Bound Matrix Approach

    cs.CL 2025-07 reject novelty 4.0 of 10

    A Jacobian-magnitude regularization called GBM improves empirical robustness of CNN/LSTM/S4 text classifiers to synonym-substitution attacks, but the claimed certified robustness is not delivered.

  4. On Adversarial Robustness of Language Models in Transfer Learning

    cs.CL 2024-12 reject novelty 4.0 of 10

    Sequential fine-tuning across related bias-detection tasks tends to raise adversarial attack success rates, but the size-resilience pattern the paper highlights is not borne out by its own data.

Pith tools