Pith. sign in

REVIEW 2 cited by

Robustness of Large Language Models Against Adversarial Attacks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.17011 v1 pith:M5J7DLXX submitted 2024-12-22 cs.CL

classification cs.CL
keywords robustnessadversarialmodelsattacksllmscharacter-levelevaluationlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The increasing deployment of Large Language Models (LLMs) in various applications necessitates a rigorous evaluation of their robustness against adversarial attacks. In this paper, we present a comprehensive study on the robustness of GPT LLM family. We employ two distinct evaluation methods to assess their resilience. The first method introduce character-level text attack in input prompts, testing the models on three sentiment classification datasets: StanfordNLP/IMDB, Yelp Reviews, and SST-2. The second method involves using jailbreak prompts to challenge the safety mechanisms of the LLMs. Our experiments reveal significant variations in the robustness of these models, demonstrating their varying degrees of vulnerability to both character-level and semantic-level adversarial attacks. These findings underscore the necessity for improved adversarial training and enhanced safety mechanisms to bolster the robustness of LLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems

    cs.CR 2025-09 conditional novelty 6.0 of 10

    A systematic review that categorizes LLM threats, severity scores, and mitigations across development and operation life cycles and multiple deployment scenarios.

  2. SSH: Sparse Spectrum Adaptation via Discrete Hartley Transformation

    cs.CV 2025-02 conditional novelty 3.0 of 10

    SSH fine-tunes large models by learning sparse Hartley-spectrum coefficients selected by energy of the pretrained weights, matching or beating LoRA and FourierFT with fewer parameters.

Pith tools