Pith. sign in

REVIEW 2 major objections 5 minor 4 cited by

LLM Security: Vulnerabilities, Attacks, Defenses, and Countermeasures

T0 review · 2 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read This survey maps attacks on large language models across training and deployment and concludes that only a handful of defenses work well, while most existing countermeasures have exploitable limitations.

desk verdict A competent, readable survey whose only real new contribution—the Table 2 effectiveness matrix—lacks a method and contains at least one High rating that its own cited source seems to contradict. read the letter →

arxiv 2505.01177 v1 pith:F53A2KXO submitted 2025-05-02 cs.CR cs.AIcs.LGcs.NE

classification cs.CRcs.AIcs.LGcs.NE
keywords largelanguagemodelsLLMsecurityadversarialattackspromptinjectionjailbreakingmembershipinferencedefensemechanismssurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Large language models are vulnerable to attacks both while they are being trained and after they are deployed, spanning backdoors and data poisoning on the training side and jailbreaking, prompt injection, data extraction, membership inference, and inversion attacks at inference time. This survey organizes that landscape by classifying defenses into prevention-based and detection-based mechanisms, and it maps each defense to the attacks it mitigates. Its central assessment is that only a handful of defenses—such as retokenization, SmoothLLM, differential privacy against membership inference, fine-tuning against backdoors, and dimensional masking against embedding inversion—are highly effective. Most other countermeasures have inherent limitations that informed and skilled attackers can exploit, which matters because it suggests that securing real LLM applications requires either stronger defenses or layered combinations rather than reliance on any single mechanism.

What carries the argument

The organizing device is a two-axis classification scheme. Attacks are split by lifecycle phase into causative attacks that corrupt the model during training (backdoor, data poisoning, gradient leakage) and exploratory attacks that target the trained model at inference (adversarial inputs, prompt hacking, model inversion, data extraction, membership inference, embedding inversion). Defenses are split into prevention-based mechanisms that harden or transform the input or model, and detection-based mechanisms that flag suspicious prompts or responses. The argument is carried by a coverage matrix (Table 2) listing each attack, the defenses that apply to it, and an effectiveness rating of High, Moderate, or Low.

What would settle it

A direct check would be to take a fixed set of representative attacks (for example, a standardized jailbreak suite and a prompt-injection benchmark) and measure attack success rates against a fixed model with and without each rated defense; if, say, retokenization or SmoothLLM fails to keep success rates near zero, the 'High' rating would be overturned.

Watch

Extended reading notes

Core claim

The paper's core claim is that the current defensive toolbox for LLMs is mostly partial rather than comprehensive. In the authors' assessment, highly effective defenses exist for specific attack types: retokenization for adversarial inputs, jailbreaking, and prompt injection; SmoothLLM for jailbreaking; fine-tuning and fine-mixing for backdoor removal; anomaly detection for data poisoning; differential privacy for membership inference; and dimensional masking for embedding inversion. For the remaining attack-defense combinations, effectiveness is moderate or low, and for model inversion attacks the survey finds no known defense at all. The authors therefore conclude that only a handful of defenses are highly effective and that most existing mechanisms present inherent limitations that can be exploited by informed and sufficiently-skilled malicious actors.

Load-bearing premise

The load-bearing assumption is that the effectiveness ratings in Table 2—High, Moderate, or Low for each defense—are trustworthy, but they are assigned without a described scoring rubric, per-cell citations, or shared benchmarks, so the survey's practical guidance stands or falls on an assessment the reader cannot yet verify.

Editorial extensions

If this is right

  • Practitioners should expect single defenses to be bypassable: the paper rates most countermeasures as moderate, meaning an informed attacker can work around them.
  • For high-stakes deployments, the results imply combining defenses—for example, retokenization with detection—rather than relying on one mechanism.
  • Model inversion attacks currently have no known defense, so applications exposing detailed model outputs should treat reconstruction of training data as an open risk.
  • Differential privacy is singled out as the effective path against membership inference, but the survey notes its practical effectiveness against gradient leakage is low, so privacy guarantees depend on the specific attack.
  • Human-crafted, low-perplexity attacks evade perplexity-based detection, so detection defenses are not a substitute for input hardening.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The effectiveness ratings in the coverage matrix are presented without a scoring rubric or per-cell benchmarks; a natural next step would be a standardized evaluation that reproduces the ratings under fixed attack suites and models.
  • If the ratings hold, retokenization's 'High' rating suggests token-boundary perturbation is a particularly robust input transformation, and it would be worth testing against adaptive adversaries that optimize triggers after seeing the defense.
  • The absence of any known defense against model inversion points to a concrete research gap: mechanisms that limit confidence or embedding information leakage, perhaps along the lines of the invertibility-interference noise effect the survey mentions, deserve attention.
  • The survey's taxonomy could be extended to emerging multimodal and agentic LLM systems, where the input space is larger and the injection surface includes tool calls and retrieval results.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper is a survey of security threats to large language models (LLMs) that organizes attacks according to whether they occur during training (backdoor, data poisoning, gradient leakage) or after deployment (adversarial inputs, prompt hacking, model inversion, data extraction, membership inference, embedding inversion). It then reviews defense mechanisms, partitioned into prevention-based and detection-based categories, and presents two tables: Table 1 summarizes attacks and their targets (model integrity vs. data privacy), and Table 2 maps defenses to attacks and assigns qualitative effectiveness ratings (High, Moderate, Low). The paper concludes that only a handful of defenses are highly effective and that most suffer from inherent limitations. The distinctive contribution is the coverage matrix with effectiveness ratings.

Significance. If the effectiveness ratings were substantiated, the survey would be a useful structured guide for practitioners and researchers, because it brings together a broad range of attacks and defenses in one framework and provides mathematical descriptions of key mechanisms (e.g., backdoor poisoning, gradient leakage, SmoothLLM). The organization by lifecycle phase is clear and the coverage is extensive relative to prior surveys. However, the paper's distinctive contribution—the effectiveness assessment—is asserted rather than demonstrated, and at least one rating conflicts with its cited source. Since the conclusion is a direct restatement of these ratings, the practical value of the survey is currently limited. The paper does not provide machine-checked proofs or reproducible code; its claims are qualitative and would need an added evidence-based rubric to be actionable.

major comments (2)
  1. [Section 5, Table 2] The Effectiveness column in Table 2 is the paper's central contribution, but the High/Moderate/Low ratings are assigned without a scoring rubric, per-cell citations, or shared benchmarks. The 'quick recap' in Section 5 provides only mechanism-level rationales; for example, retokenization is called 'highly effective' against jailbreaking because 'changing token segmentation alters how the model interprets the input at a more fundamental level,' which is a claim about mechanism, not a measurement. More concretely, the High rating for retokenization against jailbreaking appears to conflict with the primary source cited for baseline defenses, Jain et al. (2023, arXiv:2309.00614, reference [65]), which evaluates retokenization as a baseline and finds that it does not reliably neutralize GCG-style attacks, whereas SmoothLLM is the more effective defense in that study. Because the conclusion that 'only a handful of defenses are highly effective' is a direct restatement of these ratings, the central claim is currently unverifiable. The authors should specify a transparent scoring method, provide per-cell evidence with citations, and correct ratings that conflict with their sources.
  2. [Section 3.2.4] This section, titled 'Data extraction attacks,' opens by describing attacks that 'exploit publicly accessible prediction APIs' to 'replicate a model functionality' and reconstruct a 'high-fidelity surrogate model.' That is model extraction/stealing, not data extraction, and it conflicts with the scope statement in Section 2.2, which explicitly excludes model stealing attacks. The section then shifts to extraction of memorized text instances, which is data extraction proper. These are different attack classes with different threat models, and the conflation makes the taxonomy in Table 1 and the defense mapping in Table 2 ambiguous. The authors should either remove the model-extraction material (consistent with the stated scope) or split it into a separate category and adjust the coverage discussion accordingly.
minor comments (5)
  1. [Title and headings] The string 'A/t_tacks' appears in the title, in the Section 3 and Section 4 headings, and in the captions of Tables 1 and 2; this should be 'Attacks.' The title also contains 'Count ermeasures' instead of 'Countermeasures.'
  2. [Section 2.1] The equations describing token probability and generation are garbled (e.g., 'p(⟨t1..tn⟩) = ˆp_t_{n+1}' and 'φ(input) = ˆoutput, for short'); the notation needs to be typeset cleanly so that the definitions are readable.
  3. [Table 1 caption] The caption contains a typo: 'whether they taget model integrity and/or data privacy' should read 'target.'
  4. [Section 5] The phrase 'providing an eagled-eyed view view' has a duplicated word and should be 'an eagle-eyed view.'
  5. [References] Reference [106] appears to have an incorrect author list for the GPT-1 paper (it lists 'Aditya Kwon, Marloes Koot' rather than the actual authors), and reference [8] ends with 'Just Accepted' rather than full publication details. The reference list should be carefully checked against the sources.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey's conclusions are aggregative summaries of cited literature, and the authors' self-citations are background only.

full rationale

This is a survey paper, not a derivation or prediction pipeline. It fits no parameters, proves no theorems, and presents no new empirical evaluation. The central conclusion in Section 6, that 'only a handful of defenses are highly effective' and that most defenses have 'inherent limitations', is an aggregative summary of the qualitative Effectiveness column in Table 2. Aggregating one's own evaluative table into a conclusion is a summary, not a circular derivation, because the table's ratings are not themselves defined in terms of the conclusion. The two self-citations, references [12] and [13] by the second author, are cited only in the introduction as general deep-learning background ('deep learning in identifying complex patterns within large datasets [46][12][13]'); they are not load-bearing for any attack, defense, or effectiveness claim. No uniqueness theorem, ansatz, or fitted input is imported from the authors' prior work, and no technical claim reduces to its own inputs by construction. The main legitimate concern is that Table 2's effectiveness ratings (e.g., retokenization 'High', SmoothLLM 'High', differential privacy 'High' for membership inference) are presented without a scoring rubric, per-cell citations, or shared benchmarks, and at least one rating may conflict with its cited source. That is a support/reproducibility and correctness risk, not circularity, because the ratings are not derived from the paper's own outputs. Under the circularity standard requiring a specific reduction (Eq. X = Eq. Y by construction, or a fitted parameter renamed as a prediction), no such step exists here.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

As a survey, the paper introduces no fitted parameters and no new theoretical entities. Its claims rest on the accuracy of the cited primary literature, the completeness of the informal taxonomy, and the validity of the authors' qualitative effectiveness ratings.

assumptions (3)
  • domain assumption The cited primary papers accurately report the effectiveness of their proposed attacks and defenses.
    The survey does not re-implement any attack or defense; every technical claim in Sections 3 and 4 rests on the trustworthiness of the cited literature.
  • domain assumption The taxonomy (training-time vs. deployed attacks, prevention vs. detection defenses) is exhaustive enough to organize the LLM security landscape.
    The paper gives no systematic search or inclusion criteria, so the completeness of the categories and the selected references is assumed rather than demonstrated.
  • ad hoc to paper Qualitative effectiveness levels such as High and Moderate can be meaningfully assigned without a common benchmark.
    Table 2 assigns these levels across heterogeneous threat models without a scoring rubric, per-cell citations, or quantitative metrics, making the ratings an authorial judgment.

how reviews work

0 comments
Cite this review

Pith. "Pith review of LLM Security: Vulnerabilities, Attacks, Defenses, and Countermeasures." pith.science (2026). https://pith.science/paper/F53A2KXO

@misc{pith2026250501177,
  author       = {Pith},
  title        = {Pith review of: LLM Security: Vulnerabilities, Attacks, Defenses, and Countermeasures},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/F53A2KXO}},
  note         = {Machine review of arXiv:2505.01177}
}
read the original abstract

As large language models (LLMs) continue to evolve, it is critical to assess the security threats and vulnerabilities that may arise both during their training phase and after models have been deployed. This survey seeks to define and categorize the various attacks targeting LLMs, distinguishing between those that occur during the training phase and those that affect already trained models. A thorough analysis of these attacks is presented, alongside an exploration of defense mechanisms designed to mitigate such threats. Defenses are classified into two primary categories: prevention-based and detection-based defenses. Furthermore, our survey summarizes possible attacks and their corresponding defense strategies. It also provides an evaluation of the effectiveness of the known defense mechanisms for the different security threats. Our survey aims to offer a structured framework for securing LLMs, while also identifying areas that require further research to improve and strengthen defenses against emerging security challenges.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Uncovering Vulnerabilities of LLM-Assisted Cyber Threat Intelligence

    cs.CR 2025-09 conditional novelty 6.0 of 10

    LLMs assisting cyber threat intelligence fail mainly due to spurious correlations, contradictory knowledge, and constrained generalization that stem from the threat landscape itself.

  2. Secure Multi-LLM Agentic AI and Agentification for Edge General Intelligence by Zero-Trust: A Survey

    cs.NI 2025-08 conditional novelty 4.0 of 10

    A survey proposing zero-trust architecture for multi-LLM systems in edge computing, with a taxonomy of model- and system-level defenses and a conceptual framework.

  3. SoK: Semantic Privacy in Large Language Models

    cs.CR 2025-06 conditional novelty 4.0 of 10

    A systematization of knowledge arguing that LLM privacy threats extend beyond data leakage to semantically inferred attributes, and that current defenses only partially address them.

  4. A Survey on Model Extraction Attacks and Defenses for Large Language Models

    cs.CR 2025-06 conditional novelty 4.0 of 10

    A taxonomy of model extraction attacks and defenses for large language models, with proposed evaluation metrics and future research directions.

Reference graph

Works this paper leans on

182 extracted references · 5 canonical work pages · cited by 4 Pith papers

  1. [127]

    Deep Learning: Learning Techniques and Applic ations

    Jorge Sánchez-González. 2020. Sentiment Analysis in Twitter, in “Deep Learning: Learning Techniques and Applic ations” [in Spanish]. B.Eng. capstone project, University of Granada

  2. [65]

    Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Some palli, John Kirchenbauer, Ping yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein. 2023. Baseline Defenses for Adv ersarial Attacks Against Aligned Language Models. arXiv: 2309.00614 [cs.LG] https://arxiv.org/abs/2309.00614

  3. [1]

    Brendan McMah an, Ilya Mironov, Kunal Talwar, and Li Zhang

    Martin Abadi, Andy Chu, Ian Goodfellow, H. Brendan McMah an, Ilya Mironov, Kunal Talwar, and Li Zhang. 2016. Deep Learning with Differential Privacy. In Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security (Vienna, Austria) (CCS ’16) . Association for Computing Machinery, New York, NY, USA, 308–318. https://doi.org/10.1145/...

  4. [2]

    Sara Abdali, Richard Anarfi, CJ Barberan, and Jia He. 2024 . Securing Large Language Models: Threats, Vulnerabilitie s and Responsible Practices. arXiv:2403.12503 [cs.CR] https://arxiv.org/abs/2403.12503

  5. [3]

    Sara Abdali, Jia He, CJ Barberan, and Richard Anarfi. 2024 . Can LLMs be Fooled? Investigating Vulnerabilities in LLMs . arXiv: 2407.20529 [cs.LG] https://arxiv.org/abs/2407.20529

  6. [4]

    Microsoft Research AI4Science and Microsoft Azure Quan tum. 2023. The Impact of Large Language Models on Scientific Discovery: a Preliminary Study using GPT-4. arXiv: 2311.07361 [cs.CL] https://arxiv.org/abs/2311.07361

  7. [5]

    Ali Al-Kaswan, Maliheh Izadi, and Arie van Deursen. 2023 . Targeted Attack on GPT-Neo for the SATML Language Model Dat a Extraction Challenge. arXiv: 2302.07735 [cs.CL] https://arxiv.org/abs/2302.07735

  8. [6]

    Gabriel Alon and Michael Kamfonas. 2023. Detecting Lang uage Model Attacks with Perplexity. arXiv: 2308.14132 [cs.CL] https://arxiv.org/abs/2308.14132

Show all 182 references
  1. [7]

    Anthropic. 2024. Introducing Claude 3.5 Sonnet. https://www.anthropic.com/news/claude-3-5-sonnet . Accessed: 2024-11-10

  2. [8]

    Li Bai, Haibo Hu, Qingqing Ye, Haoyang Li, Leixia Wang, an d Jianliang Xu. 2024. Membership Inference Attacks and Defe nses in Federated Learning: A Survey. https://doi.org/10.1145/3704633 Just Accepted

  3. [9]

    Mislav Balunovic, Dimitar Dimitrov, Nikola Jovanović, and Martin Vechev. 2022. LAMP: Extracting Text from Gradients with Language Model Priors. In Advances in Neural Information Processing Systems , S. Koyejo, S. Mo- hamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vo...

  4. [10]

    Joseph, and J

    Marco Barreno, Blaine Nelson, Russell Sears, Anthony D . Joseph, and J. D. Tygar. 2006. Can machine learning be secure?. In Proceedings of the 2006 ACM Symposium on Information, Computer and Communications Security (Taipei, Taiwan) (ASIACCS ’06). Association for Computing Mach...

  5. [11]

    Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Ch ristian Jauvin. 2003. A Neural Probabilistic Language Mode l. Journal of Machine Learning Research 3 (2003), 1137–1155

  6. [12]

    Fernando Berzal. 2019. Redes Neuronales & Deep Learning - Volumen I: Entrenamiento de redes neuronales artificiales [Neural Networks and Deep Learning - Volume I: Training Artificial Neural Networks, in Spanish]. Amazon KDP, Granada, Spain. https://deep-learning.ikor.org/

  7. [13]

    Fernando Berzal. 2019. Redes Neuronales & Deep Learning - Volumen II: Regularizaci ón, optimización & arquitecturas especializadas [Neural N et- works and Deep Learning - Volume II: Regularization, Optimi zation, and Specialized Architectures, in Spanish] . Amazon KDP, Granada...

  8. [14]

    Satwik Bhattamishra, Arkil Patel, Phil Blunsom, and Va run Kanade. 2023. Understanding In-Context Learning in Tra nsformers and LLMs by Learning to Learn Discrete Functions. arXiv: 2310.03016 [cs.LG] https://arxiv.org/abs/2310.03016

  9. [15]

    Battista Biggio, Igino Corona, Davide Maiorca, Blaine Nelson, Nedim Šrndić, Pavel Laskov, Giorgio Giacinto, and Fabio Roli. 2013. Evasion Attacks against Machine Learning at Test Time. In Machine Learning and Knowledge Discovery in Databases, Hendrik Blockeel, Kristian Kerstin...

  10. [16]

    Lewis Birch, William Hackett, Stefan Trawicki, Neeraj Suri, and Peter Garraghan. 2023. Model Leeching: An Extract ion Attack Targeting LLMs. arXiv:2309.10544 [cs.LG] https://arxiv.org/abs/2309.10544

  11. [17]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbia h, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...

  12. [18]

    Xiangrui Cai, Haidong Xu, Sihan Xu, Ying Zhang, and Xiao jie Yuan. 2022. BadPrompt: Backdoor Attacks on Continuous P rompts. arXiv:2211.14719 [cs.CL] https://arxiv.org/abs/2211.14719

  13. [19]

    Nicholas Carlini, Daphne Ippolito, Matthew Jagielski , Katherine Lee, Florian Tramer, and Chiyuan Zhang. 2023. Quantifying Memorization Across Neural Language Models. arXiv: 2202.07646 [cs.LG] https://arxiv.org/abs/2202.07646

  14. [20]

    C hoquette-Choo, Daniel Paleka, Will Pearce, Hyrum Anderson, Andreas Terzis, Kurt Thomas, and Florian Tramèr

    Nicholas Carlini, Matthew Jagielski, Christopher A. C hoquette-Choo, Daniel Paleka, Will Pearce, Hyrum Anderson, Andreas Terzis, Kurt Thomas, and Florian Tramèr. 2024. Poisoning Web-Scale Training Dat asets is Practical. arXiv: 2302.10149 [cs.CR] https://arxiv.org/abs/2302.10149

  15. [21]

    Nicholas Carlini, Florian Tramèr, Eric Wallace, Matth ew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Robe rts, Tom Brown, Dawn Song, Úlfar Erlingsson, Alina Oprea, and Colin R affel. 2021. Extracting Training Data from Large Language Models. In 30th USENIX Security Sympo...

  16. [22]

    Tarek Chaalan, Shaoning Pang, Joarder Kamruzzaman, Iq bal Gondal, and Xuyun Zhang. 2024. The Path to Defence: A Road map to Characterising Data Poisoning Attacks on Victim Models. ACM Comput. Surv. 56, 7, Article 175 (April 2024), 39 pages. https://doi.org/10.1145/3627536

  17. [23]

    Wenhan Chang and Tianqing Zhu. 2024. Gradient-based de fense methods for data leakage in vertical federated learni ng. Computers & Security 139 (2024), 103744. https://doi.org/10.1016/j.cose.2024.103744

  18. [24]

    Yu, Qiang Yang, and Xing Xie

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang , Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wa ng, Wei Ye, Yue Zhang, Yi Chang, Philip S. Yu, Qiang Yang, and Xing Xie. 2024. A Survey on Evaluation of Large Language Models. ACM Trans. Intell. Syst. Techn...

  19. [25]

    Shreyas Chaudhari, Pranjal Aggarwal, Vishvak Murahar i, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, A meet Deshpande, and Bruno Castro da Silva. 2024. RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs. arXiv:2404.08555 [cs.LG...

  20. [26]

    Yiyi Chen, Heather Lent, and Johannes Bjerva. 2024. Tex t Embedding Inversion Security for Multilingual Language M odels. arXiv:2401.12192 [cs.CL] https://arxiv.org/abs/2401.12192

  21. [27]

    Gonzalez, and Ion Stoica

    Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica. 2024. Chatbot Arena: An O pen Platform for Evaluating LLMs by Human Preference. arXiv :2403.04132 [cs.AI]

  22. [28]

    Junjie Chu, Yugeng Liu, Ziqing Yang, Xinyue Shen, Micha el Backes, and Yang Zhang. 2024. Comprehensive Assessment o f Jailbreak Attacks Against LLMs. arXiv: 2402.05668 [cs.CR] https://arxiv.org/abs/2402.05668

  23. [29]

    Consens, Cameron Dufault, Michael Wainberg , Duncan Forster, Mehran Karimzadeh, Hani Goodarzi, Fabian J

    Micaela E. Consens, Cameron Dufault, Michael Wainberg , Duncan Forster, Mehran Karimzadeh, Hani Goodarzi, Fabian J. Theis, Alan Moses, and Bo Wang. 2023. To Transformers and Beyond: Large L anguage Models for the Genome. arXiv: 2311.07621 [q-bio.GN] https://arxiv.org/abs/2311.07621

  24. [30]

    Avisha Das, Amara Tariq, Felipe Batalini, Boddhisattw a Dhara, and Imon Banerjee. 2024. Exposing Vulnerabilities in Clinical LLMs Through Data Poisoning Attacks: Case Study in Breast Cancer . medRxiv 1, 1 (2024), 1–1. https://doi.org/10.1101/2024.03.20.24304627 arXiv:https://w...

  25. [31]

    Hadi Amini, and Yanzhao Wu

    Badhan Chandra Das, M. Hadi Amini, and Yanzhao Wu. 2024. Security and Privacy Challenges of Large Language Models: A Survey. arXiv:2402.00888 [cs.CL] https://arxiv.org/abs/2402.00888

  26. [32]

    Ms D Deepa et al. 2021. Bidirectional encoder represent ations from transformers (BERT) language model for sentime nt analysis task. Turkish Journal of Computer and Mathematics Education (TUR COMAT) 12, 7 (2021), 1708–1721. URL: 22 Aguilera-Martínez and Berzal https://turcomat...

  27. [33]

    Jieren Deng, Yijue Wang, Ji Li, Chao Shang, Hang Liu, San guthevar Rajasekaran, and Caiwen Ding. 2021. TAG: Gradient Attack on Transformer- based Language Models. arXiv: 2103.06819 [cs.CR] https://arxiv.org/abs/2103.06819

  28. [34]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv: 1810.04805 [cs.CL] https://arxiv.org/abs/1810.04805

  29. [35]

    Hanze Dong, Wei Xiong, Deepanshu Goyal, Yihan Zhang, Wi nnie Chow, Rui Pan, Shizhe Diao, Jipeng Zhang, Kashun Shum, a nd Tong Zhang

  30. [36]

    Michael Duan, Anshuman Suri, Niloofar Mireshghallah, Sewon Min, Weijia Shi, Luke Zettlemoyer, Yulia Tsvetkov, Ye jin Choi, David Evans, and Hannaneh Hajishirzi. 2024. Do Membership Infere nce Attacks Work on Large Language Models? arXiv: 2402.07841 [cs.CL] https://arxiv.org/ab...

  31. [37]

    Cynthia Dwork. 2006. Differential Privacy. In Automata, Languages and Programming , Michele Bugliesi, Bart Preneel, Vladimiro Sassone, and Ingo Wegener (Eds.). Springer Berlin Heidelberg, Berlin, H eidelberg, 1–12

  32. [38]

    Cynthia Dwork, Frank McSherry, Kobbi Nissim, and Adam S mith. 2006. Calibrating Noise to Sensitivity in Private Dat a Analysis. In Theory of Cryptography, Shai Halevi and Tal Rabin (Eds.). Springer Berlin Heidelbe rg, Berlin, Heidelberg, 265–284

  33. [39]

    Cynthia Dwork and Kobbi Nissim. 2004. Privacy-Preserv ing Datamining on Vertically Partitioned Databases. InAdvances in Cryptology – CRYPTO 2004, Matt Franklin (Ed.). Springer Berlin Heidelberg, Berlin, Heidelberg, 528–544

  34. [40]

    Amine Elhafsi, Rohan Sinha, Christopher Agia, Edward S chmerling, Issa A. D. Nesnas, and Marco Pavone. 2023. Semant ic anomaly detection with large language models. Autonomous Robots 47, 8 (2023), 1035–1055. https://doi.org/10.1007/s10514-023-10132-6

  35. [41]

    Aysan Esmradi, Daniel Wankit Yip, and Chun Fai Chan. 202 4. A Comprehensive Survey of Attack Techniques, Implementa tion, and Mitigation Strategies in Large Language Models. In Ubiquitous Security, Guojun Wang, Haozhe Wang, Geyong Min, Nektarios Georgalas , and Weizhi Meng (Ed...

  36. [42]

    Verykios

    Georgios Feretzakis, Konstantinos Papaspyridis, Ari s Gkoulalas-Divanis, and Vassilios S. Verykios. 2024. Priv acy-Preserving Techniques in Gen- erative AI and Large Language Models: A Narrative Review. Information 15, 11 (2024). https://doi.org/10.3390/info15110697

  37. [43]

    Matt Fredrikson, Somesh Jha, and Thomas Ristenpart. 20 15. Model Inversion Attacks that Exploit Confidence Informa tion and Basic Counter- measures. In Proceedings of the 22nd ACM SIGSAC Conference on Computer and Communications Security (Denver, Colorado, USA) (CCS ’15) . Asso...

  38. [44]

    Xiaopeng Fu, Zhaoquan Gu, Weihong Han, Yaguan Qian, Bin Wang, and Federico Tramarin. 2021. Exploring Security Vuln erabilities of Deep Learning Models by Adversarial Attacks. Wirel. Commun. Mob. Comput. 2021 (Jan. 2021), 9 pages. https://doi.org/10.1155/2021/9969867

  39. [45]

    Jonas Geiping, Hartmut Bauermeister, Hannah Dröge, an d Michael Moeller. 2020. Inverting Gradients – How easy is it to break privacy in federated learning? arXiv: 2003.14053 [cs.CV] https://arxiv.org/abs/2003.14053

  40. [46]

    Ian Goodfellow, Yoshua Bengio, and Aaron Courville. 20 16. Deep Learning. MIT Press, USA. http://www.deeplearningbook.org

  41. [47]

    Goodfellow, Jonathon Shlens, and Christian Szeg edy

    Ian J. Goodfellow, Jonathon Shlens, and Christian Szeg edy. 2015. Explaining and Harnessing Adversarial Examples . arXiv: 1412.6572 [stat.ML] https://arxiv.org/abs/1412.6572

  42. [48]

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Ab hinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal , Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravank umar, Artem Korenev, A...

  43. [49]

    Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Chris toph Endres, Thorsten Holz, and Mario Fritz. 2023. Not what y ou’ve signed up for: Com- promising Real-World LLM-Integrated Applications with In direct Prompt Injection. arXiv: 2302.12173 [cs.CR] https://arxiv.org/abs/2302.12173

  44. [50]

    Tianyu Gu, Brendan Dolan-Gavitt, and Siddharth Garg. 2 019. BadNets: Identifying Vulnerabilities in the Machine L earning Model Supply Chain. arXiv:1708.06733 [cs.CR] https://arxiv.org/abs/1708.06733

  45. [51]

    Chuan Guo, Alexandre Sablayrolles, Hervé Jégou, and Do uwe Kiela. 2021. Gradient-based Adversarial Attacks again st Text Transformers. arXiv:2104.13733 [cs.CL] https://arxiv.org/abs/2104.13733

  46. [52]

    Muhammad Usman Hadi, Rizwan Qureshi, Abbas Shah, Muham mad Irfan, Anas Zafar, Muhammad Bilal Shaikh, Naveed Akhtar, Jia Wu, Seyedali Mirjalili, et al. 2023. A survey on large language models: Ap plications, challenges, limitations, and practical usage

  47. [53]

    Hampshire and A

    J.B. Hampshire and A. Waibel. 1992. The Meta-Pi network : building distributed knowledge representations for robu st multisource pattern recog- nition. IEEE Transactions on Pattern Analysis and Machine Intellige nce 14, 7 (1992), 751–769. https://doi.org/10.1109/34.142911

  48. [54]

    Douglas M. Hawkins. 1980. Identification of Outliers . Chapman and Hall, Kluwer Academic Publishers, Boston/Dor drecht/London

  49. [55]

    Jamie Hayes, Luca Melis, George Danezis, and Emiliano D e Cristofaro. 2018. LOGAN: Membership Inference Attacks Ag ainst Generative Models. arXiv:1705.07663 [cs.CR] https://arxiv.org/abs/1705.07663

  50. [56]

    Kai He, Rui Mao, Qika Lin, Yucheng Ruan, Xiang Lan, Mengl ing Feng, and Erik Cambria. 2024. A Survey of Large Language Models for Healthcare: from Data, Technology, and Applications to Accountability and Ethics. arXiv: 2310.05694 [cs.CL] https://arxiv.org/abs/2310.05694

  51. [57]

    Keegan Hines, Gary Lopez, Matthew Hall, Federico Zarfa ti, Yonatan Zunger, and Emre Kiciman. 2024. Defending Against Indirect Prompt Injection Attacks With Spotlighting. arXiv: 2403.14720 [cs.CR] https://arxiv.org/abs/2403.14720 24 Aguilera-Martínez and Berzal

  52. [58]

    Yu, and Xuyun Zhang

    Hongsheng Hu, Zoran Salcic, Lichao Sun, Gillian Dobbie , Philip S. Yu, and Xuyun Zhang. 2022. Membership Inference A ttacks on Machine Learning: A Survey. ACM Comput. Surv. 54, 11s, Article 235 (Sept. 2022), 37 pages. https://doi.org/10.1145/3523273

  53. [59]

    Li Hu, Anli Yan, Hongyang Yan, Jin Li, Teng Huang, Yingyi ng Zhang, Changyu Dong, and Chunsheng Yang. 2023. Defenses t o Membership Inference Attacks: A Survey. ACM Comput. Surv. 56, 4, Article 92 (Nov. 2023), 34 pages. https://doi.org/10.1145/3620667

  54. [60]

    Hai Huang, Zhengyu Zhao, Michael Backes, Yun Shen, and Y ang Zhang. 2024. Composite Backdoor Attacks Against Large L anguage Models. arXiv:2310.07676 [cs.CR] https://arxiv.org/abs/2310.07676

  55. [61]

    Jie Huang, Hanyin Shao, and Kevin Chen-Chuan Chang. 202 2. Are Large Pre-Trained Language Models Leaking Your Perso nal Information?. In Findings of the Association for Computational Linguistics : EMNLP 2022 , Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). As sociatio...

  56. [62]

    Yunpeng Huang, Jingwei Xu, Junyu Lai, Zixu Jiang, Taolu e Chen, Zenan Li, Yuan Yao, Xiaoxing Ma, Lijuan Yang, Hao Chen , Shupeng Li, and Penghao Zhao. 2024. Advancing Transformer Architec ture in Long-Context Large Language Models: A Comprehensiv e Survey. arXiv:2311.12351 [cs...

  57. [63]

    Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Run gta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, B rian Fuller, Davide Tes- tuggine, and Madian Khabsa. 2023. Llama Guard: LLM-based In put-Output Safeguard for Human-AI Conversations. arXiv: 2312.06674 [cs.CL] h...

  58. [64]

    Raisa Islam and Imtiaz Ahmed. 2024. Gemini-the most pow erful LLM: Myth or Truth. In 2024 5th Information Communication Technologies Conference (ICTC). IEEE, Socorro, New Mexico, USA, 303–308. https://doi.org/10.1109/ICTC61510.2024.10602253

  59. [66]

    Balder Janryd and Tim Johansson. 2024. Preventing Heal th Data from Leaking in a Machine Learning System : Implement ing code analysis with LLM and model privacy evaluation testing. , 84 pages

  60. [67]

    Jelinek, R

    F. Jelinek, R. L. Mercer, L. R. Bahl, and J. K. Baker. 1977 . Perplexity—A measure of the difficulty of speech recogni- tion tasks. The Journal of the Acoustical Society of America 62, S1 (1977), S63–S63. https://doi.org/10.1121/1.2016299 arXiv:https://pubs.aip.org/asa/jasa/arti...

  61. [68]

    Joseph, Blaine Nelson, Benjamin I

    Anthony D. Joseph, Blaine Nelson, Benjamin I. P. Rubins tein, and J. D. Tygar. 2019. Adversarial Machine Learning (1st ed.). Cambridge University Press, USA

  62. [69]

    Nikhil Kandpal, Eric Wallace, and Colin Raffel. 2022. De duplicating Training Data Mitigates Privacy Risks in Langu age Models. arXiv:2202.06539 [cs.CR] https://arxiv.org/abs/2202.06539

  63. [70]

    Enkelejda Kasneci, Kathrin Sessler, Stefan Küchemann , Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gass er, Georg Groh, Stephan Günnemann, Eyke Hüllermeier, Stephan Krusche, Gitta Kutyn iok, Tilman Michaeli, Claudia Nerdel, Jürgen Pfeffer, Oleks andra Poquet, Michael S...

  64. [71]

    Kassem, Omar Mahmoud, Niloofar Mireshghallah, H yunwoo Kim, Yulia Tsvetkov, Yejin Choi, Sherif Saad, and San tu Rana

    Aly M. Kassem, Omar Mahmoud, Niloofar Mireshghallah, H yunwoo Kim, Yulia Tsvetkov, Yejin Choi, Sherif Saad, and San tu Rana. 2024. Alpaca against Vicuna: Using LLMs to Uncover Memorization of LLMs. arXiv:2403.04801 [cs.CL] https://arxiv.org/abs/2403.04801

  65. [72]

    John Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu , Khalid Saifullah, Kezhi Kong, Kasun Fernando, Aniruddha S aha, Micah Goldblum, and Tom Goldstein. 2024. On the Reliability of Watermarks for La rge Language Models. arXiv: 2306.04634 [cs.LG] https://arxiv.org/abs/2306.04634

  66. [73]

    Alexey Kurakin, Natalia Ponomareva, Umar Syed, Liam Ma cDermed, and Andreas Terzis. 2024. Harnessing large-langu age models to generate private synthetic text. arXiv: 2306.01684 [cs.LG] https://arxiv.org/abs/2306.01684

  67. [74]

    Minhyeok Lee. 2023. A mathematical investigation of ha llucination and creativity in GPT models. Mathematics 11, 10 (2023), 2320

  68. [75]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio P etroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler , Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. R etrieval-Augmented Generation for Knowledge-Intensive N LP Tasks. In Advances...

  69. [76]

    Haoran Li, Yulin Chen, Jinglong Luo, Yan Kang, Xiaojin Z hang, Qi Hu, Chunkit Chan, and Yangqiu Song. 2023. Privacy in Large Language Models: Attacks, Defenses and Future Directions. arXiv: 2310.10383 [cs.CL] https://arxiv.org/abs/2310.10383

  70. [77]

    Linyang Li, Demin Song, Xiaonan Li, Jiehang Zeng, Ruoti an Ma, and Xipeng Qiu. 2021. Backdoor Attacks on Pre-trained Models by Layerwise Weight Poisoning. arXiv: 2108.13888 [cs.CR] https://arxiv.org/abs/2108.13888

  71. [78]

    Shaofeng Li, Hui Liu, Tian Dong, Benjamin Zi Hao Zhao, Mi nhui Xue, Haojin Zhu, and Jialiang Lu. 2021. Hidden Backdoor s in Human-Centric Language Models. arXiv: 2105.00164 [cs.CL] https://arxiv.org/abs/2105.00164

  72. [79]

    Frank Weizhen Liu and Chenhui Hu. 2024. Exploring Vulne rabilities and Protections in Large Language Models: A Surv ey. arXiv:2406.00240 [cs.LG] https://arxiv.org/abs/2406.00240

  73. [80]

    Xiaogeng Liu, Zhiyuan Yu, Yizhe Zhang, Ning Zhang, and C haowei Xiao. 2024. Automatic and Universal Prompt Injectio n Attacks against Large Language Models. arXiv: 2403.04957 [cs.AI] https://arxiv.org/abs/2403.04957 LLM Security: Vulnerabilities, Attacks, Defenses, and Cou nte...

  74. [81]

    Yi Liu, Gelei Deng, Yuekang Li, Kailong Wang, Zihao Wang , Xiaofeng Wang, Tianwei Zhang, Yepang Liu, Haoyu Wang, Yan Z heng, and Yang Liu. 2024. Prompt Injection attack against LLM-integrated Applications. arXiv: 2306.05499 [cs.CR] https://arxiv.org/abs/2306.05499

  75. [82]

    Yupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia, and Neil Zhenqiang Gong. 2024. Formalizing and Benchmarking Prompt Injection Attacks and Defenses. arXiv: 2310.12815 [cs.CR] https://arxiv.org/abs/2310.12815

  76. [83]

    Yingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee, Ju an Zhai, Weihang Wang, and Xiangyu Zhang. 2018. Trojaning At tack on Neural Networks. In 25th Annual Network and Distributed System Security Sympos ium, NDSS 2018, San Diego, California, USA, February 18-221 , 2018. The I...

  77. [84]

    Ziyao Liu, Huanyi Ye, Chen Chen, Yongsen Zheng, and Kwok -Yan Lam. 2024. Threats, Attacks, and Defenses in Machine Un learning: A Survey. arXiv:2403.13682 [cs.CR] https://arxiv.org/abs/2403.13682

  78. [85]

    Udara Piyasena Liyanage and Nimnaka Dilshan Ranaweera . 2023. Ethical Considerations and Potential Risks in the De ployment of Large Language Models in Diverse Societal Contexts. Journal of Computational Social Dynamics 8, 11 (Nov 2023), 15–25. https://vectoral.org/index.php/J...

  79. [86]

    Mayra Macas, Chunming Wu, and Walter Fuertes. 2024. Adv ersarial examples: A survey of attacks and defenses in deep l earning-enabled cybersecurity systems. Expert Systems with Applications 238 (2024), 122223. https://doi.org/10.1016/j.eswa.2023.122223

  80. [87]

    Nitin Madnani and Bonnie J. Dorr. 2010. Generating Phra sal and Sentential Paraphrases: A Survey of Data- Driven Methods. Computational Linguistics 36, 3 (09 2010), 341–387. https://doi.org/10.1162/coli_a_00002 arXiv:https://direct.mit.edu/coli/article-pdf/36/3/341/1812631/col...

  81. [88]

    John McCarthy. 2004. What is Artificial Intelligence? C omputer Science Department, Stanford University, Stanford, CA 94305. Revised November 12, 2007. URL: http://www-formal.stanford.edu/jmc/

  82. [89]

    Dongyu Meng and Hao Chen. 2017. MagNet: a Two-Pronged De fense against Adversarial Examples. arXiv: 1705.09064 [cs.CR] https://arxiv.org/abs/1705.09064

  83. [90]

    Yash More, Prakhar Ganesh, and Golnoosh Farnadi. 2024. Towards More Realistic Extraction Attacks: An Adversarial Perspective. arXiv:2407.02596 [cs.CR] https://arxiv.org/abs/2407.02596

  84. [91]

    Morris, Volodymyr Kuleshov, Vitaly Shmatikov, and Alexander M

    John X. Morris, Volodymyr Kuleshov, Vitaly Shmatikov, and Alexander M. Rush. 2023. Text Embeddings Reveal (Almost ) As Much As Text. arXiv:2310.06816 [cs.CL] https://arxiv.org/abs/2310.06816

  85. [92]

    Morris, Wenting Zhao, Justin T

    John X. Morris, Wenting Zhao, Justin T. Chiu, Vitaly Shm atikov, and Alexander M. Rush. 2023. Language Model Inversi on. arXiv:2311.13647 [cs.CL] https://arxiv.org/abs/2311.13647

  86. [93]

    Maximilian Mozes, Xuanli He, Bennett Kleinberg, and Le wis D. Griffin. 2023. Use of LLMs for Illicit Purposes: Threats , Prevention Measures, and Vulnerabilities. arXiv: 2308.12833 [cs.CL] https://arxiv.org/abs/2308.12833

  87. [94]

    Feder Cooper, Daphne Ippolito, Christopher A

    Milad Nasr, Nicholas Carlini, Jonathan Hayase, Matthe w Jagielski, A. Feder Cooper, Daphne Ippolito, Christopher A. Choquette-Choo, Eric Wal- lace, Florian Tramèr, and Katherine Lee. 2023. Scalable Extraction of Training Data from (Production) Language Models. arXiv: 2311.1703...

  88. [95]

    Humza Naveed, Asad Ullah Khan, Shi Qiu, Muhammad Saqib, Saeed Anwar, Muhammad Usman, Naveed Akhtar, Nick Barnes, an d Ajmal Mian

  89. [96]

    Daryna Oliynyk, Rudolf Mayer, and Andreas Rauber. 2023 . I Know What You Trained Last Summer: A Survey on Stealing Mac hine Learning Models and Defences. ACM Comput. Surv. 55, 14s, Article 324 (July 2023), 41 pages. https://doi.org/10.1145/3595292

  90. [97]

    OpenAI, :, Aaron Hurst, Adam Lerer, Adam P. Goucher, Ada m Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Wel ihinda, Alan Hayes, Alec Radford, Aleksander Mądry, Alex Baker-Whitcomb, Alex Beutel, Alex Borzunov, Alex Carney, Alex Chow, Alex Kirillo v, Alex Nichol, Alex ...

  91. [98]

    Fábio Perez and Ian Ribeiro. 2022. Ignore Previous Prom pt: Attack Techniques For Language Models. arXiv: 2211.09527 [cs.CL] https://arxiv.org/abs/2211.09527

  92. [99]

    Mansi Phute, Alec Helbling, Matthew Hull, ShengYun Pen g, Sebastian Szyller, Cory Cornelius, and Duen Horng Chau. 2 024. LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked. arXiv :2308.07308 [cs.CL] https://arxiv.org/abs/2308.07308

  93. [100]

    Ivan Provilkov, Dmitrii Emelianenko, and Elena Voita . 2020. BPE-Dropout: Simple and Effective Subword Regularization. In Proceedings of the 58th Annual Meeting of the Association for Computational Lingui stics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetre ault (E...

  94. [101]

    Xiao Pu, Mingqi Gao, and Xiaojun Wan. 2023. Summarizat ion is (Almost) Dead. arXiv: 2309.09558 [cs.CL] https://arxiv.org/abs/2309.09558

  95. [102]

    Fanchao Qi, Yangyi Chen, Xurui Zhang, Mukai Li, Zhiyua n Liu, and Maosong Sun. 2021. Mind the Style of Text! Adversar ial and Backdoor Attacks Based on Text Style Transfer. arXiv: 2110.07139 [cs.CL] https://arxiv.org/abs/2110.07139

  96. [103]

    Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Peter Hend erson, Mengdi Wang, and Prateek Mittal. 2023. Visual Advers arial Examples Jailbreak Aligned Large Language Models. arXiv: 2306.13213 [cs.CR] https://arxiv.org/abs/2306.13213

  97. [104]

    Zhenting Qi, Hanlin Zhang, Eric Xing, Sham Kakade, and Himabindu Lakkaraju. 2024. Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems. arXiv:2402.17840 [cs.CL] https://arxiv.org/abs/2402.17840

  98. [105]

    Baha Rababah, Shang, Wu, Matthew Kwiatkowski, Carson Leung, and Cuneyt Gurcan Akcora. 2024. SoK: Prompt Hacking o f Large Language Models. arXiv: 2410.13901 [cs.CR] https://arxiv.org/abs/2410.13901

  99. [106]

    Alec Radford, Aditya Kwon, Marloes Koot, et al. 2018. Improving Language Understanding by Generative Pre-Traini ng. Technical Report. OpenAI. https://openai.com/research/language-unsupervised

  100. [107]

    Saddam Hossain Mukta , Kaniz Fatema, Nur Mohammad Fahad, Sadman Sakib, Most Maruf atul Jannat Mim, Jubaer Ahmad, Mohammed Eunus Ali, and Sami Azam

    Mohaimenul Azam Khan Raiaan, Md. Saddam Hossain Mukta , Kaniz Fatema, Nur Mohammad Fahad, Sadman Sakib, Most Maruf atul Jannat Mim, Jubaer Ahmad, Mohammed Eunus Ali, and Sami Azam. 2024. A Revi ew on Large Language Models: Architectures, Applications, Taxonomies, Open Issues a...

  101. [108]

    Alexander Robey, Eric Wong, Hamed Hassani, and George J. Pappas. 2024. SmoothLLM: Defending Large Language Model s Against Jailbreaking Attacks. arXiv: 2310.03684 [cs.LG] https://arxiv.org/abs/2310.03684

  102. [109]

    Rudin, Stanley Osher, and Emad Fatemi

    Leonid I. Rudin, Stanley Osher, and Emad Fatemi. 1992. Nonlinear total variation based noise removal algorithms. Physica D: Nonlinear Phenomena 60, 1 (1992), 259–268. https://doi.org/10.1016/0167-2789(92)90242-F

  103. [110]

    Stuart J Russell and Peter Norvig. 2016. Artificial Intelligence: a Modern Approach . Pearson, USA

  104. [111]

    Vinu Sankar Sadasivan, Shoumik Saha, Gaurang Srirama nan, Priyatham Kattakinda, Atoosa Chegini, and Soheil Feiz i. 2024. Fast Adversarial Attacks on Language Models In One GPU Minute. arXiv: 2402.15570 [cs.CR] https://arxiv.org/abs/2402.15570

  105. [112]

    Sander Schulhoff. 2023. Instruction Defense. https://learnprompting.org/docs/prompt_hacking/defensive_measures/instruction Accessed: 2024- 12-04

  106. [113]

    Leo Schwinn, David Dobre, Stephan Günnemann, and Gaut hier Gidel. 2023. Adversarial attacks and defenses in large language models: Old and new threats. arXiv preprint arXiv:2310.19737 239 (2023), 103–117. LLM Security: Vulnerabilities, Attacks, Defenses, and Cou ntermeasures 27

  107. [114]

    Pasi Shailendra, Rudra Chandra Ghosh, Rajdeep Kumar, and Nitin Sharma. 2024. Survey of Large Language Models for A nswering Questions Across Various Fields. , 520-527 pages. https://doi.org/10.1109/ICACCS60874.2024.10717078

  108. [115]

    Erfan Shayegani, Md Abdullah Al Mamun, Yu Fu, Pedram Za ree, Yue Dong, and Nael Abu-Ghazaleh. 2023. Survey of Vulner abilities in Large Language Models Revealed by Adversarial Attacks. arXiv: 2310.10844 [cs.CL] https://arxiv.org/abs/2310.10844

  109. [116]

    Do Anything Now

    Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, an d Yang Zhang. 2024. "Do Anything Now": Characterizing and Ev aluating In-The-Wild Jailbreak Prompts on Large Language Models. arXiv: 2308.03825 [cs.CR] https://arxiv.org/abs/2308.03825

  110. [117]

    Jingzhe Shi, Jialuo Li, Qinwei Ma, Zaiwen Yang, Huan Ma , and Lei Li. 2024. CHOPS: CHat with custOmer Profile Systems f or Customer Service with LLMs. arXiv: 2404.01343 [cs.CL] https://arxiv.org/abs/2404.01343

  111. [118]

    Reza Shokri, Marco Stronati, Congzheng Song, and Vita ly Shmatikov. 2017. Membership Inference Attacks against M achine Learning Models. arXiv:1610.05820 [cs.CR] https://arxiv.org/abs/1610.05820

  112. [119]

    Victoria Smith, Ali Shahin Shamsabadi, Carolyn Ashur st, and Adrian Weller. 2024. Identifying and Mitigating Pri vacy Risks Stemming from Language Models: A Survey. arXiv: 2310.01424 [cs.CL] https://arxiv.org/abs/2310.01424

  113. [120]

    Congzheng Song and Ananth Raghunathan. 2020. Informa tion Leakage in Embedding Models. In Proceedings of the 2020 ACM SIGSAC Conference on Computer and Communications Security (Virtual Event, USA) (CCS ’20). Association for Computing Machinery, New York, NY, USA, 37 7–390. htt...

  114. [121]

    Congzheng Song and Vitaly Shmatikov. 2019. Auditing D ata Provenance in Text-Generation Models. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Min ing (Anchorage, AK, USA) (KDD ’19). Association for Computing Machinery, New York, N...

  115. [122]

    Feifan Song, Bowen Yu, Minghao Li, Haiyang Yu, Fei Huan g, Yongbin Li, and Houfeng Wang. 2024. Preference Ranking Op timization for Human Alignment. arXiv: 2306.17492 [cs.CL] https://arxiv.org/abs/2306.17492

  116. [123]

    Junzhe Song and Dmitry Namiot. 2022. A Survey of Model I nversion Attacks and Countermeasures. Lomonosov Moscow St ate University, GSP-1, Leninskie Gory, Moscow, 119991. URL: https://damdid2022.frccsc.ru/files/article/DAMDID_2022_paper_1040.pdf

  117. [124]

    Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing X ia, Jilong Xue, Jianyong Wang, and Furu Wei. 2023. Retentive Network: A Successor to Transformer for Large Language Models. arXiv: 2307.08621 [cs.CL] https://arxiv.org/abs/2307.08621

  118. [125]

    Xuchen Suo. 2024. Signed-Prompt: A New Approach to Pre vent Prompt Injection Attacks Against LLM-Integrated Appl ications. arXiv:2401.07612 [cs.CR] https://arxiv.org/abs/2401.07612

  119. [126]

    Christian Szegedy, Wojciech Zaremba, Ilya Sutskever , Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. 2014. Intriguing properties of neural networks. arXiv: 1312.6199 [cs.CV] https://arxiv.org/abs/1312.6199

  120. [128]

    Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Ba ptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, A ndrew M. Dai, Anja Hauth, Katie Millican, David Silver, Melvin Johnson, Ioannis Antonoglou, Julian Schrittwieser, Amelia Glaese, Jilin Chen, E mily Pitler, Timothy ...

  121. [129]

    Stephen Burabari Tete. 2024. Threat Modelling and Ris k Analysis for Large Language Model (LLM)-Powered Applicat ions. arXiv:2406.11007 [cs.CR] https://arxiv.org/abs/2406.11007

  122. [130]

    Zhiyi Tian, Lei Cui, Jie Liang, and Shui Yu. 2022. A Comp rehensive Survey on Poisoning Attacks and Countermeasures in Machine Learning. ACM Comput. Surv. 55, 8, Article 166 (Dec. 2022), 35 pages. https://doi.org/10.1145/3551636

  123. [131]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert , Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, ...

  124. [132]

    Reite r, and Thomas Ristenpart

    Florian Tramèr, Fan Zhang, Ari Juels, Michael K. Reite r, and Thomas Ristenpart. 2016. Stealing Machine Learning M od- els via Prediction APIs. In 25th USENIX Security Symposium (USENIX Security 16) . USENIX Association, Austin, TX, 601–618. https://www.usenix.org/conference/u...

  125. [133]

    Alan M. Turing. 1950. Computing Machinery and Intelli gence. Mind 59, 236 (1950), 433–460. http://www.jstor.org/stable/2251299

  126. [134]

    Bibek Upadhayay and Vahid Behzadan. 2024. Sandwich at tack: Multi-language Mixture Adaptive Attack on LLMs. , 208 –226 pages. https://doi.org/10.18653/v1/2024.trustnlp-1.18

  127. [135]

    Neeraj Varshney, Pavel Dolin, Agastya Seth, and Chitt a Baral. 2023. The Art of Defending: A Systematic Evaluation and Analysis of LLM Defense Strategies on Safety and Over-Defensiveness. arXiv: 2401.00287 [cs.CL] https://arxiv.org/abs/2401.00287

  128. [136]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszk oreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention Is All You Need. arXiv: 1706.03762 [cs.CL] https://arxiv.org/abs/1706.03762

  129. [137]

    Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2021. Universal Adversarial Triggers for Attacking and Analyzing NLP. arXiv:1908.07125 [cs.CL] https://arxiv.org/abs/1908.07125

  130. [138]

    Zhao, Shi Feng, and Sameer Singh

    Eric Wallace, Tony Z. Zhao, Shi Feng, and Sameer Singh. 2021. Concealed Data Poisoning Attacks on NLP Models. arXiv :2010.12563 [cs.CL] https://arxiv.org/abs/2010.12563

  131. [139]

    Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein . 2023. Poisoning Language Models During Instruction Tuning. arXiv: 2305.00944 [cs.CL] https://arxiv.org/abs/2305.00944

  132. [140]

    Jiongxiao Wang, Zichen Liu, Keun Hee Park, Zhuojun Jia ng, Zhaoheng Zheng, Zhuofeng Wu, Muhao Chen, and Chaowei Xiao. 2023. Adversarial Demonstration Attacks on Large Language Models. arXiv: 2305.14950 [cs.CL] https://arxiv.org/abs/2305.14950

  133. [141]

    Shang Wang, Tianqing Zhu, Bo Liu, Ming Ding, Xu Guo, Day ong Ye, Wanlei Zhou, and Philip S. Yu. 2024. Unique Security a nd Privacy Threats of Large Language Model: A Comprehensive Survey. arXiv: 2406.07973 [cs.CR] https://arxiv.org/abs/2406.07973

  134. [142]

    Yue Wang, Hung Le, Akhilesh Deepak Gotmare, Nghi D. Q. B ui, Junnan Li, and Steven C. H. Hoi. 2023. CodeT5+: Open Code L arge Language Models for Code Understanding and Generation. arXiv: 2305.07922 [cs.CL] https://arxiv.org/abs/2305.07922

  135. [143]

    Yihan Wang, Zhouxing Shi, Andrew Bai, and Cho-Jui Hsie h. 2024. Defending LLMs against Jailbreaking Attacks via Ba cktranslation. arXiv:2402.16459 [cs.CL] https://arxiv.org/abs/2402.16459

  136. [144]

    Zhibo Wang, Jingjing Ma, Xue Wang, Jiahui Hu, Zhan Qin, and Kui Ren. 2022. Threats to Training: A Survey of Poisoning Attacks and Defenses on Machine Learning Systems. ACM Comput. Surv. 55, 7, Article 134 (Dec. 2022), 36 pages. https://doi.org/10.1145/3538707

  137. [145]

    Zezhong Wang, Fangkai Yang, Lu Wang, Pu Zhao, Hongru Wa ng, Liang Chen, Qingwei Lin, and Kam-Fai Wong. 2024. Self-Gu ard: Empower the LLM to Safeguard Itself. arXiv: 2310.15851 [cs.CL] https://arxiv.org/abs/2310.15851

  138. [146]

    Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023. Jailbroken: How Does LLM Safety Training Fail? arXiv: 2307.02483 [cs.LG] https://arxiv.org/abs/2307.02483

  139. [147]

    Bryce Wong. 2024. Finetuning as a Defense Against LLM Secret-leaking . Master’s thesis. EECS Department, University of Californ ia, Berkeley. http://www2.eecs.berkeley.edu/Pubs/TechRpts/2024/EECS-2024-135.html LLM Security: Vulnerabilities, Attacks, Defenses, and Cou ntermeasures 31

  140. [148]

    Fangzhou Wu, Ning Zhang, Somesh Jha, Patrick McDaniel , and Chaowei Xiao. 2024. A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems. arXiv: 2402.18649 [cs.CR] https://arxiv.org/abs/2402.18649

  141. [149]

    Na ughton

    Xi Wu, Matthew Fredrikson, Somesh Jha, and Jeffrey F. Na ughton. 2016. A Methodology for Formalizing Model-Inversi on Attacks. In 2016 IEEE 29th Computer Security Foundations Symposium (CS F). Institute of Electrical and Electronics Engineers, Lisbo n, Portugal, 355–370. https:...

  142. [150]

    2023, 2024

    xAI. 2023, 2024. Grok: A Next-Generation Large Langua ge Model by xAI. https://x.ai/blog/grok-os. Accessed: 2024-11-10

  143. [151]

    Zhaohan Xi, Tianyu Du, Changjiang Li, Ren Pang, Shouli ng Ji, Jinghui Chen, Fenglong Ma, and Ting Wang. 2023. Defending Pre-trained Language Models as Few-shot Learners against Backdoor Attacks. arXi v:2309.13256 [cs.LG] https://arxiv.org/abs/2309.13256

  144. [152]

    Haoran Xu, Amr Sharaf, Yunmo Chen, Weiting Tan, Lingfe ng Shen, Benjamin Van Durme, Kenton Murray, and Young Jin Kim . 2024. Contrastive Preference Optimization: Pushing the Boundar ies of LLM Performance in Machine Translation. arXiv: 2401.08417 [cs.CL] https://arxiv.org/abs/...

  145. [153]

    Jiashu Xu, Mingyu Derek Ma, Fei Wang, Chaowei Xiao, and Muhao Chen. 2024. Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models. arXiv: 2305.14710 [cs.CL] https://arxiv.org/abs/2305.14710

  146. [154]

    Yudong Xu, Wenhao Li, Pashootan Vaezipoor, Scott Sann er, and Elias B. Khalil. 2024. LLMs and the Abstraction and Re asoning Corpus: Successes, Failures, and the Importance of Object-based Representati ons. arXiv: 2305.18354 [cs.CL] https://arxiv.org/abs/2305.18354

  147. [155]

    Zihao Xu, Yi Liu, Gelei Deng, Yuekang Li, and Stjepan Pi cek. 2024. A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models. arXiv: 2402.13457 [cs.CR] https://arxiv.org/abs/2402.13457

  148. [156]

    Biwei Yan, Kun Li, Minghui Xu, Yueyan Dong, Yue Zhang, Z haochun Ren, and Xiuzhen Cheng. 2024. On Protecting the Data Privacy of Large Language Models (LLMs): A Survey. arXiv: 2403.05156 [cs.CR] https://arxiv.org/abs/2403.05156

  149. [157]

    Haomiao Yang, Mengyu Ge, Dongyun Xue, Kunlan Xiang, Ho ngwei Li, and Rongxing Lu. 2024. Gradient Leakage Attacks in Federated Learning: Research Frontiers, Taxonomy, and Future Directions. IEEE Network 38, 2 (2024), 247–254. https://doi.org/10.1109/MNET.001.2300140

  150. [158]

    Haomiao Yang, Kunlan Xiang, Mengyu Ge, Hongwei Li, Ron gxing Lu, and Shui Yu. 2023. A Comprehensive Overview of Back door Attacks in Large Language Models within Communication Networks. arXi v:2308.14367 [cs.CR] https://arxiv.org/abs/2308.14367

  151. [159]

    Jingfeng Yang, Hongye Jin, Ruixiang Tang, Xiaotian Ha n, Qizhang Feng, Haoming Jiang, Bing Yin, and Xia Hu. 2023. Ha rnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond. arXiv: 2304.13712 [cs.CL] https://arxiv.org/abs/2304.13712

  152. [160]

    Wenkai Yang, Lei Li, Zhiyuan Zhang, Xuancheng Ren, Xu S un, and Bin He. 2021. Be Careful about Poisoned Word Embeddin gs: Exploring the Vulnerability of the Embedding Layers in NLP Models. arXiv: 2103.15543 [cs.CL] https://arxiv.org/abs/2103.15543

  153. [161]

    Wenkai Yang, Yankai Lin, Peng Li, Jie Zhou, and Xu Sun. 2 021. RAP: Robustness-Aware Perturbations for Defending ag ainst Backdoor Attacks on NLP Models. arXiv: 2110.07831 [cs.CL] https://arxiv.org/abs/2110.07831

  154. [162]

    Yifan Yao, Jinhao Duan, Kaidi Xu, Yuanfang Cai, Zhibo S un, and Yue Zhang. 2024. A survey on large language model (LLM ) security and privacy: The Good, The Bad, and The Ugly. High-Confidence Computing 4, 2 (2024), 100211. https://doi.org/10.1016/j.hcc.2024.100211

  155. [163]

    Vasilak os, and Thippa Reddy Gadekallu

    Gokul Yenduri, Ramalingam M, Chemmalar Selvi G, Supri ya Y, Gautam Srivastava, Praveen Kumar Reddy Maddikunta, De epti Raj G, Rutvij H Jhaveri, Prabadevi B, Weizheng Wang, Athanasios V. Vasilak os, and Thippa Reddy Gadekallu. 2023. Generative Pre-train ed Transformer: A Compre...

  156. [164]

    Samuel Yeom, Irene Giacomelli, Matt Fredrikson, and S omesh Jha. 2018. Privacy Risk in Machine Learning: Analyzin g the Connection to Over- fitting. arXiv: 1709.01604 [cs.CR] https://arxiv.org/abs/1709.01604

  157. [165]

    Yifan Zeng, Yiran Wu, Xiao Zhang, Huazheng Wang, and Qi ngyun Wu. 2024. AutoDefense: Multi-Agent LLM Defense again st Jailbreak Attacks. arXiv:2403.04783 [cs.LG] https://arxiv.org/abs/2403.04783

  158. [166]

    Ruisi Zhang, Seira Hidano, and Farinaz Koushanfar. 20 22. Text Revealer: Private Text Reconstruction via Model In version Attacks against Trans- formers. arXiv: 2209.10505 [cs.CL] https://arxiv.org/abs/2209.10505

  159. [167]

    Xinyang Zhang, Zheng Zhang, Shouling Ji, and Ting Wang . 2021. Trojaning Language Models for Fun and Profit. arXiv: 2008.00312 [cs.CR] https://arxiv.org/abs/2008.00312

  160. [168]

    Yiming Zhang, Nicholas Carlini, and Daphne Ippolito. 2024. Effective Prompt Extraction from Language Models. arX iv:2307.06865 [cs.CL] https://arxiv.org/abs/2307.06865

  161. [169]

    Yu Zhang, Xiusi Chen, Bowen Jin, Sheng Wang, Shuiwang J i, Wei Wang, and Jiawei Han. 2024. A Comprehensive Survey of S cientific Large Language Models and Their Applications in Scientific Discov ery. arXiv: 2406.10833 [cs.CL] https://arxiv.org/abs/2406.10833

  162. [170]

    Yuheng Zhang, Ruoxi Jia, Hengzhi Pei, Wenxiao Wang, Bo Li, and Dawn Song. 2020. The Secret Revealer: Generative Mod el-Inversion Attacks Against Deep Neural Networks. arXiv: 1911.07135 [cs.LG] https://arxiv.org/abs/1911.07135

  163. [171]

    Zhiyuan Zhang, Lingjuan Lyu, Xingjun Ma, Chenguang Wa ng, and Xu Sun. 2022. Fine-mixing: Mitigating Backdoors in F ine-tuned Language Models. arXiv: 2210.09545 [cs.CL] https://arxiv.org/abs/2210.09545

  164. [172]

    Ziyang Zhang, Qizhen Zhang, and Jakob Foerster. 2024. PARDEN, Can You Repeat That? Defending against Jailbreaks v ia Repetition. arXiv:2405.07932 [cs.CL] https://arxiv.org/abs/2405.07932

  165. [173]

    Haiyan Zhao, Hanjie Chen, Fan Yang, Ninghao Liu, Huiqi Deng, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, and Mengnan Du. 2024. Explainability for Large Language Models: A Survey. ACM Trans. Intell. Syst. Technol. 15, 2, Article 20 (feb 2024), 38 pages. https://doi.org/10.1145/3639...

  166. [174]

    Shuai Zhao, Jinming Wen, Anh Luu, Junbo Zhao, and Jie Fu . 2023. Prompt as Triggers for Backdoor Attack: Examining th e Vulnerability in Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Na tural Language Processing . Association for Computational ...

  167. [175]

    Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaole i Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhan g, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiy ang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yu n Nie, and Ji-Ro...

  168. [176]

    Ligeng Zhu, Zhijian Liu, and Song Han. 2019. Deep Leaka ge from Gradients. arXiv: 1906.08935 [cs.LG] https://arxiv.org/abs/1906.08935

  169. [177]

    Wenhao Zhu, Hongyi Liu, Qingxiu Dong, Jingjing Xu, Shu jian Huang, Lingpeng Kong, Jiajun Chen, and Lei Li. 2024. Mul tilingual Machine Translation with Large Language Models: Empirical Results and Analysis. arXiv: 2304.04675 [cs.CL] https://arxiv.org/abs/2304.04675

  170. [178]

    Yutao Zhu, Huaying Yuan, Shuting Wang, Jiongnan Liu, W enhan Liu, Chenlong Deng, Haonan Chen, Zhicheng Dou, and Ji- Rong Wen. 2024. Large Language Models for Information Retrieval: A Survey. arXiv:2308.07107 [cs.CL] https://arxiv.org/abs/2308.07107

  171. [179]

    Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B

    Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Br own, Alec Radford, Dario Amodei, Paul Christiano, and Geoffr ey Irving. 2020. Fine- Tuning Language Models from Human Preferences. arXiv: 1909.08593 [cs.CL] https://arxiv.org/abs/1909.08593

  172. [180]

    Zico Kolter, and Matt Fredrikson

    Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson. 2023. Universal and Trans ferable Adversarial Attacks on Aligned Language Models. arXiv: 2307.15043 [cs.CL] https://arxiv.org/abs/2307.15043 Received October 2024

  173. [2023]

    arXiv: 2304.06767 [cs.LG] https://arxiv.org/abs/2304.06767

    RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment. arXiv: 2304.06767 [cs.LG] https://arxiv.org/abs/2304.06767

  174. [2024]

    ar Xiv:2307.06435 [cs.CL] https://arxiv.org/abs/2307.06435

    A Comprehensive Overview of Large Language Models. ar Xiv:2307.06435 [cs.CL] https://arxiv.org/abs/2307.06435

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.