A relevance-guided, fully automatic attack using open-source tools extracts most of a RAG system's private knowledge base without any access to the target's internals.
Adversarial Examples in Modern Machine Learning: A Review
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Recent research has found that many families of machine learning models are vulnerable to adversarial examples: inputs that are specifically designed to cause the target model to produce erroneous outputs. In this survey, we focus on machine learning models in the visual domain, where methods for generating and detecting such examples have been most extensively studied. We explore a variety of adversarial attack methods that apply to image-space content, real world adversarial attacks, adversarial defenses, and the transferability property of adversarial examples. We also discuss strengths and weaknesses of various methods of adversarial attack and defense. Our aim is to provide an extensive coverage of the field, furnishing the reader with an intuitive understanding of the mechanics of adversarial attack and defense mechanisms and enlarging the community of researchers studying this fundamental set of problems.
fields
cs.AI 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Pirates of the RAG: Adaptively Attacking LLMs to Leak Knowledge Bases
A relevance-guided, fully automatic attack using open-source tools extracts most of a RAG system's private knowledge base without any access to the target's internals.