Hidden rules injected into a file, a webpage, or a custom GPT's system prompt can make ChatGPT produce biased recommendations and judgments that follow the attacker's instructions.
Proadvprompter: A two-stage journey to effective adversarial prompting for llms
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Breaking the Prompt Wall (I): A Real-World Case Study of Attacking ChatGPT via Lightweight Prompt Injection
Hidden rules injected into a file, a webpage, or a custom GPT's system prompt can make ChatGPT produce biased recommendations and judgments that follow the attacker's instructions.