Pith. sign in

REVIEW 4 cited by

Assessing Prompt Injection Risks in 200+ Custom GPTs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.11538 v2 pith:FLEWLMDL submitted 2023-11-20 cs.CR cs.AI

classification cs.CRcs.AI
keywords promptinjectionmodelssecurityattackschatgptcustomizationgpts
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the rapidly evolving landscape of artificial intelligence, ChatGPT has been widely used in various applications. The new feature - customization of ChatGPT models by users to cater to specific needs has opened new frontiers in AI utility. However, this study reveals a significant security vulnerability inherent in these user-customized GPTs: prompt injection attacks. Through comprehensive testing of over 200 user-designed GPT models via adversarial prompts, we demonstrate that these systems are susceptible to prompt injections. Through prompt injection, an adversary can not only extract the customized system prompts but also access the uploaded files. This paper provides a first-hand analysis of the prompt injection, alongside the evaluation of the possible mitigation of such attacks. Our findings underscore the urgent need for robust security frameworks in the design and deployment of customizable GPT models. The intent of this paper is to raise awareness and prompt action in the AI community, ensuring that the benefits of GPT customization do not come at the cost of compromised security and privacy.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Framework for Auditing Chatbots for Dialect-Based Quality-of-Service Harms

    cs.CY 2025-06 conditional novelty 6.0 of 10

    A query-only audit framework using dynamically generated dialect prompts finds that Amazon Rufus gives lower-quality and more incorrect responses to minoritized English dialects, with typos making the gap worse.

  2. Privacy and Security Threat for OpenAI GPTs

    cs.CR 2025-06 conditional novelty 6.0 of 10

    A large-scale study finds that over 98.8% of sampled OpenAI custom GPTs leak their system instructions to crafted adversarial prompts, and hundreds of GPTs transmit user conversation data to third parties.

  3. When GPT Spills the Tea: Comprehensive Assessment of Knowledge File Leakage in GPTs

    cs.CR 2025-05 conditional novelty 6.0 of 10

    A measurement of 651,022 GPTs identifies five knowledge-file leakage vectors, and the Code Interpreter tool enables direct download of original files in 95.95% of tested GPTs that enable it.

  4. System Prompt Extraction Attacks and Defenses in Large Language Models

    cs.CR 2025-05 conditional novelty 4.0 of 10

    A benchmarking study shows that chain-of-thought, few-shot, and modified sandwich queries can recover LLM system prompts with high similarity-based success, and output filtering is the most reliable tested defense.

Pith tools