Pith. sign in

REVIEW 2 cited by

Privacy Preserving Large Language Models: ChatGPT Case Study Based Vision and Framework

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.12523 v1 pith:2KIAGT7P submitted 2023-10-19 cs.CR

classification cs.CR
keywords privacyprivatellmsmodeltrainingdifferentialinformationpreserving
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The generative Artificial Intelligence (AI) tools based on Large Language Models (LLMs) use billions of parameters to extensively analyse large datasets and extract critical private information such as, context, specific details, identifying information etc. This have raised serious threats to user privacy and reluctance to use such tools. This article proposes the conceptual model called PrivChatGPT, a privacy-preserving model for LLMs that consists of two main components i.e., preserving user privacy during the data curation/pre-processing together with preserving private context and the private training process for large-scale data. To demonstrate its applicability, we show how a private mechanism could be integrated into the existing model for training LLMs to protect user privacy; specifically, we employed differential privacy and private training using Reinforcement Learning (RL). We measure the privacy loss and evaluate the measure of uncertainty or randomness once differential privacy is applied. It further recursively evaluates the level of privacy guarantees and the measure of uncertainty of public database and resources, during each update when new information is added for training purposes. To critically evaluate the use of differential privacy for private LLMs, we hypothetically compared other mechanisms e..g, Blockchain, private information retrieval, randomisation, for various performance measures such as the model performance and accuracy, computational complexity, privacy vs. utility etc. We conclude that differential privacy, randomisation, and obfuscation can impact utility and performance of trained models, conversely, the use of ToR, Blockchain, and PIR may introduce additional computational complexity and high training latency. We believe that the proposed model could be used as a benchmark for proposing privacy preserving LLMs for generative AI tools.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Blockchain Meets LLMs: A Living Survey on Bidirectional Integration

    cs.CR 2024-11 reject novelty 1.0 of 10

    A literature survey on bidirectional integration of blockchain and large language models, presenting no new experimental or theoretical results.

  2. Privacy-Preserving Large Language Models: Mechanisms, Applications, and Future Directions

    cs.CR 2024-12 conditional

    A high-level survey of privacy-preserving mechanisms for LLMs, with no new technical contributions.

Pith tools