Pith. sign in

REVIEW 8 cited by

Identifying and Mitigating Privacy Risks Stemming from Language Models: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.01424 v2 pith:7MN3OT4N submitted 2023-09-27 cs.CL cs.AI

classification cs.CLcs.AI
keywords llmsdatamodelsprivacyattackstrainingdimensionsexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) have shown greatly enhanced performance in recent years, attributed to increased size and extensive training data. This advancement has led to widespread interest and adoption across industries and the public. However, training data memorization in Machine Learning models scales with model size, particularly concerning for LLMs. Memorized text sequences have the potential to be directly leaked from LLMs, posing a serious threat to data privacy. Various techniques have been developed to attack LLMs and extract their training data. As these models continue to grow, this issue becomes increasingly critical. To help researchers and policymakers understand the state of knowledge around privacy attacks and mitigations, including where more work is needed, we present the first SoK on data privacy for LLMs. We (i) identify a taxonomy of salient dimensions where attacks differ on LLMs, (ii) systematize existing attacks, using our taxonomy of dimensions to highlight key trends, (iii) survey existing mitigation strategies, highlighting their strengths and limitations, and (iv) identify key gaps, demonstrating open problems and areas for concern.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems

    cs.CR 2025-09 conditional novelty 6.0 of 10

    A systematic review that categorizes LLM threats, severity scores, and mitigations across development and operation life cycles and multiple deployment scenarios.

  2. Look Twice before You Leap: A Rational Framework for Localized Adversarial Anonymization

    cs.CR 2025-12 unverdicted novelty 5.0 of 10

    RLAA is a localized adversarial anonymization framework that adds an arbitrator to filter ghost leaks and enforce rational early stopping, yielding superior privacy-utility trade-offs on benchmarks compared to greedy ...

  3. DevLicOps: A Framework for Mitigating Licensing Risks in AI-Generated Code

    cs.SE 2025-08 conditional novelty 4.0 of 10

    DevLicOps integrates license-compliance controls into the SDLC to reduce risk from AI-generated code, using policies, automated scans, manual audits, and indemnity-aware practices.

  4. Assessing Privacy Preservation and Utility in Online Vision-Language Models

    cs.CV 2026-04 unverdicted novelty 3.0 of 10

    The work proposes and evaluates techniques to reduce PII exposure from image context in online vision-language models while preserving utility for downstream applications.

  5. LLM Harms: A Taxonomy and Discussion

    cs.CY 2025-12 unverdicted novelty 3.0 of 10

    This paper proposes a taxonomy of LLM harms in five categories and suggests mitigation strategies plus a dynamic auditing system for responsible development.

  6. LLM Harms: A Taxonomy and Discussion

    cs.CY 2025-12 reject novelty 3.0 of 10

    Proposes a five-bucket taxonomy of LLM harms and calls for dynamic auditing, but the systematic review behind it is not reproducible and contains mismatched citations.

  7. Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs

    cs.CL 2025-09 conditional novelty 3.0 of 10

    A survey categorizing prompt-based attacks on LLMs into four classes and proposing aspirational goals of un-distillable, un-finetunable, and un-editable models.

  8. A Survey on the Memory Mechanism of Large Language Model based Agents

    cs.AI 2024-04 accept novelty 3.0 of 10

    A systematic review of memory designs, evaluation methods, applications, limitations, and future directions for LLM-based agents.

Pith tools