Pith. sign in

REVIEW 10 cited by

Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.17884 v2 pith:REQ33IVC submitted 2023-10-27 cs.AI cs.CLcs.CR

classification cs.AIcs.CLcs.CR
keywords llmsprivacymodelsreasoningworkcontextualcriticaleven
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The interactive use of large language models (LLMs) in AI assistants (at work, home, etc.) introduces a new set of inference-time privacy risks: LLMs are fed different types of information from multiple sources in their inputs and are expected to reason about what to share in their outputs, for what purpose and with whom, within a given context. In this work, we draw attention to the highly critical yet overlooked notion of contextual privacy by proposing ConfAIde, a benchmark designed to identify critical weaknesses in the privacy reasoning capabilities of instruction-tuned LLMs. Our experiments show that even the most capable models such as GPT-4 and ChatGPT reveal private information in contexts that humans would not, 39% and 57% of the time, respectively. This leakage persists even when we employ privacy-inducing prompts or chain-of-thought reasoning. Our work underscores the immediate need to explore novel inference-time privacy-preserving approaches, based on reasoning and theory of mind.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 12 citations worldwide. Full citation record

  1. CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment

    cs.AI 2026-08 conditional novelty 7.0 of 10

    CIDER is a new dataset of individualized disclosure boundaries showing that in-context personalization helps LLMs predict a user's sharing choices, but with uneven false-positive and false-negative improvements across models.

  2. The Impact of Security and Privacy Controls on Users' Emotional Engagement with Generative AI Chatbots

    cs.HC 2026-07 accept novelty 7.0 of 10

    In a vignette study of 354 U.S. participants, deletion-based privacy controls outperformed all other controls in increasing willingness to engage with GenAI chatbots for emotional support, while technically complex co...

  3. PANOPTICON: A PII-Based Assemblage of Naturalistic Output Tokens for Investigating Privacy Leakage Within LLM Context Window

    cs.AI 2026-07 conditional novelty 6.0 of 10

    PANOPTICON is a synthetic benchmark of 67,718 PII-laden prompts for measuring inference-time privacy leakage in LLMs, but its realism and label accuracy are not externally validated.

  4. ENSI: Efficient Non-Interactive Secure Inference for Large Language Models

    cs.CR 2025-09 conditional novelty 6.0 of 10

    ENSI performs secure, non-interactive LLM inference by co-designing CKKS homomorphic encryption with BitNet's ternary weights, achieving up to 8x faster matrix multiplication and 2.6x faster softmax than prior work.

  5. MAGPIE: A dataset for Multi-AGent contextual PrIvacy Evaluation

    cs.AI 2025-06 conditional novelty 6.0 of 10

    MAGPIE is a 158-scenario benchmark showing large language model agents misclassify and leak contextually private information in multi-agent collaboration, even under explicit privacy instructions.

  6. Expert Survey: AI Reliability & Security Research Priorities

    cs.CY 2025-05 conditional novelty 6.0 of 10

    Expert ratings place capability forecasting and dangerous-capability evaluations at the top of a 105-area AI reliability and security research priority list.

  7. CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models

    cs.CL 2025-01 conditional novelty 6.0 of 10

    A new benchmark shows that human safety judgments about LLM responses shift strongly with context, and that current LLMs, especially commercial ones, often fail to match those judgments.

  8. Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A Context Reasoner pipeline that cold-starts LLMs on distilled legal reasoning and applies PPO with a rule-based compliance reward improves performance on CI-based legal compliance benchmarks and transfers to general ...

  9. Robust Privacy: Inference-Stage Privacy through Certified Robustness

    cs.LG 2026-01 reject novelty 4.0 of 10

    Certified output invariance is reframed as inference-stage privacy, expanding the range of sensitive attribute values compatible with a prediction and disrupting label-only model inversion attacks.

  10. A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense

    cs.CR 2024-12 conditional novelty 4.0 of 10

    A multi-stage LLM-based attack/defense dataset pipeline improves reported safety scores of Llama-3.2-1B after SFT, but the evaluation is partly circular and lacks statistical baselines.

Pith tools