REVIEW 10 cited by
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The interactive use of large language models (LLMs) in AI assistants (at work, home, etc.) introduces a new set of inference-time privacy risks: LLMs are fed different types of information from multiple sources in their inputs and are expected to reason about what to share in their outputs, for what purpose and with whom, within a given context. In this work, we draw attention to the highly critical yet overlooked notion of contextual privacy by proposing ConfAIde, a benchmark designed to identify critical weaknesses in the privacy reasoning capabilities of instruction-tuned LLMs. Our experiments show that even the most capable models such as GPT-4 and ChatGPT reveal private information in contexts that humans would not, 39% and 57% of the time, respectively. This leakage persists even when we employ privacy-inducing prompts or chain-of-thought reasoning. Our work underscores the immediate need to explore novel inference-time privacy-preserving approaches, based on reasoning and theory of mind.
Forward citations
Cited by 10 Pith papers
-
CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment
CIDER is a new dataset of individualized disclosure boundaries showing that in-context personalization helps LLMs predict a user's sharing choices, but with uneven false-positive and false-negative improvements across models.
-
The Impact of Security and Privacy Controls on Users' Emotional Engagement with Generative AI Chatbots
In a vignette study of 354 U.S. participants, deletion-based privacy controls outperformed all other controls in increasing willingness to engage with GenAI chatbots for emotional support, while technically complex co...
-
PANOPTICON: A PII-Based Assemblage of Naturalistic Output Tokens for Investigating Privacy Leakage Within LLM Context Window
PANOPTICON is a synthetic benchmark of 67,718 PII-laden prompts for measuring inference-time privacy leakage in LLMs, but its realism and label accuracy are not externally validated.
-
ENSI: Efficient Non-Interactive Secure Inference for Large Language Models
ENSI performs secure, non-interactive LLM inference by co-designing CKKS homomorphic encryption with BitNet's ternary weights, achieving up to 8x faster matrix multiplication and 2.6x faster softmax than prior work.
-
MAGPIE: A dataset for Multi-AGent contextual PrIvacy Evaluation
MAGPIE is a 158-scenario benchmark showing large language model agents misclassify and leak contextually private information in multi-agent collaboration, even under explicit privacy instructions.
-
Expert Survey: AI Reliability & Security Research Priorities
Expert ratings place capability forecasting and dangerous-capability evaluations at the top of a 105-area AI reliability and security research priority list.
-
CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models
A new benchmark shows that human safety judgments about LLM responses shift strongly with context, and that current LLMs, especially commercial ones, often fail to match those judgments.
-
Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning
A Context Reasoner pipeline that cold-starts LLMs on distilled legal reasoning and applies PPO with a rule-based compliance reward improves performance on CI-based legal compliance benchmarks and transfers to general ...
-
Robust Privacy: Inference-Stage Privacy through Certified Robustness
Certified output invariance is reframed as inference-stage privacy, expanding the range of sensitive attribute values compatible with a prediction and disrupting label-only model inversion attacks.
-
A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense
A multi-stage LLM-based attack/defense dataset pipeline improves reported safety scores of Llama-3.2-1B after SFT, but the evaluation is partly circular and lacks statistical baselines.
Discussion (0). Continue with ORCID to comment.