A new probing framework detects moderate parametric memorization signals in tabular in-context learning models under single-task fine-tuning, strongest on low-cardinality tasks, but signals largely disappear under realistic training.
Privacy in large language models: Attacks, defenses and future directions
9 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
Users show limited grasp of GenAI smartphone data practices but heightened privacy concerns across collection, storage, and control, calling for better transparency and user controls.
COMPASS uses semantic clustering on multilingual embeddings to select auxiliary data for PEFT adapters, outperforming linguistic-similarity baselines on multilingual benchmarks while supporting continual adaptation.
Safe-SAIL supplies a pre-explanation metric and segment-level simulation to interpret 1758 safety SAE features across pornography, politics, violence, and terror, with public models and tools released.
Topic modeling of 15,000+ reviews from 59 AI healthcare chatbot apps reveals three main user-reported breakdowns and associates privacy issues with the most negative experiences.
Privacy attacks on LLMs show strong signals for membership inference and backdoors but weaker performance for attribute inference and data extraction, with risks highly dependent on system configuration.
LLM-CEG applies differential privacy during fine-tuning of DistilGPT-2 to reduce membership inference attack success by 71.5% while increasing out-of-distribution utility by 47-50%.
Sentra-Guard reports 99.96% detection of adversarial LLM prompts with AUC 1.00 and ASR of 0.004% using a hybrid SBERT-FAISS and transformer classifier architecture with multilingual translation and human feedback.
citing papers explorer
-
Probing Memorization of Tabular In-Context Learning
A new probing framework detects moderate parametric memorization signals in tabular in-context learning models under single-task fine-tuning, strongest on low-cardinality tasks, but signals largely disappear under realistic training.
-
Understanding User Privacy Perceptions of GenAI Smartphones
Users show limited grasp of GenAI smartphone data practices but heightened privacy concerns across collection, storage, and control, calling for better transparency and user controls.
-
COMPASS: COntinual Multilingual PEFT with Adaptive Semantic Sampling
COMPASS uses semantic clustering on multilingual embeddings to select auxiliary data for PEFT adapters, outperforming linguistic-similarity baselines on multilingual benchmarks while supporting continual adaptation.
-
Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework
Safe-SAIL supplies a pre-explanation metric and segment-level simulation to interpret 1758 safety SAE features across pornography, politics, violence, and terror, with public models and tools released.
-
AI Healthcare Chatbots as Information Infrastructure: A Large-Scale Study of User-Reported Breakdowns
Topic modeling of 15,000+ reviews from 59 AI healthcare chatbot apps reveals three main user-reported breakdowns and associates privacy issues with the most negative experiences.
-
On the Privacy of LLMs: An Ablation Study
Privacy attacks on LLMs show strong signals for membership inference and backdoors but weaker performance for attribute inference and data extraction, with risks highly dependent on system configuration.
-
LLM-CEG: Extending the Classification Error Gauge Framework for Privacy Auditing of Large Language Models
LLM-CEG applies differential privacy during fine-tuning of DistilGPT-2 to reduce membership inference attack success by 71.5% while increasing out-of-distribution utility by 47-50%.
-
Sentra-Guard: A Real-Time Multilingual Defense Against Adversarial LLM Prompts
Sentra-Guard reports 99.96% detection of adversarial LLM prompts with AUC 1.00 and ASR of 0.004% using a hybrid SBERT-FAISS and transformer classifier architecture with multilingual translation and human feedback.
- RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion