ContinuousBench shows non-private synthetic text transfers corpus-specific capabilities while state-of-the-art DP methods fail to do so even at ε=100.
Mixed citations
Title resolution pending
Mixed citation behavior. Most common role is background (43%).
citation-role summary
citation-polarity summary
representative citing papers
Exact composition of mechanisms under multiple simultaneous DP constraints is represented as a mixture of heterogeneous compositions using a structural lemma for binary hypothesis tests.
DoHFuse achieves 88.05% closed-world accuracy on 449 classes and strong open-world detection using a new DoH/3 traffic dataset.
TIGER turns the low-rank attention gradient subspace into a differentiable objective for continuous embedding optimization, improving reconstruction quality and robustness over prior discrete token tests especially under noise or DP.
CheckMIABench converts LLMs with intermediate checkpoints into clean MIA testbeds by using pre- and post-checkpoint training data from the same distribution and evaluates published attacks on Pythia and OLMo models while releasing an open-source library.
Alignment defenses adapted from DPO and GRPO mitigate property inference attacks on LLMs while preserving utility.
Gaussian mechanism is asymptotically optimal for high-dimensional DP additive noise; new Spherical Generalized Gamma family outperforms it and the ℓ2 mechanism in some low-dimensional cases with tight composition.
Fair fine-tuning under Equalized Odds yields a tight bound Adv(A, M_f) ≤ Δ_EO · W on adversarial advantage in distribution inference attacks, with empirical reductions below detection threshold across six datasets.
The paper establishes that the optimal excess risk for ε-unlearning is the usual statistical error plus an unlearning penalty that interpolates between retraining-from-scratch and an exponentially smaller term as ε/d grows, with matching bounds for mean estimation.
Clipped and normalized SGD converge without bias in overparameterized interpolating models under (L0,L1)-smoothness, with improved rates and extensions to heavy-tailed noise and weaker smoothness.
ALU uses public data to suppress unlearning cost quadratically while characterizing distribution mismatch effects, enabling mass unlearning with maintained utility.
PACZero achieves zero mutual information privacy in LLM fine-tuning via sign-quantized subset-aggregated ZO gradients, delivering near non-private accuracy on SST-2 at I=0.
Derives tight closed-form bounds on f-DP trade-off functions for DP-SGD with random shuffling subsampling in the high-noise regime σ ≥ √(3/ln M), plus an asymptotic convergence result to the random guessing diagonal with O(√E) δ dependence.
DP-OPD achieves lower perplexity than DP fine-tuning and synthesis-based private distillation under ε=2.0 by enforcing DP-SGD solely on the student during on-policy training with a frozen teacher.
SynBench benchmarks DP text generators across nine datasets and uses a new MIA to show that public pre-training on portions of private data overestimates synthetic text quality and breaks DP privacy bounds.
DPQuant uses epoch-wise probabilistic layer rotation and DP loss sensitivity to quantize only a changing subset of layers, reducing accuracy degradation from quantization noise in DP-SGD and delivering up to 2.21x throughput gains with under 2% accuracy drop.
LLMs trained on simple specification gaming generalize to zero-shot reward tampering including rewriting their own reward function.
Tabular encoder choice reorders multimodal rankings, can erase apparent fusion gains, and requires non-vanilla extraction for in-context learning models to avoid train-test representation shift.
Post-processing fairness methods (ROC and EqOdds) provide the most stable fairness-utility trade-offs when training classifiers on DP synthetic tabular data, outperforming pre- and in-processing interventions.
TL++ recovers centralized mini-batch gradients via virtual batches in split learning and adds secret sharing for cut-layer tensors, achieving 91.41% accuracy on CIFAR-10 with 13x lower communication than full-model sync.
FLKit is a new toolkit with four lifecycle stages, eleven role-specific entry points, a glossary, FL Story template, and tool directory to support federated learning projects in health and life sciences.
Cond-DP conditions DPSGD on public features with decaying spectra to achieve faster convergence guarantees and better empirical performance in label-DP regression.
InfoShield uses TimeAwareMINE to minimize mutual information between speech representations and sensitive attributes, cutting gender inference from 92.6% to 55.5% and age inference from 55.7% to 30.3% while dropping depression F1 by only 6%.
Introduces Bayesian Membership Privacy (BMP) as a sampling-aware node-level privacy definition for GNNs quantified by posterior membership probability, plus an auditing method and benchmark experiments.
citing papers explorer
-
ContinuousBench: Can Differentially Private Synthetic Text Improve Capabilities?
ContinuousBench shows non-private synthetic text transfers corpus-specific capabilities while state-of-the-art DP methods fail to do so even at ε=100.
-
Composition Theorems for Multiple Differential Privacy Constraints
Exact composition of mechanisms under multiple simultaneous DP constraints is represented as a mixture of heterogeneous compositions using a structural lemma for binary hypothesis tests.
-
DoHFuse: A Dual-Branch Architecture with DMAGLSTM for Website Fingerprinting over DNS over HTTPS/3
DoHFuse achieves 88.05% closed-world accuracy on 449 classes and strong open-world detection using a new DoH/3 traffic dataset.
-
TIGER: Inverting Transformer Gradients via Embedding-Subspace Distance Optimization
TIGER turns the low-rank attention gradient subspace into a differentiable objective for continuous embedding optimization, improving reconstruction quality and robustness over prior discrete token tests especially under noise or DP.
-
CheckMIABench: Firm Foundations For Membership Inference Attacks on Language Models
CheckMIABench converts LLMs with intermediate checkpoints into clean MIA testbeds by using pre- and post-checkpoint training data from the same distribution and evaluates published attacks on Pythia and OLMo models while releasing an open-source library.
-
Alignment Defends LLMs from Property Inference Attacks
Alignment defenses adapted from DPO and GRPO mitigate property inference attacks on LLMs while preserving utility.
-
Asymptotic Optimality of the High-Dimensional Gaussian Mechanism and Improved Low-Dimensional Mechanisms for Differential Privacy
Gaussian mechanism is asymptotically optimal for high-dimensional DP additive noise; new Spherical Generalized Gamma family outperforms it and the ℓ2 mechanism in some low-dimensional cases with tight composition.
-
Fair Finetuning Mitigates Distribution Inference Attacks
Fair fine-tuning under Equalized Odds yields a tight bound Adv(A, M_f) ≤ Δ_EO · W on adversarial advantage in distribution inference attacks, with empirical reductions below detection threshold across six datasets.
-
Near-Optimal Pure Machine Unlearning for Smooth Strongly Convex Losses
The paper establishes that the optimal excess risk for ε-unlearning is the usual statistical error plus an unlearning penalty that interpolates between retraining-from-scratch and an exponentially smaller term as ε/d grows, with matching bounds for mean estimation.
-
Avoiding Bias in Clipped SGD for Overparameterized Models under Generalized Smoothness
Clipped and normalized SGD converge without bias in overparameterized interpolating models under (L0,L1)-smoothness, with improved rates and extensions to heavy-tailed noise and weaker smoothness.
-
Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data
ALU uses public data to suppress unlearning cost quadratically while characterizing distribution mismatch effects, enabling mass unlearning with maintained utility.
-
PACZero: PAC-Private Fine-Tuning of Language Models via Sign Quantization
PACZero achieves zero mutual information privacy in LLM fine-tuning via sign-quantized subset-aggregated ZO gradients, delivering near non-private accuracy on SST-2 at I=0.
-
Trade-off Functions for DP-SGD with Subsampling based on Random Shuffling: Tight Upper and Lower Bounds
Derives tight closed-form bounds on f-DP trade-off functions for DP-SGD with random shuffling subsampling in the high-noise regime σ ≥ √(3/ln M), plus an asymptotic convergence result to the random guessing diagonal with O(√E) δ dependence.
-
DP-OPD: Differentially Private On-Policy Distillation for Language Models
DP-OPD achieves lower perplexity than DP fine-tuning and synthesis-based private distillation under ε=2.0 by enforcing DP-SGD solely on the student during on-policy training with a frozen teacher.
-
SynBench: A Benchmark for Differentially Private Text Generation
SynBench benchmarks DP text generators across nine datasets and uses a new MIA to show that public pre-training on portions of private data overestimates synthetic text quality and breaks DP privacy bounds.
-
DPQuant: Efficient and Differentially-Private Model Training via Dynamic Quantization Scheduling
DPQuant uses epoch-wise probabilistic layer rotation and DP loss sensitivity to quantize only a changing subset of layers, reducing accuracy degradation from quantization noise in DP-SGD and delivering up to 2.21x throughput gains with under 2% accuracy drop.
-
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
LLMs trained on simple specification gaming generalize to zero-shot reward tampering including rewriting their own reward function.
-
The Importance of Encoder Choice:A Tabular-Image Study
Tabular encoder choice reorders multimodal rankings, can erase apparent fusion gains, and requires non-vanilla extraction for in-context learning models to avoid train-test representation shift.
-
Where to Intervene? Benchmarking Fairness-Aware Learning on Differentially Private Synthetic Tabular Data
Post-processing fairness methods (ROC and EqOdds) provide the most stable fairness-utility trade-offs when training classifiers on DP synthetic tabular data, outperforming pre- and in-processing interventions.
-
TL++: Accuracy and Privacy Preserving Traversal Learning for Distributed Intelligent Systems
TL++ recovers centralized mini-batch gradients via virtual batches in split learning and adds secret sharing for cut-layer tensors, achieving 91.41% accuracy on CIFAR-10 with 13x lower communication than full-model sync.
-
Development and Design of FLKit: A Structured Onboarding Toolkit for Federated Learning in Health and Life Sciences
FLKit is a new toolkit with four lifecycle stages, eleven role-specific entry points, a glossary, FL Story template, and tool directory to support federated learning projects in health and life sciences.
-
Private Learning with Public Feature Conditioning
Cond-DP conditions DPSGD on public features with decaying spectra to achieve faster convergence guarantees and better empirical performance in label-DP regression.
-
InfoShield: Privacy-Preserving Speech Representations for Mental Health Screening via Information-Theoretic Optimization
InfoShield uses TimeAwareMINE to minimize mutual information between speech representations and sensitive attributes, cutting gender inference from 92.6% to 55.5% and age inference from 55.7% to 30.3% while dropping depression F1 by only 6%.
-
Bayesian Membership Privacy for Graph Neural Networks
Introduces Bayesian Membership Privacy (BMP) as a sampling-aware node-level privacy definition for GNNs quantified by posterior membership probability, plus an auditing method and benchmark experiments.
-
An exponential mechanism based on quadratic approximations for fine-tuning machine learning models with privacy guarantees
A differentially private fine-tuning method that constructs a quadratic utility function to allow exact sampling from a multivariate normal distribution while providing theoretical privacy guarantees.
-
Auditing Privacy in Multi-Tenant RAG under Account Collusion
Account collusion in multi-tenant RAG raises retrieval leakage to Θ(√k ε_acc) under Gaussian noise-then-select; the work supplies a matching attack and a verifier-runnable audit protocol for coalitions up to k_max.
-
Keyed Nonlinear Transform: Lightweight Privacy-Enhancing Feature Sharing for Medical Image Analysis
KNT applies key-conditioned nonlinear obfuscation to split-inference features, cutting re-identification AUC from 0.635 to 0.586 with 0.15 ms overhead and under 1 pp accuracy loss.
-
DP-KFC: Data-Free Preconditioning for Privacy-Preserving Deep Learning
DP-KFC approximates the Fisher Information Matrix for KFAC preconditioning via synthetic noise probes and modality frequency statistics, matching private-data performance without consuming privacy budget or introducing distribution shift.
-
On What We Can Learn from Low-Resolution Data
Low-resolution data improves high-resolution model performance when high-resolution samples are limited, via KL-divergence bounds and experiments on vision transformers and CNNs.
-
Deep Learning under Fractional-Order Differential Privacy
FO-DP-SGD adds fractional-order memory to the private gradient release in DP-SGD, achieving better test accuracy on SVHN, CIFAR-10, and CIFAR-100 while using standard Rényi DP accounting with adjusted sensitivity βC.
-
Benchmarking Large Language Models for IoC Recovery under Adversarial Code Obfuscation and Encryption
LLMs recover IoCs from lightweight-obfuscated JavaScript but performance collapses under encryption in a new benchmark of 336 programs across 12 concealment levels.
-
SoK: Practical Aspects of Releasing Differentially Private Graphs
A modular systemisation plus practitioner objectives framework organises DP graph release methods and yields an open social-network benchmark of SotA edge- and node-DP algorithms.
-
Cyclic Adaptive Private Synthesis for Sharing Real-World Data in Education
CAPS provides an iterative differentially private synthesis method that outperforms one-shot baselines on authentic educational real-world data.
-
Crowding Out The Noise: Algorithmic Collective Action Under Differential Privacy
Differential privacy reduces algorithmic collective action effectiveness, with formal lower bounds on success probability depending on collective size and privacy parameters, plus experimental verification on neural nets.
-
Ethical and social risks of harm from Language Models
The authors provide a detailed taxonomy of 21 risks associated with language models, covering discrimination, information leaks, misinformation, malicious applications, interaction harms, and societal impacts like job loss and environmental costs.
-
Private training in quantum machine learning
Deterministic gradient-norm bounds in variational QML control DP-SGD clipping bias, so quantum models retain more accuracy than matched classical models under the same privacy budget.
-
CHRONOS: Temporally-Aware Multi-Agent Coordination for Evolving Data Marketplaces
CHRONOS is a three-layer system for evolving data marketplaces that applies neural-ODE temporal decay, changepoint-aware Shapley valuation, and EXP3-IX private coordination to achieve 0.937 recall, 2.74 qps, 161 ms latency, and epsilon 4.25 at delta 10^-6.
-
Measuring Database Unfairness via Dependency Quantification Under Differential Privacy
Proposes a formal DP-compatible framework with three unfairness measures (mutual information with TV proxy, MaxSAT-based repair, top-k tuple contribution) that satisfy positivity, monotonicity, and computability.
-
UMEDA: Unified Multi-modal Efficient Data Fusion for Privacy-Preserving Graph Federated Learning via Spectral-Gated Attention and Diffusion-Based Operator Alignment
UMEDA is a new graph federated learning method that uses low-rank spectral filtering and diffusion over a shared integral operator to fuse multi-modal data privately, outperforming baselines on MM-Fi and RELI11D under high heterogeneity and tight privacy budgets.
-
Class-Aware Adaptive Differential Privacy in Deep Learning for Sensor-Based Fall Detection
CA-ADP adjusts differential privacy noise per mini-batch class composition to improve F-scores by 3.3-8.5% over standard DP on three fall-detection datasets while claiming formal (ε,δ) guarantees.
-
Protecting and Preserving Protest Dynamics for Responsible Analysis
A responsible computing framework substitutes real protest imagery with labeled synthetic reproductions from conditional image synthesis to enable privacy-aware analysis of collective action patterns.
-
Secure, Verifiable, and Scalable Multi-Client Data Sharing via Consensus-Based Privacy-Preserving Data Distribution
CPPDD is a new consensus-based protocol for privacy-preserving multi-client data sharing that achieves unanimous-release confidentiality, linear scalability, and high-probability malicious deviation detection.
-
On Optimal Hyperparameters for Differentially Private Deep Transfer Learning
Empirical study of DP transfer learning reveals that larger clipping bounds outperform under tight privacy and cumulative DP noise explains batch-size effects better than existing heuristics.
-
LLM-Assisted Web Measurements
LLMs achieve strong performance on website classification tasks relevant to web measurements and support a practical two-step methodology for targeted studies from the Tranco list.
-
Defending Diffusion Models Against Membership Inference Attacks via Higher-Order Langevin Dynamics
Introduces higher-order Langevin dynamics with auxiliary variables as a defense that mixes randomness early to reduce membership inference success on diffusion models, measured via AUROC and FID on toy and speech data.
-
SFBD Flow: A Continuous-Optimization Framework for Training Diffusion Models with Noisy Samples
SFBD Flow converts the iterative SFBD approach into a continuous optimization framework for diffusion models on noisy samples, with its Online SFBD instantiation outperforming baselines.
-
Revisiting Privacy Preservation in Brain-Computer Interfaces: Conceptual Boundaries, Risk Pathways, and a Protection-Strength Grading Framework
The paper synthesizes BCI privacy risks and introduces a three-dimensional framework that grades existing protection methods into four strength levels while flagging mental privacy as an unresolved neuroethical issue.
-
Towards the Anonymization of the Language Modeling
Authors introduce MLM and CLM specialization methods that avoid memorizing identifiers in sensitive training data while aiming for a privacy-utility tradeoff on medical datasets.
-
A Novel Approach for Detection and Ranking of Trendy and Emerging Cyber Threat Events in Twitter Streams
Unsupervised ML approach detects and ranks novel and developing cyber threat events in Twitter by extracting named entities and keywords weighted by user influence, evaluated against human annotators.
- Fundamental Limitations of Favorable Privacy-Utility Guarantees for DP-SGD