The first systematization of reconstruction attacks on synthetic tabular data finds that generator choice dominates privacy risk over attack choice, with differential privacy effective only at low budgets and most leakage reflecting population structure rather than memorization.
Title resolution pending
42 Pith papers cite this work, alongside 173 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
representative citing papers
Face-Feature Tuning is a label-free logit remapping method that reduces FPR/TPR gaps across groups in deepfake detection while preserving overall accuracy.
Concatenated MIA score evaluation is uncalibrated across per-sample FPRs and the efficient LiRA has finite population bias; a post-processing calibration is proposed.
Energy shields are adaptive probabilistic controllers using energy functions to ensure runtime fairness with short-term safety and long-term liveness guarantees.
Algorithm uses neural network verification to compute arbitrarily tight bounds on exact SHAP values for neural networks, recovering the exact values and scaling to larger feature spaces than prior exact methods.
FML-Bench shows a simple greedy hill-climber nearly matches tree search on dense-opportunity tasks while an adaptive agent that broadens search on stagnation outperforms six baselines across 18 tasks.
ReMIA offers a practical privacy metric for synthetic data by training two generators and using a classifier to detect source dataset membership, achieving sensitivity comparable to standard MIAs with far less computation.
SOGAR learns Pareto-optimal recourse summaries by solving a bi-objective decision tree optimization that partitions populations and assigns shared low-cost actions per subgroup.
MSP quantifies the minimum changes to analyst choices required to falsify a causal claim by making its confidence interval contain zero, providing information orthogonal to dispersion-based robustness summaries.
Differential subgroups identify specific feature combinations where population differences in outcomes are most extreme, found via a new optimization objective and the DiffSub method.
FairTree audits ML models for subgroup fairness by decomposing performance disparities into systematic bias and variance using permutation-based and fluctuation tests adapted from psychometric methods.
CMRM adds a conformal quantile regularization on prediction margins to any loss, improving noisy-label classification accuracy up to 3.39% across methods and benchmarks while preserving performance at zero noise.
S2-WEF detects dynamic free-riders in federated learning by simulating attack WEF patterns from prior global models, combining them with mutual deviation scores, and using two-dimensional clustering without proxy data or pre-training.
A new framework evaluates privacy metrics for synthetic tabular data by inserting controlled risks and testing detection under no-box threat models on public datasets.
TAB-DRW embeds detectable watermarks in the frequency domain of normalized synthetic tabular data via DFT and rank-based pseudorandom bits, achieving robustness to attacks while preserving fidelity and supporting mixed data types.
LLM-TabLogic extracts inter-column logical constraints using LLMs and conditions a score-based latent diffusion model on them to generate synthetic tabular data that preserves those relationships.
MALLM-GAN uses multi-agent LLMs to emulate GAN architecture for generating higher-quality synthetic tabular data from small samples than prior models, while preserving privacy.
Post-processing fairness methods (ROC and EqOdds) provide the most stable fairness-utility trade-offs when training classifiers on DP synthetic tabular data, outperforming pre- and in-processing interventions.
Multistage defer trees chain sparse decision trees with deferral to match complex ensemble accuracy while routing most samples through one or few interpretable trees.
P²CE is a model-agnostic algorithm for plausible Pareto-optimal counterfactual explanations that uses isolation forest for plausibility and SHAP for efficiency, claiming better quality and speed on three datasets.
Introduces IFSC framework modeling peer imitation in individual fairness-aware strategic classification to improve fairness consistency under interdependent manipulations.
TabChange produces more proximal and valid counterfactuals on tabular data by relationship-based flipping or adversarial latent-space attribute removal compared to baselines on seven datasets.
TDS uses per-tree prediction trajectories to derive instance difficulty scores that rank errors better than prior hardness measures and improve active learning, selective prediction, and Mondrian conformal prediction on tabular data.
PACE-GGM selects poorly approximated covariance entries, measures them privately, and reconstructs the full matrix with a maximum-entropy objective to produce a Gaussian graphical model, yielding lower estimation error than uniform perturbation.
citing papers explorer
-
SoK: Reconstruction Attacks on Synthetic Tabular Data (Insights from Winning the NIST CRC)
The first systematization of reconstruction attacks on synthetic tabular data finds that generator choice dominates privacy risk over attack choice, with differential privacy effective only at low budgets and most leakage reflecting population structure rather than memorization.
-
Toward Calibrated, Fair, and accurate Deepfake Detection
Face-Feature Tuning is a label-free logit remapping method that reduces FPR/TPR gaps across groups in deepfake detection while preserving overall accuracy.
-
On Reliability of Efficient Membership Inference Vulnerability Evaluation
Concatenated MIA score evaluation is uncalibrated across per-sample FPRs and the efficient LiRA has finite population bias; a post-processing calibration is proposed.
-
Energy Shields for Fairness
Energy shields are adaptive probabilistic controllers using energy functions to ensure runtime fairness with short-term safety and long-term liveness guarantees.
-
Verified SHAP: Provable Bounds for Exact Shapley Values of Neural Networks
Algorithm uses neural network verification to compute arbitrarily tight bounds on exact SHAP values for neural networks, recovering the exact values and scaling to larger feature spaces than prior exact methods.
-
FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics
FML-Bench shows a simple greedy hill-climber nearly matches tree search on dense-opportunity tasks while an adaptive agent that broadens search on stagnation outperforms six baselines across 18 tasks.
-
ReMIA: a Powerful and Efficient Alternative to Membership Inference Attacks against Synthetic Data Generators
ReMIA offers a practical privacy metric for synthetic data by training two generators and using a classifier to detect source dataset membership, achieving sensitivity comparable to standard MIAs with far less computation.
-
Optimal Recourse Summaries via Bi-Objective Decision Tree Learning
SOGAR learns Pareto-optimal recourse summaries by solving a bi-objective decision tree optimization that partitions populations and assigns shared low-cost actions per subgroup.
-
Minimum Specification Perturbation: Robustness as Distance-to-Falsification in Causal Inference
MSP quantifies the minimum changes to analyst choices required to falsify a causal claim by making its confidence interval contain zero, providing information orthogonal to dispersion-based robustness summaries.
-
Differential Subgroup Discovery: Characterizing Where Two Populations Differ, and Why
Differential subgroups identify specific feature combinations where population differences in outcomes are most extreme, found via a new optimization objective and the DiffSub method.
-
FairTree: Subgroup Fairness Auditing of Machine Learning Models with Bias-Variance Decomposition
FairTree audits ML models for subgroup fairness by decomposing performance disparities into systematic bias and variance using permutation-based and fluctuation tests adapted from psychometric methods.
-
Conformal Margin Risk Minimization: An Envelope Framework for Robust Learning under Label Noise
CMRM adds a conformal quantile regularization on prediction margins to any loss, improving noisy-label classification accuracy up to 3.39% across methods and benchmarks while preserving performance at zero noise.
-
Dynamic Free-Rider Detection in Federated Learning via Simulated Attack Patterns
S2-WEF detects dynamic free-riders in federated learning by simulating attack WEF patterns from prior global models, combining them with mutual deviation scores, and using two-dimensional clustering without proxy data or pre-training.
-
Empirical Evaluation of Structured Synthetic Data Privacy Metrics: Novel experimental framework
A new framework evaluates privacy metrics for synthetic tabular data by inserting controlled risks and testing detection under no-box threat models on public datasets.
-
Robust Spectral Watermark for Synthetic Tabular Data
TAB-DRW embeds detectable watermarks in the frequency domain of normalized synthetic tabular data via DFT and rank-based pseudorandom bits, achieving robustness to attacks while preserving fidelity and supporting mixed data types.
-
LLM-TabLogic: Preserving Inter-Column Logical Relationships in Synthetic Tabular Data via Prompt-Guided Latent Diffusion
LLM-TabLogic extracts inter-column logical constraints using LLMs and conditions a score-based latent diffusion model on them to generate synthetic tabular data that preserves those relationships.
-
MALLM-GAN: Multi-Agent Large Language Model as Generative Adversarial Network for Synthesizing Tabular Data
MALLM-GAN uses multi-agent LLMs to emulate GAN architecture for generating higher-quality synthetic tabular data from small samples than prior models, while preserving privacy.
-
Where to Intervene? Benchmarking Fairness-Aware Learning on Differentially Private Synthetic Tabular Data
Post-processing fairness methods (ROC and EqOdds) provide the most stable fairness-utility trade-offs when training classifiers on DP synthetic tabular data, outperforming pre- and in-processing interventions.
-
Multistage Defer Trees for Hybrid Interpretability: If at First You Can't Succeed, Tree Again
Multistage defer trees chain sparse decision trees with deferral to match complex ensemble accuracy while routing most samples through one or few interpretable trees.
-
P$^2$CE: Model-Agnostic Plausible Pareto-Optimal Counterfactual Explanations
P²CE is a model-agnostic algorithm for plausible Pareto-optimal counterfactual explanations that uses isolation forest for plausibility and SHAP for efficiency, claiming better quality and speed on three datasets.
-
Beyond Independent Manipulation: Individual Fairness-aware Strategic Classification with Peer Imitation
Introduces IFSC framework modeling peer imitation in individual fairness-aware strategic classification to improve fairness consistency under interdependent manipulations.
-
TabChange: Precise Attribute Changes in Tabular Data
TabChange produces more proximal and valid counterfactuals on tabular data by relationship-based flipping or adversarial latent-space attribute removal compared to baselines on seven datasets.
-
Trajectory-Based Difficulty Scoring for Reliable Learning on Tabular Data
TDS uses per-tree prediction trajectories to derive instance difficulty scores that rank errors better than prior hardness measures and improve active learning, selective prediction, and Mondrian conformal prediction on tabular data.
-
Private Adaptive Covariance Estimation via Gaussian Graphical Models
PACE-GGM selects poorly approximated covariance entries, measures them privately, and reconstructs the full matrix with a maximum-entropy objective to produce a Gaussian graphical model, yielding lower estimation error than uniform perturbation.
-
Provable Fairness Repair for Deep Neural Networks
ProF repairs DNNs for individual fairness by using interval bound propagation to bound outputs over input sets and solving a MILP to adjust the model with guarantees on those sets.
-
R\'enyi Pufferfish Privacy with Gaussian-based Priors: From Single Gaussian to Mixture Model
Gaussian mechanisms for Rényi Pufferfish Privacy under Gaussian and mixture priors deliver exact divergence derivations, closed-form sufficient conditions, and 48.9% less noise than additive baselines on statistical and model queries.
-
TabSCM: A practical Framework for Generating Realistic Tabular Data
TabSCM produces causally consistent tabular data by orienting a CPDAG into a DAG, fitting root marginals with KDE, and using conditional diffusion plus trees for child nodes, outperforming GANs and diffusion baselines on fidelity, utility, and privacy across seven datasets.
-
How to quantify direct correlations between variables
Regularized Jensen-Shannon analogues of KL-based direct-correlation measures are constructed, bounded in [0,1], shown to satisfy the metric property, and their closed-form maxima under fixed alphabet sizes are derived.
-
$\lambda$-GELU: Learning Gating Hardness for Controlled ReLU-ization in Deep Networks
λ-GELU learns layer-wise hardness parameters via constrained reparameterization to allow controlled post-training conversion from smooth GELU to ReLU activations.
-
Logic of Hypotheses: from Zero to Full Knowledge in Neurosymbolic Integration
LoH adds a learnable choice operator to propositional logic, compiles formulas to differentiable graphs via fuzzy logic, subsumes prior NeSy models, and supports discretization to Boolean functions via the Gödel trick.
-
Unleash the Power of Ellipsis: Accuracy-enhanced Sparse Vector Technique with Exponential Noise
New privacy analysis for SVT enables exponential noise plus threshold correction and appending, raising precision and recall up to 50%.
-
GLANCE: Global Actions in a Nutshell for Counterfactual Explainability
GLANCE is a new agglomerative algorithm that jointly clusters in feature and counterfactual-action spaces to produce few, low-cost, high-coverage global recourse actions.
-
Data Collaboration Analysis with Orthonormal Basis Selection and Alignment
Orthonormal Data Collaboration (ODC) enforces orthonormal secret and target bases so that alignment reduces to the Orthogonal Procrustes problem, yielding O(acl^2) complexity, orthogonal concordance, and downstream performance invariant to the choice of target basis.
-
Data-dependent Evaluations for Budgeted Submodular Maximization
New data-dependent upper bounds for budgeted submodular maximization that dominate OPT and empirically tighten optimality certificates on real datasets.
-
Cluster-Specific Localized Drift Detection for Efficient Batch Model Adaptation under Controlled Distribution Shift
A cluster-induced distribution shift simulation framework is proposed and used to evaluate six batch adaptation strategies including cluster-local ADWIN on five benchmark datasets.
-
Interactive Pareto navigation for deep multi-task learning
PPE is a novel predictor-corrector method for interactive Pareto set exploration in deep multi-task learning that approximates tangent spaces via Krylov subspace iterations using only matrix-vector products from automatic differentiation.
-
HE-DAP: Homomorphic Encryption-based Dynamic Adaptive Parameter Optimization for Statistical Computation
HE-DAP adaptively tunes Chebyshev polynomial degree and iteration count for encrypted inverse square root to achieve up to 2.35x speedup across Lattigo, HEaaN-CPU, and HEaaN-GPU while keeping MRE below 3.1e-8.
-
Linear Strategic Classification with Endogenous Improvements
In a linear strategic-classification model where manipulation can genuinely improve outcomes, the optimal strategic classifier is a parallel shift of the Bayes boundary, and it is a provably better proxy for the improvement-aware objective than the Bayes classifier.
-
Learning with Conflicts of Interest
A game-theoretic framework and algorithms are introduced to maximize beneficial information from ML systems while minimizing biased influences arising from conflicts of interest.
-
Cluster Exploration using Informative Manifold Projections
A manifold optimization method combining contrastive PCA and kurtosis projection pursuit generates embeddings that discount prior knowledge structures while revealing underlying cluster separation in high-dimensional data.
-
Data Bias Mitigation under Coverage Constraints & The Price of Fairness
Extends bias mitigation with coverage constraints, casts it as an ILP, and defines the price of fairness as the minimum data modification cost as a function of allowed bias tolerance.
- From Uniform to Learned Knots: A Study of Spline-Based Numerical Encodings for Tabular Deep Learning