REVIEW 3 cited by
Threshold KNN-Shapley: A Linear-Time and Privacy-Friendly Approach to Data Valuation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Threshold KNN-Shapley: A Linear-Time and Privacy-Friendly Approach to Data Valuation
read the original abstract
Data valuation aims to quantify the usefulness of individual data sources in training machine learning (ML) models, and is a critical aspect of data-centric ML research. However, data valuation faces significant yet frequently overlooked privacy challenges despite its importance. This paper studies these challenges with a focus on KNN-Shapley, one of the most practical data valuation methods nowadays. We first emphasize the inherent privacy risks of KNN-Shapley, and demonstrate the significant technical difficulties in adapting KNN-Shapley to accommodate differential privacy (DP). To overcome these challenges, we introduce TKNN-Shapley, a refined variant of KNN-Shapley that is privacy-friendly, allowing for straightforward modifications to incorporate DP guarantee (DP-TKNN-Shapley). We show that DP-TKNN-Shapley has several advantages and offers a superior privacy-utility tradeoff compared to naively privatized KNN-Shapley in discerning data quality. Moreover, even non-private TKNN-Shapley achieves comparable performance as KNN-Shapley. Overall, our findings suggest that TKNN-Shapley is a promising alternative to KNN-Shapley, particularly for real-world applications involving sensitive data.
Forward citations
Cited by 3 Pith papers
-
Is Data Shapley Not Better than Random in Data Selection? Ask NASH
NASH improves Shapley-based data selection by decomposing the utility function into Shapley-informative components and aggregating them non-linearly.
-
Is Data Shapley Not Better than Random in Data Selection? Ask NASH
NASH decomposes the validation utility into Shapley-informative component functions and aggregates them non-linearly to make Data Shapley-based data selection consistently effective.
-
Local Shapley: Model-Induced Locality and Optimal Reuse in Data Valuation
Local Shapley restricts data valuation to per-test support sets and reuses subset trainings, but the claimed exactness and concentration bounds are flawed.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.