Pith. sign in

REVIEW 4 cited by

LAVA: Data Valuation without Pre-Specified Learning Algorithms

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.00054 v3 pith:QXAT4QGH submitted 2023-04-28 cs.LG cs.AIstat.ML

LAVA: Data Valuation without Pre-Specified Learning Algorithms

classification cs.LG cs.AIstat.ML
keywords datalearningalgorithmdistanceperformancetrainingvalidationvaluation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Traditionally, data valuation (DV) is posed as a problem of equitably splitting the validation performance of a learning algorithm among the training data. As a result, the calculated data values depend on many design choices of the underlying learning algorithm. However, this dependence is undesirable for many DV use cases, such as setting priorities over different data sources in a data acquisition process and informing pricing mechanisms in a data marketplace. In these scenarios, data needs to be valued before the actual analysis and the choice of the learning algorithm is still undetermined then. Another side-effect of the dependence is that to assess the value of individual points, one needs to re-run the learning algorithm with and without a point, which incurs a large computation burden. This work leapfrogs over the current limits of data valuation methods by introducing a new framework that can value training data in a way that is oblivious to the downstream learning algorithm. Our main results are as follows. (1) We develop a proxy for the validation performance associated with a training set based on a non-conventional class-wise Wasserstein distance between training and validation sets. We show that the distance characterizes the upper bound of the validation performance for any given model under certain Lipschitz conditions. (2) We develop a novel method to value individual data based on the sensitivity analysis of the class-wise Wasserstein distance. Importantly, these values can be directly obtained for free from the output of off-the-shelf optimization solvers when computing the distance. (3) We evaluate our new data valuation framework over various use cases related to detecting low-quality data and show that, surprisingly, the learning-agnostic feature of our framework enables a significant improvement over SOTA performance while being orders of magnitude faster.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training

    cs.LG 2026-06 conditional novelty 6.0

    HPO learns language-agent policies by using the Wasserstein distance between the current policy and a hindsight distribution in an intent embedding space, producing low-variance step-level advantages.

  2. Buying Data of Unknown Quality: Fisher Information Procurement Auctions

    cs.GT 2026-04 unverdicted novelty 6.0

    The paper introduces second-score procurement mechanisms for data markets that achieve truthful cost reporting and approximately truthful quality reporting via ex-post statistical verification, with misreporting devia...

  3. Choose Wisely and Privately: Proactive Client Selection for Fair and Efficient Federated Learning

    cs.LG 2026-05 unverdicted novelty 5.0

    Proposes proactive client selection via differentially private mutual information and Potential Federation Loss optimized by simulated annealing to achieve faster, fairer, and more accurate federated models than unifo...

  4. Choose Wisely and Privately: Proactive Client Selection for Fair and Efficient Federated Learning

    cs.LG 2026-05 unverdicted novelty 5.0

    Proactive client selection in federated learning via differentially private mutual information and simulated annealing to optimize Potential Federation Loss for utility and fairness.