Pith. sign in

REVIEW 4 cited by

MetaShift: A Dataset of Datasets for Evaluating Contextual Distribution Shifts and Training Conflicts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.06523 v1 pith:BHAUKIF4 submitted 2022-02-14 cs.LG cs.AIcs.CLcs.CV

classification cs.LGcs.AIcs.CLcs.CV
keywords shiftsdatametashiftacrossdistributionnaturalsetstraining
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Understanding the performance of machine learning models across diverse data distributions is critically important for reliable applications. Motivated by this, there is a growing focus on curating benchmark datasets that capture distribution shifts. While valuable, the existing benchmarks are limited in that many of them only contain a small number of shifts and they lack systematic annotation about what is different across different shifts. We present MetaShift--a collection of 12,868 sets of natural images across 410 classes--to address this challenge. We leverage the natural heterogeneity of Visual Genome and its annotations to construct MetaShift. The key construction idea is to cluster images using its metadata, which provides context for each image (e.g. "cats with cars" or "cats in bathroom") that represent distinct data distributions. MetaShift has two important benefits: first, it contains orders of magnitude more natural data shifts than previously available. Second, it provides explicit explanations of what is unique about each of its data sets and a distance score that measures the amount of distribution shift between any two of its data sets. We demonstrate the utility of MetaShift in benchmarking several recent proposals for training models to be robust to data shifts. We find that the simple empirical risk minimization performs the best when shifts are moderate and no method had a systematic advantage for large shifts. We also show how MetaShift can help to visualize conflicts between data subsets during model training.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 13 citations worldwide. Full citation record

  1. Right Regions, Wrong Labels: Semantic Label Flips in Segmentation under Correlation Shift

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    Under category–scene correlation shift, segmentation models often preserve foreground extent but swap confusable class identities; Flip, FG-Corr/Flip/Miss, and entropy flip-risk make that failure measurable and monitorable.

  2. Tab-MIA: A Benchmark Dataset for Membership Inference Attacks on Tabular Data in LLMs

    cs.CR 2025-07 conditional novelty 6.0 of 10

    Tab-MIA shows LLMs fine-tuned on tabular data are vulnerable to membership inference attacks, with AUROC up to 97.7% after three epochs and encoding format strongly affecting leakage.

  3. Diverse Prototypical Ensembles Improve Robustness to Subpopulation Shift

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A prototype-based ensemble head with explicit diversity regularization improves worst-group accuracy on nine subpopulation shift benchmarks, often matching or beating prior state of the art.

  4. Measuring Time-Series Dataset Similarity using Wasserstein Distance

    cs.LG 2025-07 conditional novelty 4.0 of 10

    Time-series dataset similarity is defined via the Wasserstein distance between fitted multivariate normal distributions, and the distance shows partial correlation with foundation model inference loss.

Pith tools