Pith. sign in

REVIEW 1 cited by

Dataset Interfaces: Diagnosing Model Failures Using Controllable Counterfactual Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.07865 v2 pith:LMYREBZ6 submitted 2023-02-15 cs.LG

classification cs.LG
keywords datasetdistributionshiftinterfacemodelinputcounterfactualexhibit
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Distribution shift is a major source of failure for machine learning models. However, evaluating model reliability under distribution shift can be challenging, especially since it may be difficult to acquire counterfactual examples that exhibit a specified shift. In this work, we introduce the notion of a dataset interface: a framework that, given an input dataset and a user-specified shift, returns instances from that input distribution that exhibit the desired shift. We study a number of natural implementations for such an interface, and find that they often introduce confounding shifts that complicate model evaluation. Motivated by this, we propose a dataset interface implementation that leverages Textual Inversion to tailor generation to the input distribution. We then demonstrate how applying this dataset interface to the ImageNet dataset enables studying model behavior across a diverse array of distribution shifts, including variations in background, lighting, and attributes of the objects. Code available at https://github.com/MadryLab/dataset-interfaces.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DILLEMA: Diffusion and Large Language Models for Multi-Modal Augmentation

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A pipeline that uses LLM-generated counterfactual captions and control-conditioned diffusion to produce label-preserving synthetic test images that expose vision model weaknesses and improve retraining.

Pith tools