REVIEW 5 cited by
Model Patching: Closing the Subgroup Performance Gap with Data Augmentation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Classifiers in machine learning are often brittle when deployed. Particularly concerning are models with inconsistent performance on specific subgroups of a class, e.g., exhibiting disparities in skin cancer classification in the presence or absence of a spurious bandage. To mitigate these performance differences, we introduce model patching, a two-stage framework for improving robustness that encourages the model to be invariant to subgroup differences, and focus on class information shared by subgroups. Model patching first models subgroup features within a class and learns semantic transformations between them, and then trains a classifier with data augmentations that deliberately manipulate subgroup features. We instantiate model patching with CAMEL, which (1) uses a CycleGAN to learn the intra-class, inter-subgroup augmentations, and (2) balances subgroup performance using a theoretically-motivated subgroup consistency regularizer, accompanied by a new robust objective. We demonstrate CAMEL's effectiveness on 3 benchmark datasets, with reductions in robust error of up to 33% relative to the best baseline. Lastly, CAMEL successfully patches a model that fails due to spurious features on a real-world skin cancer dataset.
Forward citations
Cited by 5 Pith papers
-
Improving Group Robustness on Spurious Correlation via Evidential Alignment
Evidential Alignment improves worst-group accuracy by upweighting a biased model's high-uncertainty errors and retraining the last layer with a calibration set, without group annotations.
-
Fundamental Challenges in Evaluating Text2SQL Solutions and Detecting Their Limitations
Aggregate Text2SQL benchmark numbers are distorted by ambiguous single labels and by the SQL-equivalence match functions, a problem the paper organizes into a taxonomy with concrete Spider examples.
-
Debiasing Classifiers by Amplifying Bias with Latent Diffusion and Large Language Models
A training-free pipeline that uses captions of high-loss images to drive a latent diffusion model to synthesize bias-conflict samples, improving classifier debiasing.
-
Elastic Representation: Mitigating Spurious Correlations for Group Robustness
Elastic Representation regularizes the last-layer representation with nuclear and Frobenius norms, improving worst-group accuracy on CelebA, Waterbirds, and CivilComments without needing group labels.
-
FairDropout: Using Example-Tied Dropout to Enhance Generalization of Minority Groups
An example-tied dropout layer that drops per-example memorizing neurons at inference improves worst-group accuracy across five spurious-correlation benchmarks.
Discussion (0). Continue with ORCID to comment.