Pith. sign in

REVIEW 1 cited by

Distributionally Robust Models with Parametric Likelihood Ratios

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.06340 v1 pith:ZHGVGSYV submitted 2022-04-13 cs.LG

classification cs.LG
keywords modelstrainingdistributionlikelihoodparametricrobustdistributionallydistributions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As machine learning models are deployed ever more broadly, it becomes increasingly important that they are not only able to perform well on their training distribution, but also yield accurate predictions when confronted with distribution shift. The Distributionally Robust Optimization (DRO) framework proposes to address this issue by training models to minimize their expected risk under a collection of distributions, to imitate test-time shifts. This is most commonly achieved by instance-level re-weighting of the training objective to emulate the likelihood ratio with possible test distributions, which allows for estimating their empirical risk via importance sampling (assuming that they are subpopulations of the training distribution). However, re-weighting schemes in the literature are usually limited due to the difficulty of keeping the optimization problem tractable and the complexity of enforcing normalization constraints. In this paper, we show that three simple ideas -- mini-batch level normalization, a KL penalty and simultaneous gradient updates -- allow us to train models with DRO using a broader class of parametric likelihood ratios. In a series of experiments on both image and text classification benchmarks, we find that models trained with the resulting parametric adversaries are consistently more robust to subpopulation shifts when compared to other DRO approaches, and that the method performs reliably well with little hyper-parameter tuning. Code to reproduce our experiments can be found at https://github.com/pmichel31415/P-DRO.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Decision Making under the Exponential Family: Distributionally Robust Optimisation with Bayesian Ambiguity Sets

    cs.LG 2024-11 conditional novelty 6.0 of 10

    A Bayesian posterior is used to center KL-divergence ambiguity sets for distributionally robust optimization, and for conjugate exponential families the worst-case problem reduces to a single-stage stochastic program.

Pith tools