Pith. sign in

REVIEW 5 cited by

Distributionally Robust Language Modeling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.02060 v1 pith:RLKKXGDL submitted 2019-09-04 cs.CL cs.LGstat.ML

classification cs.CLcs.LGstat.ML
keywords distributionreviewstestlanguagetrainingapproachdistributionallymixture
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Language models are generally trained on data spanning a wide range of topics (e.g., news, reviews, fiction), but they might be applied to an a priori unknown target distribution (e.g., restaurant reviews). In this paper, we first show that training on text outside the test distribution can degrade test performance when using standard maximum likelihood (MLE) training. To remedy this without the knowledge of the test distribution, we propose an approach which trains a model that performs well over a wide range of potential test distributions. In particular, we derive a new distributionally robust optimization (DRO) procedure which minimizes the loss of the model over the worst-case mixture of topics with sufficient overlap with the training distribution. Our approach, called topic conditional value at risk (topic CVaR), obtains a 5.5 point perplexity reduction over MLE when the language models are trained on a mixture of Yelp reviews and news and tested only on reviews.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unconstrained Robust Online Convex Optimization

    cs.LG 2025-06 conditional novelty 7.0 of 10

    New algorithms achieve regret roughly ||u||G(sqrt(T)+k) under adversarial gradient corruption in unconstrained online convex optimization, with a matching lower bound when G is known.

  2. Multi-Task GRPO: Reliable LLM Reasoning Across Tasks

    cs.CL 2026-02 conditional novelty 6.0 of 10

    MT-GRPO reweights tasks by reward and improvement and enforces those weights after zero-gradient filtering, improving worst-task accuracy by 6–28% over GRPO/DAPO baselines on 3- and 9-task setups.

  3. BLISS: A Lightweight Bilevel Influence Scoring Method for Data Selection in Language Model Pretraining

    cs.LG 2025-10 conditional novelty 6.0 of 10

    A bilevel optimization method ranks pretraining data by training a small proxy model on weighted samples, yielding modest downstream-task gains without external pretrained models.

  4. Learning Credal Ensembles via Distributionally Robust Optimization

    cs.LG 2026-02 conditional novelty 5.0 of 10

    An ensemble trained with distributionally robust optimization at several reweighting intensities yields credal predictions whose uncertainty better separates in-distribution from out-of-distribution samples.

  5. Group Distributionally Robust Machine Learning under Group Level Distributional Uncertainty

    cs.LG 2025-09 reject novelty 5.0 of 10

    A min-max-sup extension of Group DRO that adds a Wasserstein ball around each group's empirical distribution, with a descent-mirror-ascent algorithm and Adult income experiments.

Pith tools