Pith. sign in

REVIEW 4 cited by

Omnipredictors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.05389 v1 pith:M36VHTVQ submitted 2021-09-11 cs.LG stat.ML

classification cs.LGstat.ML
keywords lossfunctionlearningmathcaldifferentminimizationomnipredictorspredictor
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Loss minimization is a dominant paradigm in machine learning, where a predictor is trained to minimize some loss function that depends on an uncertain event (e.g., "will it rain tomorrow?''). Different loss functions imply different learning algorithms and, at times, very different predictors. While widespread and appealing, a clear drawback of this approach is that the loss function may not be known at the time of learning, requiring the algorithm to use a best-guess loss function. We suggest a rigorous new paradigm for loss minimization in machine learning where the loss function can be ignored at the time of learning and only be taken into account when deciding an action. We introduce the notion of an (${\mathcal{L}},\mathcal{C}$)-omnipredictor, which could be used to optimize any loss in a family ${\mathcal{L}}$. Once the loss function is set, the outputs of the predictor can be post-processed (a simple univariate data-independent transformation of individual predictions) to do well compared with any hypothesis from the class $\mathcal{C}$. The post processing is essentially what one would perform if the outputs of the predictor were true probabilities of the uncertain events. In a sense, omnipredictors extract all the predictive power from the class $\mathcal{C}$, irrespective of the loss function in $\mathcal{L}$. We show that such "loss-oblivious'' learning is feasible through a connection to multicalibration, a notion introduced in the context of algorithmic fairness. In addition, we show how multicalibration can be viewed as a solution concept for agnostic boosting, shedding new light on past results. Finally, we transfer our insights back to the context of algorithmic fairness by providing omnipredictors for multi-group loss minimization.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Representative Language Generation

    cs.CL 2025-05 conditional novelty 7.0 of 10

    A new 'representative generation' requirement is formalized, characterized by a group closure dimension, with a feasibility result under finite support and a membership-query impossibility.

  2. Discretization-free Multicalibration through Loss Minimization over Tree Ensembles

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A one-shot ERM over depth-two tree ensembles on the base predictor and group indicators yields multicalibration whenever squared loss is saturated, a condition verified empirically on six datasets.

  3. Persuasive Prediction via Decision Calibration

    cs.GT 2025-05 reject novelty 6.0 of 10

    A data-driven sender can learn a near-optimal decision-calibrated predictor without knowing the prior, but the proof as written has a critical Lagrangian error and the Bayesian benchmark is restricted by construction.

  4. Calibration through the Lens of Indistinguishability

    cs.LG 2025-09 accept novelty 2.0 of 10

    A survey arguing that calibration error is best understood as the degree to which two worlds, the predictor's and nature's, can be distinguished, and that this view unifies ECE, smooth calibration, CDL, and distance t...

Pith tools