Pith. sign in

Scoring Rules and Calibration for Imprecise Probabilities

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

What does it mean to say that, for example, the probability for rain tomorrow is between 20% and 30%? The theory for the evaluation of precise probabilistic forecasts is well-developed and is grounded in the key concepts of proper scoring rules and calibration. For the case of imprecise probabilistic forecasts (sets of probabilities), such theory is still lacking. In this work, we therefore generalize proper scoring rules and calibration to the imprecise case. We develop these concepts as relative to data models and decision problems. As a consequence, the imprecision is embedded in a clear context. We establish a close link to the paradigm of (group) distributional robustness and in doing so provide new insights for it. We argue that proper scoring rules and calibration serve two distinct goals, which are aligned in the precise case, but intriguingly are not necessarily aligned in the imprecise case. The concept of decision-theoretic entropy plays a key role for both goals. Finally, we demonstrate the theoretical insights in machine learning practice, in particular we illustrate subtle pitfalls relating to the choice of loss function in distributional robustness.

citation-role summary

background 1

citation-polarity summary

fields

cs.LG 1

years

2024 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

support 1

representative citing papers

On Calibration in Multi-Distribution Learning

cs.LG · 2024-12-18 · conditional · novelty 4.0

In multi-distribution learning, the minimax-optimal predictor is calibrated only for the worst-case distribution, creating calibration disparities across other distributions.

citing papers explorer

Showing 1 of 1 citing paper.

  • On Calibration in Multi-Distribution Learning cs.LG · 2024-12-18 · conditional · none · ref 23 · internal anchor

    In multi-distribution learning, the minimax-optimal predictor is calibrated only for the worst-case distribution, creating calibration disparities across other distributions.