Pith. sign in

REVIEW 3 cited by

Reconciling Model Multiplicity for Downstream Decision Making

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.19667 v1 pith:BX5HIDIR submitted 2024-05-30 cs.LG cs.AI

classification cs.LGcs.AI
keywords downstreammodelspredictiveprobabilitybest-responsedecision-makingdistributionindividual
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We consider the problem of model multiplicity in downstream decision-making, a setting where two predictive models of equivalent accuracy cannot agree on the best-response action for a downstream loss function. We show that even when the two predictive models approximately agree on their individual predictions almost everywhere, it is still possible for their induced best-response actions to differ on a substantial portion of the population. We address this issue by proposing a framework that calibrates the predictive models with regard to both the downstream decision-making problem and the individual probability prediction. Specifically, leveraging tools from multi-calibration, we provide an algorithm that, at each time-step, first reconciles the differences in individual probability prediction, then calibrates the updated models such that they are indistinguishable from the true probability distribution to the decision-maker. We extend our results to the setting where one does not have direct access to the true probability distribution and instead relies on a set of i.i.d data to be the empirical distribution. Finally, we provide a set of experiments to empirically evaluate our methods: compared to existing work, our proposed algorithm creates a pair of predictive models with both improved downstream decision-making losses and agrees on their best-response actions almost everywhere.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Tractable Agreement Protocols

    cs.LG 2024-11 conditional novelty 8.0 of 10

    Conversation-calibrated agents, efficiently constructible from any ML model, reach approximate agreement in few rounds while improving accuracy, generalizing Aumann-Aaronson theorems to d dimensions and action feedback.

  2. Resolving Predictive Multiplicity for the Rashomon Set

    cs.LG 2026-01 unverdicted novelty 6.0 of 10

    Post-hoc editing of predictions from the Rashomon set—outlier correction, local patching, and pairwise reconciliation—sharply reduces predictive multiplicity while maintaining accuracy.

  3. Predictive Multiplicity in Survival Models: A Method for Quantifying Model Uncertainty in Predictive Maintenance Applications

    cs.LG 2025-04 conditional novelty 5.0 of 10

    Near-optimal survival models can give conflicting failure-risk estimates for the same equipment, and the proposed ambiguity, discrepancy, and obscurity metrics quantify this on CMAPSS engine data.

Pith tools