Pith. sign in

REVIEW 1 cited by

Your 2 is My 1, Your 3 is My 9: Handling Arbitrary Miscalibrations in Ratings

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1806.05085 v2 pith:IAHD4J7I submitted 2018-06-13 stat.ML cs.AIcs.ITcs.LGmath.IT

classification stat.MLcs.AIcs.ITcs.LGmath.IT
keywords cardinalmiscalibrationsrankingscoresestimatorsmiscalibrationonlyapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Cardinal scores (numeric ratings) collected from people are well known to suffer from miscalibrations. A popular approach to address this issue is to assume simplistic models of miscalibration (such as linear biases) to de-bias the scores. This approach, however, often fares poorly because people's miscalibrations are typically far more complex and not well understood. In the absence of simplifying assumptions on the miscalibration, it is widely believed by the crowdsourcing community that the only useful information in the cardinal scores is the induced ranking. In this paper, inspired by the framework of Stein's shrinkage, empirical Bayes, and the classic two-envelope problem, we contest this widespread belief. Specifically, we consider cardinal scores with arbitrary (or even adversarially chosen) miscalibrations which are only required to be consistent with the induced ranking. We design estimators which despite making no assumptions on the miscalibration, strictly and uniformly outperform all possible estimators that rely on only the ranking. Our estimators are flexible in that they can be used as a plug-in for a variety of applications, and we provide a proof-of-concept for A/B testing and ranking. Our results thus provide novel insights in the eternal debate between cardinal and ordinal data.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Personalized Recommendations via Active Utility-based Pairwise Sampling

    cs.IR 2025-08 conditional novelty 5.0 of 10

    A utility-based active sampling strategy for pairwise preference learning picks the questions that most improve expected recommendation quality, outperforming random and uncertainty-based baselines in two experiments.

Pith tools