Pith. sign in

REVIEW 1 cited by

Joint Training for Selective Prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.24029 v1 pith:G74ZR2EN submitted 2024-10-31 cs.CL cs.LG

classification cs.CLcs.LG
keywords classifiermodelconfidencedeferraljointlearnedperformanceprediction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Classifier models are prevalent in natural language processing (NLP), often with high accuracy. Yet in real world settings, human-in-the-loop systems can foster trust in model outputs and even higher performance. Selective Prediction (SP) methods determine when to adopt a classifier's output versus defer to a human. Previous SP approaches have addressed how to improve softmax as a measure of model confidence, or have developed separate confidence estimators. One previous method involves learning a deferral model based on engineered features. We introduce a novel joint-training approach that simultaneously optimizes learned representations used by the classifier module and a learned deferral policy. Our results on four classification tasks demonstrate that joint training not only leads to better SP outcomes over two strong baselines, but also improves the performance of both modules.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On the Limits of Selective AI Prediction: A Case Study in Clinical Decision Making

    cs.HC 2025-08 conditional novelty 6.0 of 10

    Selective prediction keeps clinicians' overall accuracy roughly intact but shifts errors toward underdiagnosis (18% more missed) and undertreatment (35% more missed) when the AI abstains.

Pith tools