Pith. sign in

REVIEW 2 cited by

Calibration-compatible Listwise Distillation of Privileged Features for CTR Prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.08727 v1 pith:2SCI5CNP submitted 2023-12-14 cs.IR

classification cs.IR
keywords featuresdistillationmodelprivilegedabilitylosscalibration-compatiblelistwise
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In machine learning systems, privileged features refer to the features that are available during offline training but inaccessible for online serving. Previous studies have recognized the importance of privileged features and explored ways to tackle online-offline discrepancies. A typical practice is privileged features distillation (PFD): train a teacher model using all features (including privileged ones) and then distill the knowledge from the teacher model using a student model (excluding the privileged features), which is then employed for online serving. In practice, the pointwise cross-entropy loss is often adopted for PFD. However, this loss is insufficient to distill the ranking ability for CTR prediction. First, it does not consider the non-i.i.d. characteristic of the data distribution, i.e., other items on the same page significantly impact the click probability of the candidate item. Second, it fails to consider the relative item order ranked by the teacher model's predictions, which is essential to distill the ranking ability. To address these issues, we first extend the pointwise-based PFD to the listwise-based PFD. We then define the calibration-compatible property of distillation loss and show that commonly used listwise losses do not satisfy this property when employed as distillation loss, thus compromising the model's calibration ability, which is another important measure for CTR prediction. To tackle this dilemma, we propose Calibration-compatible LIstwise Distillation (CLID), which employs carefully-designed listwise distillation loss to achieve better ranking ability than the pointwise-based PFD while preserving the model's calibration ability. We theoretically prove it is calibration-compatible. Extensive experiments on public datasets and a production dataset collected from the display advertising system of Alibaba further demonstrate the effectiveness of CLID.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Hybrid Cross-Stage Coordination Pre-ranking Model for Online Recommendation Systems

    cs.IR 2025-02 conditional novelty 6.0 of 10

    A hybrid pre-ranking model that combines ranking-sequence consistency training with margin-based contrastive learning on unexposed items improves recommendation accuracy, especially for long-tail items.

  2. IU4Rec: Interest Unit-Based Product Organization and Recommendation for E-Commerce Platform

    cs.IR 2025-02 conditional novelty 4.0 of 10

    Recommending groups of similar products (interest units) instead of single items improves CTR and transactions on a C2C platform where individual items have limited stock.

Pith tools