Pith. sign in

REVIEW 1 cited by

Probabilistic Rank and Reward: A Scalable Model for Slate Recommendation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.06263 v3 pith:7T5L3QZX submitted 2022-08-10 cs.IR cs.LGstat.ML

Probabilistic Rank and Reward: A Scalable Model for Slate Recommendation

classification cs.IR cs.LGstat.ML
keywords slaterewardprobabilisticrankscalableallowsitemmodel
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We introduce Probabilistic Rank and Reward (PRR), a scalable probabilistic model for personalized slate recommendation. Our approach allows off-policy estimation of the reward in the scenario where the user interacts with at most one item from a slate of K items. We show that the probability of a slate being successful can be learned efficiently by combining the reward, whether the user successfully interacted with the slate, and the rank, the item that was selected within the slate. PRR outperforms existing off-policy reward optimizing methods and is far more scalable to large action spaces. Moreover, PRR allows fast delivery of recommendations powered by maximum inner product search (MIPS), making it suitable in low latency domains such as computational advertising.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Off-Policy Learning in Large Action Spaces: Optimization Matters More Than Estimation

    stat.ML 2025-09 conditional novelty 4.0

    Reward-weighted log-likelihood objectives outperform complex off-policy estimators in large action spaces because their optimization landscapes are much easier to navigate.