Pith. sign in

REVIEW 2 cited by

Top-K Off-Policy Correction for a REINFORCE Recommender System

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1812.02353 v3 pith:AA6BBL3E submitted 2018-12-06 cs.LG cs.IRstat.ML

classification cs.LGcs.IRstat.ML
keywords recommenderfeedbackbiasescorrectionlearningloggedmultipleoff-policy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Industrial recommender systems deal with extremely large action spaces -- many millions of items to recommend. Moreover, they need to serve billions of users, who are unique at any point in time, making a complex user state space. Luckily, huge quantities of logged implicit feedback (e.g., user clicks, dwell time) are available for learning. Learning from the logged feedback is however subject to biases caused by only observing feedback on recommendations selected by the previous versions of the recommender. In this work, we present a general recipe of addressing such biases in a production top-K recommender system at Youtube, built with a policy-gradient-based algorithm, i.e. REINFORCE. The contributions of the paper are: (1) scaling REINFORCE to a production recommender system with an action space on the orders of millions; (2) applying off-policy correction to address data biases in learning from logged feedback collected from multiple behavior policies; (3) proposing a novel top-K off-policy correction to account for our policy recommending multiple items at a time; (4) showcasing the value of exploration. We demonstrate the efficacy of our approaches through a series of simulations and multiple live experiments on Youtube.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Self-Evolving Recommendation System: End-To-End Autonomous Model Optimization With LLM Agents

    cs.LG 2026-02 conditional novelty 6.0 of 10

    LLM agents acting as ML engineers autonomously generated optimizer, architecture, and reward changes that produced small live metric gains at YouTube when deployed through a dual offline/online loop.

  2. Session-Level Optimization for Large-Scale Retrieval using REINFORCE with Multi-Step Off-Policy Correction

    cs.IR 2026-07 conditional novelty 5.5 of 10

    Off-policy REINFORCE with up to 10 importance-weight factors raises estimated discounted session reward over next-item and positive-only baselines in offline evaluation on the Yambda-5B dataset.

Pith tools