REVIEW 2 cited by
Learned Ranking Function: From Short-term Behavior Predictions to Long-term User Satisfaction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present the Learned Ranking Function (LRF), a system that takes short-term user-item behavior predictions as input and outputs a slate of recommendations that directly optimizes for long-term user satisfaction. Most previous work is based on optimizing the hyperparameters of a heuristic function. We propose to model the problem directly as a slate optimization problem with the objective of maximizing long-term user satisfaction. We also develop a novel constraint optimization algorithm that stabilizes objective trade-offs for multi-objective optimization. We evaluate our approach with live experiments and describe its deployment on YouTube.
Forward citations
Cited by 2 Pith papers
-
Self-Evolving Recommendation System: End-To-End Autonomous Model Optimization With LLM Agents
LLM agents acting as ML engineers autonomously generated optimizer, architecture, and reward changes that produced small live metric gains at YouTube when deployed through a dual offline/online loop.
-
ACT: Automated Constraint Targeting for Multi-Objective Recommender Systems
ACT automatically finds minimal hyperparameter adjustments to satisfy recommender guardrails, with a YouTube deployment showing a severely degraded secondary metric restored toward neutral.
Discussion (0). Sign in to comment.