Pith. sign in

REVIEW 2 cited by

Learned Ranking Function: From Short-term Behavior Predictions to Long-term User Satisfaction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.06512 v1 pith:MYNRJCWH submitted 2024-08-12 cs.LG cs.AIcs.IR

classification cs.LGcs.AIcs.IR
keywords functionlong-termoptimizationsatisfactionuserbehaviordirectlylearned
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present the Learned Ranking Function (LRF), a system that takes short-term user-item behavior predictions as input and outputs a slate of recommendations that directly optimizes for long-term user satisfaction. Most previous work is based on optimizing the hyperparameters of a heuristic function. We propose to model the problem directly as a slate optimization problem with the objective of maximizing long-term user satisfaction. We also develop a novel constraint optimization algorithm that stabilizes objective trade-offs for multi-objective optimization. We evaluate our approach with live experiments and describe its deployment on YouTube.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Self-Evolving Recommendation System: End-To-End Autonomous Model Optimization With LLM Agents

    cs.LG 2026-02 conditional novelty 6.0 of 10

    LLM agents acting as ML engineers autonomously generated optimizer, architecture, and reward changes that produced small live metric gains at YouTube when deployed through a dual offline/online loop.

  2. ACT: Automated Constraint Targeting for Multi-Objective Recommender Systems

    cs.IR 2025-09 conditional novelty 5.0 of 10

    ACT automatically finds minimal hyperparameter adjustments to satisfy recommender guardrails, with a YouTube deployment showing a severely degraded secondary metric restored toward neutral.

Pith tools