Pith. sign in

REVIEW 1 cited by

Multi-Objective Recommendation via Multivariate Policy Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.02141 v2 pith:Y3B2MANT submitted 2024-05-03 cs.IR cs.LG

classification cs.IRcs.LG
keywords rewardpolicysignalslearninglowermaximisemethodsmultivariate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Real-world recommender systems often need to balance multiple objectives when deciding which recommendations to present to users. These include behavioural signals (e.g. clicks, shares, dwell time), as well as broader objectives (e.g. diversity, fairness). Scalarisation methods are commonly used to handle this balancing task, where a weighted average of per-objective reward signals determines the final score used for ranking. Naturally, how these weights are computed exactly, is key to success for any online platform. We frame this as a decision-making task, where the scalarisation weights are actions taken to maximise an overall North Star reward (e.g. long-term user retention or growth). We extend existing policy learning methods to the continuous multivariate action domain, proposing to maximise a pessimistic lower bound on the North Star reward that the learnt policy will yield. Typical lower bounds based on normal approximations suffer from insufficient coverage, and we propose an efficient and effective policy-dependent correction for this. We provide guidance to design stochastic data collection policies, as well as highly sensitive reward signals. Empirical observations from simulations, offline and online experiments highlight the efficacy of our deployed approach.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ACT: Automated Constraint Targeting for Multi-Objective Recommender Systems

    cs.IR 2025-09 conditional novelty 5.0 of 10

    ACT automatically finds minimal hyperparameter adjustments to satisfy recommender guardrails, with a YouTube deployment showing a severely degraded secondary metric restored toward neutral.

Pith tools