Pith. sign in

REVIEW 2 cited by

Generator and Critic: A Deep Reinforcement Learning Approach for Slate Re-ranking in E-commerce

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.12206 v1 pith:XHP5SFRE submitted 2020-05-25 cs.LG cs.IRstat.ML

classification cs.LGcs.IRstat.ML
keywords slateitemscriticgeneratorapproachlearningreinforcemente-commerce
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The slate re-ranking problem considers the mutual influences between items to improve user satisfaction in e-commerce, compared with the point-wise ranking. Previous works either directly rank items by an end to end model, or rank items by a score function that trades-off the point-wise score and the diversity between items. However, there are two main existing challenges that are not well studied: (1) the evaluation of the slate is hard due to the complex mutual influences between items of one slate; (2) even given the optimal evaluation, searching the optimal slate is challenging as the action space is exponentially large. In this paper, we present a novel Generator and Critic slate re-ranking approach, where the Critic evaluates the slate and the Generator ranks the items by the reinforcement learning approach. We propose a Full Slate Critic (FSC) model that considers the real impressed items and avoids the impressed bias of existing models. For the Generator, to tackle the problem of large action space, we propose a new exploration reinforcement learning algorithm, called PPO-Exploration. Experimental results show that the FSC model significantly outperforms the state of the art slate evaluation methods, and the PPO-Exploration algorithm outperforms the existing reinforcement learning methods substantially. The Generator and Critic approach improves both the slate efficiency(4% gmv and 5% number of orders) and diversity in live experiments on one of the largest e-commerce websites in the world.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. DIRECTOR: Dynamic Index-based Recommendation with Transport-Optimized Retrieval

    cs.IR 2026-07 conditional novelty 6.0 of 10

    A parallel non-autoregressive reranker that trains with capacity-constrained optimal transport and decodes with global hard matching improves slate recommendation quality and serving efficiency.

  2. PSG: Pair-Space Generation for Efficient Generative Reranking

    cs.IR 2026-07 conditional novelty 4.0 of 10

    PSG halves autoregressive decoding steps for list reranking by generating ordered item pairs as single tokens, claiming ~2-4x speedup and ~4x lower worst-case error, with a 1.83x latency win and 0.178% stay-time lift online.

Pith tools