REVIEW 4 major objections 4 minor 65 references
Pair-space generation halves the decoding horizon of autoregressive rerankers without losing expressive power.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
PSG halves autoregressive decoding steps for list reranking by generating ordered item pairs as single tokens, claiming ~2-4x speedup and ~4x lower worst-case error, with a 1.83x latency win and 0.178% stay-time lift online.
T0 review reviewed 2026-08-01 challenge →
load-bearing objection Solid engineering contribution with a real deployment result, but don't quote the 'nearly 4x' theory without reading Appendix C. the 4 major comments →
PSG: Pair-Space Generation for Efficient Generative Reranking
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The central claim is that elevating the generation atom from items to ordered item pairs—a per-request vocabulary of n(n−1) tokens—preserves the full set of realizable list distributions while halving the autoregressive horizon from L to L/2. The bijection between item-space and pair-space sequences and the telescoping product identity show that any item-space policy can be replicated exactly by a pair-space policy. Because reward suboptimality under outcome-only rewards scales quadratically with the horizon, halving the horizon yields a 4x reduction in the worst-case bound, provided the per-step total-variation mismatch of pair generation does not exceed that of item generation. The authors
What carries the argument
The key object is the pair token: an ordered pair (v_i, v_j) of candidate items treated as a single generation symbol. The pair-token representation is produced by a position-aware fusion module with learnable role embeddings, so (i,j) and (j,i) are distinct. The dynamic vocabulary of size n(n−1) is scored via inner product with decoder hidden states; a pretrained pair encoder supplies representations without a static n² embedding table. The theory rests on a bijection between item-space and pair-space sequences and a telescoping product identity that equates the two autoregressive factorizations.
Load-bearing premise
The 4x suboptimality improvement holds only if the per-step generation error of pair tokens is no larger than that of item tokens; the paper estimates this on order-of-magnitude grounds but never measures it.
What would settle it
Measure, on a fixed reranking task, the total-variation distance between the learned policy and the optimal policy at each decoding step for both item-space and pair-space generators. If ε_pair consistently exceeds ε_item, the claimed 4x worst-case improvement shrinks proportionally; if ε_pair exceeds 4×ε_item, the pair-space bound is no better than item-space.
If this is right
- Recommender rerankers can halve decoding latency while retaining the expressiveness of item-level autoregressive models.
- Worst-case suboptimality under outcome-only rewards improves by up to 4x when per-step error does not grow, reducing the impact of error accumulation in long lists.
- The approach generalizes to k-item tokens, but the paper shows k=2 is the industrial sweet spot; k=3 becomes infeasible for n≥60 within their latency budget.
- PSG is orthogonal to existing acceleration techniques (e.g., speculative decoding, KV-cache optimizations) and can be combined with them.
- The online deployment result suggests that pair-level action granularity can translate to measurable user-engagement gains in a production recommender.
Where Pith is reading between the lines
- A direct measurement of per-step total-variation error ε_pair and ε_item on real data would determine whether the 4x dividend survives; the paper's own Remark 1 concedes the unconditional claim is false.
- The one-step look-ahead the authors attribute to pair tokens could be reinterpreted as a form of local planning: the pair token jointly decides two positions, implicitly conditioning the first on the second.
- The method's benefit is likely largest when list utility is dominated by pairwise item interactions (complementarity, diversity); for purely additive utilities, the advantage may shrink—a testable hypothesis the paper motivates but does not quantify.
- Pair-token pretraining from exposure logs could be extended to variable-length n-grams or triples if the reward structure warrants it, though vocabulary growth imposes practical limits.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes Pair-Space Generation (PSG), a generative reranking reformulation in which the decoder emits ordered item pairs instead of individual items. For a candidate set of size n, the per-step vocabulary becomes n(n−1) and the decoding horizon halves from L to L/2; a pretrained pair-token representation module constructs pair embeddings on the fly. The authors claim three guarantees: (1) a bijection between pair-token and item-token sequences, so no expressiveness is lost; (2) a theoretical 2×–4× decoding speedup (1.83× measured in production); (3) an error-compounding bound showing that, under outcome-only rewards, worst-case suboptimality is O((L/2)^2 ε_pair), 'nearly 4×' better than item-space generation when per-step TV mismatches satisfy ε_pair≈ε_item. Training combines pair-token pretraining, next-token prediction, and GRPO. Experiments on ML-1M, Amazon-Books, and RecFlow show consistent gains over strong baselines, and an online A/B test on Kuaishou reports 0.178% stay-time lift and 1.83× decoding speedup. Appendices extend results to odd L and give proofs.
Significance. The paper has a clean and useful idea: packing two items into one token shortens the autoregressive horizon, and the constructive argument in §3.2 correctly shows that the induced distribution family over pair tokens can reproduce any item-space autoregressive distribution (for valid non-repeat pair sequences). The complexity decomposition in Theorem 4.1 is a helpful contribution, and the deployment results are credible: a 1.83× measured decoding speedup and a 0.178% stay-time lift on a large production platform, with a three-part ablation showing all training objectives matter. The main weakness is that the headline 'nearly 4× worst-case suboptimality reduction' is conditional on an unmeasured per-step error parity, and the speedup theorem does not establish the advertised 2×–4× range by itself. If the parity condition were verified or the claim appropriately qualified, the theoretical contribution would be solid; in its current form the central theoretical guarantee is oversold.
major comments (4)
- [§4.2 (Theorem 4.2), Appendix C Step 5] The advertised 'nearly 4×' suboptimality improvement is the ratio SubOpt_item/SubOpt_pair = 4·ε_item/ε_pair. The full 4× dividend requires ε_pair ≤ ε_item. This condition is never measured; Step 5 supports it only by order-of-magnitude heuristics (pretrain density ρ2≈6×10^4, GRPO group size G=32, temperature/top-k truncation), and Remark 1 concedes the unconditional claim is false. If ε_pair were only 2× ε_item, the improvement would be 2×, not 4×. Since this ratio is the load-bearing input to the theorem, the paper should either report direct TV estimates for both policies or withdraw the 'nearly 4×' headline.
- [Appendix C Step 5(ii)] The argument that pair-encoder generalization error is comparable to item-embedding error relies on reverse-pair cold-start being 'addressed by the OAR pretraining loss (§3.4).' However, §3.4 defines only L_pretrain over ordered exposure pairs (π_i,π_j) with i<j; no OAR loss for reverse pairs is described anywhere. This is an internal inconsistency in the very step that establishes ε_pair≈ε_item. Clarify the missing objective or remove the claim.
- [§4.1 (Theorem 4.1, Table 1)] The theorem decomposes FLOPs but does not prove a 2×–4× speedup: Table 1 shows the output-projection term grows by a factor n/2 in pair space, so the overall ratio depends on the relative weights of T_fixed, T_kv, and T_proj. No parameter regime is given under which the stated 2×–4× range holds, and the immediately following sentence says 'the overall decode speedup falls between 2× and 4×, reaching 1.83× in real deployment,' which is self-contradictory since 1.83 < 2. State explicit conditions for the speedup claim or reframe it as an empirical observation.
- [§3.2] The claimed bijection between Π_item and Π_pair is stated for the unrestricted vocabulary P(V)={ordered pairs with i≠j}. An arbitrary length-L/2 sequence from P(V) can unfold to an item sequence with repeats and hence not a permutation; the bijection holds only for the constrained subset of pair sequences with no repeated items. The action-mask constraint enforcing non-repetition is mentioned later (Section 1) but not incorporated into the formal equivalence. Please define the valid pair-sequence set and state the bijection for that set.
minor comments (4)
- [Abstract and §4.1] The '2× to 4× theoretical speedup' wording is difficult to reconcile with the measured 1.83×. The abstract is acceptable as a theoretical range, but §4.1's phrasing should be aligned to avoid the apparent contradiction.
- [§3.4] The pretraining loss equation is missing a closing parenthesis for the sigmoid argument, and the notation σ(1/d Σ e(...)) is unclear — specify which dimension is averaged and add parentheses.
- [Appendix C Step 5(i)] The claim that 'temperature scaling and top-k truncation (standard in GRPO)' neutralizes vocabulary-size effects is not described in the training section. If these are used, provide values; otherwise the claim is unsupported.
- [Table 5] The 'Improvement' row should specify which baseline each percentage is relative to; the 'strongest competing baseline' appears to be a different model in different columns.
Circularity Check
The 'nearly 4×' suboptimality dividend is not independently derived; it is the paper's asserted parity ε̄_pair≈ε̄_item restated as a theorem.
specific steps
-
self definitional
[Section 4.2 (Theorem 4.2) and Appendix C, Step 5]
"The net improvement ratio is SubOptitem/SubOptpair = 4·ε̄item/ε̄pair. When ε̄pair≤ε̄item (encoder-saturated regime), the full 4× dividend is realized. ... For the deployment regime n≤400, pretrain density ρ_2≫10^3, and G=32: ε̄_pair≈ε̄_item, yielding≈4× improvement."
The theorem's derived content is the standard quadratic-compounding lemma; the ratio SubOptitem/SubOptpair = 4·ε̄item/ε̄pair is an algebraic identity once horizon halving is granted. The headline 'nearly 4× improvement' therefore reduces to the asserted parity ε̄_pair≈ε̄_item. That parity is not measured from the model or from external benchmarks; Appendix C Step 5 constructs it from the heuristic pretrain density ρ_2≈6×10^4, the GRPO group size G=32, and a claimed gradient-variance match. Remark 1 concedes the unconditional claim is false. If ε̄_pair = 2 ε̄_item, the same identity gives only 2×, so the 'prediction' is exactly the assumed input, not an independent first-principles result.
full rationale
The paper has one genuinely circular/assumption-as-result element, and it is load-bearing for the strongest theoretical contribution. Theorem 4.2's bound is a conditional statement: the improvement over item-space generation is SubOptitem/SubOptpair = 4·ε̄item/ε̄pair. The 'nearly 4×' wording in the abstract, introduction, and conclusion treats this as the paper's headline guarantee, but the only way the ratio becomes ≈4 is the Appendix C Step 5 assertion ε̄_pair≈ε̄_item, which is justified by order-of-magnitude heuristics rather than measurement. This is not an external benchmark or a fitted value; it is a chosen constant that makes the headline number come out. The paper is unusually honest about this in Remark 1, but the abstract-level claim remains overstated. Other parts of the paper are not circular: the bijection and distribution-equivalence in §3.2 are constructive and mathematically self-contained; the speedup claim is conditional and separately validated by the measured 1.83× online speedup; the empirical gains on ML-1M, Amazon-Books, and RecFlow are standard experimental comparisons; and the self-citations to GoalRank and RecForest are not used to justify the paper's novel derivation. The circularity is therefore partial and localized to the suboptimality-improvement headline, which is the paper's central theoretical selling point.
Axiom & Free-Parameter Ledger
free parameters (2)
- Per-step TV mismatch ratio ε̄_pair/ε̄_item =
≈1 (asserted, not measured)
- Training hyperparameters (β, λ1, λ2, δ, G, beam) =
β=0.01, λ1=1, λ2=0.1, δ=0.2, G=16 offline / 32 online, beam=4
axioms (5)
- standard math Performance Difference Lemma and the quadratic error-compounding bound of Ross & Bagnell (2010) / Kakade & Langford (2002)
- domain assumption Reranking is a deterministic-transition MDP
- domain assumption Per-step action distribution deviates from optimal by ≤ε̄ in total variation at every step
- ad hoc to paper Encoder-saturated regime (n≤400, pretrain density ρ_2≫1e3, G=32) implies ε̄_pair≈ε̄_item
- ad hoc to paper PTR module is compute-bound and absorbed by the user-encoder wall-clock window
Cite this review
Pith. "Pith review of PSG: Pair-Space Generation for Efficient Generative Reranking." pith.science (2026). https://pith.science/paper/YMQ4IL62
@misc{pith2026260726427,
author = {Pith},
title = {Pith review of: PSG: Pair-Space Generation for Efficient Generative Reranking},
year = {2026},
howpublished = {\url{https://pith.science/paper/YMQ4IL62}},
note = {Machine review of arXiv:2607.26427}
}
abstract
Modern recommender systems adopt Generator-Evaluator (G-E) for list-wise reranking: a generator produces sequences from candidates and an evaluator scores them at sequence-level to filter out the optimal one for exposure. Auto-Regressive(AR), working as the backbone for generative recommendation, suffers two limitations. First, its complexity grows linearly with list length, forcing the system to generate fewer lists under rigorous latency constraints and thus limiting exploration. Second, teacher-forcing creates a train-test mismatch; cumulative errors worsen with length and degrade quality. To address these problems, we propose Pair-Space Generation (PSG), a reformulation that elevates the generation atom from individual items to ordered item pairs. Given $n$ candidate items, PSG operates over pair vocabulary of size $n(n-1)$ per request, generates only $L/2$ tokens. Pair token representations are produced on-the-fly by a pretrained pair-token representation module optimized over large scale exposure logs, eliminating the data sparsity that would otherwise plague a quadratic sized vocabulary. We establish three theoretical guarantees: (i) PSG is bijective with item-space generation and induces an equivalent family of sequence distributions, thus incurring no loss of expressiveness; (ii) generation in pair-token space achieves approximately a $2\times$ to $4\times$ speedup theoretically under moderate settings and $1.83\times$ in the real industrial environmental settings; and (iii) under outcome-only rewards, the worst-case suboptimality of PSG is bounded by $O((L/2)^2 \bar{\epsilon})$, representing a nearly $4\times$ improvement over item-space generation. Beyond benchmark-based validation, PSG has also been deployed on Kuaishou, delivering a 0.178\% lift in per-user stay time on the platform, which serves over 400 million daily active users.
Figures
Reference graph
Works this paper leans on
-
[1]
Qingyao Ai, Keping Bi, Jiafeng Guo, and W. Bruce Croft. 2018. Learning a Deep Listwise Context Model for Ranking Refinement. InThe 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, SIGIR 2018, Ann Arbor, MI, USA, July 08-12, 2018, Kevyn Collins-Thompson, Qiaozhu Mei, Brian D. Davison, Yiqun Liu, and Emine Yilmaz (...
arXiv 2018
-
[2]
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai. 2023. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, Houda Bouamor, Juan Pino, a...
2023
-
[3]
Irwan Bello, Sayali Kulkarni, Sagar Jain, Craig Boutilier, Ed Huai-hsin Chi, Elad Eban, Xiyang Luo, Alan Mackey, and Ofer Meshi. 2018. Seq2Slate: Re-ranking and Slate Optimization with RNNs.CoRRabs/1810.02019 (2018). arXiv:1810.02019 http://arxiv.org/abs/1810.02019
Pith/arXiv arXiv 2018
-
[4]
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. 2015. Scheduled sampling for sequence prediction with recurrent Neural networks. InProceedings of the 29th International Conference on Neural Information Processing Systems - Vol- ume 1(Montreal, Canada)(NIPS’15). MIT Press, Cambridge, MA, USA, 1171–1179
2015
-
[5]
Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T. Freeman. 2022. MaskGIT: Masked Generative Image Transformer. InIEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022. IEEE, 11305–11315. https://doi.org/10.1109/CVPR52688.2022.01103
arXiv 2022
-
[6]
Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Lau- rent Sifre, and John Jumper. 2023. Accelerating Large Language Model Decoding with Speculative Sampling.CoRRabs/2302.01318 (2023). https://doi.org/10.48550/ ARXIV.2302.01318 arXiv:2302.01318
-
[7]
Jinglin Chen, Qiwei Li, Zuchao Li, Baoyuan Qi, Guoming Liu, Haojun Ai, Hai Zhao, and Ping Wang. 2025. Faster In-Context Learning for LLMs via N-Gram Trie Speculative Decoding. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, Christos Christodoulopoulos, Tanmoy Chakrabo...
-
[8]
Xiaoyang Chen, Yanjiang Liu, Ben He, Le Sun, and Yingfei Sun. 2023. Under- standing differential search index for text retrieval. InFindings of the association for computational linguistics: ACL 2023. 10701–10717
2023
-
[9]
Paul Covington, Jay Adams, and Emre Sargin. 2016. Deep Neural Networks for YouTube Recommendations. InProceedings of the 10th ACM Conference on Recommender Systems, Boston, MA, USA, September 15-19, 2016, Shilad Sen, Werner Geyer, Jill Freyne, and Pablo Castells (Eds.). ACM, 191–198. https://doi.org/10. 1145/2959100.2959190
arXiv 2016
-
[10]
Tri Dao. 2024. FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning. InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net. https://openreview.net/forum?id=mZn2Xyh9Ec
2024
-
[11]
Fu, Stefano Ermon, Atri Rudra, and Christopher Ré
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. InAdvances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022, Sanmi Koyejo, S...
2022
-
[12]
Chao Feng, Wuchao Li, Defu Lian, Zheng Liu, and Enhong Chen. 2022. Recommender Forest for Efficient Retrieval. InAdvances in Neural Informa- tion Processing Systems 35: Annual Conference on Neural Information Process- ing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - De- cember 9, 2022, Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belg...
2022
-
[13]
Yufei Feng, Yu Gong, Fei Sun, Qingwen Liu, and Wenwu Ou. 2021. Revisit Recommender System in the Permutation Prospective.CoRRabs/2102.12057 (2021). arXiv:2102.12057 https://arxiv.org/abs/2102.12057
Pith/arXiv arXiv 2021
-
[14]
Yufei Feng, Binbin Hu, Yu Gong, Fei Sun, Qingwen Liu, and Wenwu Ou. 2021. GRN: Generative Rerank Network for Context-wise Recommendation.CoRR abs/2104.00860 (2021). arXiv:2104.00860 https://arxiv.org/abs/2104.00860
Pith/arXiv arXiv 2021
-
[15]
Courville, and Yoshua Bengio
Anirudh Goyal, Alex Lamb, Ying Zhang, Saizheng Zhang, Aaron C. Courville, and Yoshua Bengio. 2016. Professor Forcing: A New Algorithm for Training Recurrent Networks. InAdvances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems 2016, December 5-10, 2016, Barcelona, Spain, Daniel D. Lee, Masashi Sugiyam...
2016
-
[16]
Chengcheng Guo, Kuo Cai, Yu Zhou, Qiang Luo, Ruiming Tang, Han Li, Kun Gai, and Guorui Zhou. 2026. PROMISE: Process Reward Models Unlock Test-Time Scaling Laws in Generative Recommendations.arXiv preprint arXiv:2601.04674 (2026)
arXiv 2026
-
[17]
F. Maxwell Harper and Joseph A. Konstan. 2016. The MovieLens Datasets: History and Context.ACM Trans. Interact. Intell. Syst.5, 4 (2016), 19:1–19:19. https://doi.org/10.1145/2827872
doi:10.1145/2827872 2016
-
[18]
Ruining He and Julian J. McAuley. 2016. Ups and Downs: Modeling the Vi- sual Evolution of Fashion Trends with One-Class Collaborative Filtering. In Proceedings of the 25th International Conference on World Wide Web, WWW 2016, Montreal, Canada, April 11 - 15, 2016, Jacqueline Bourdeau, Jim Hendler, Roger Nkambou, Ian Horrocks, and Ben Y. Zhao (Eds.). ACM, ...
arXiv 2016
-
[19]
Yupeng Hou, Jiacheng Li, Ashley Shin, Jinsung Jeon, Abhishek Santhanam, Wei Shao, Kaveh Hassani, Ning Yao, and Julian J. McAuley. 2025. Generating Long Semantic IDs in Parallel for Recommendation. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, V.2, KDD 2025, Toronto ON, Canada, August 3-7, 2025, Luiza Antonie, Jian...
arXiv 2025
-
[20]
Xun Huang, Zhengqi Li, Guande He, Mingyuan Zhou, and Eli Shechtman. 2025. Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion. InAdvances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November ...
2025
-
[21]
Mann, and Danilo J
Ray Jiang, Sven Gowal, Yuqiu Qian, Timothy A. Mann, and Danilo J. Rezende. 2019. Beyond Greedy Ranking: Slate Optimization via List-CVAE. In7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net. https://openreview.net/forum?id=r1xX42R5Fm
2019
-
[22]
Sham Kakade and John Langford. 2002. Approximately Optimal Approximate Reinforcement Learning. InProceedings of the Nineteenth International Conference on Machine Learning (ICML ’02). Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 267–274
2002
-
[23]
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient Memory Management for Large Language Model Serving with PagedAtten- tion. InProceedings of the 29th Symposium on Operating Systems Principles, SOSP 2023, Koblenz, Germany, October 23-26, 2023, Jason Flinn, Margo I. Se...
arXiv 2023
-
[24]
Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023. Fast Inference from Transformers via Speculative Decoding. InInternational Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA (Proceedings of Ma- chine Learning Research, Vol. 202), Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jona...
2023
-
[26]
Yi Li, Jieming Zhu, Weiwen Liu, Liangcai Su, Guohao Cai, Qi Zhang, Ruiming Tang, Xi Xiao, and Xiuqiang He. 2022. PEAR: Personalized Re-ranking with Contextualized Transformer for Recommendation. InCompanion of The Web Conference 2022, Virtual Event / Lyon, France, April 25 - 29, 2022, Frédérique Laforest, Raphaël Troncy, Elena Simperl, Deepak Agarwal, Ari...
doi:10.1145/3487553 2022
-
[27]
Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2024. Let’s Verify Step by Step. InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net. https://openreview.net/forum?id=v8L0pN6EOi
2024
-
[29]
Qi Liu, Kai Zheng, Rui Huang, Wuchao Li, Kuo Cai, Yuan Chai, Yanan Niu, Yiqun Hui, Bing Han, Na Mou, Hongning Wang, Wentian Bao, Yunen Yu, Guorui Zhou, Han Li, Yang Song, Defu Lian, and Kun Gai. 2025. RecFlow: An Industrial Full Flow Recommendation Dataset. InThe Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April ...
2025
-
[30]
McAuley, Dong Zheng, Peng Jiang, and Kun Gai
Shuchang Liu, Qingpeng Cai, Zhankui He, Bowen Sun, Julian J. McAuley, Dong Zheng, Peng Jiang, and Kun Gai. 2023. Generative Flow Network for Listwise Recommendation. InProceedings of the 29th ACM SIGKDD Conference on Knowl- edge Discovery and Data Mining, KDD 2023, Long Beach, CA, USA, August 6- 10, 2023, Ambuj K. Singh, Yizhou Sun, Leman Akoglu, Dimitrio...
arXiv 2023
-
[31]
Yue Meng, Cheng Guo, Yi Cao, Tong Liu, and Bo Zheng. 2025. A Generative Re-ranking Model for List-level Multi-objective Optimization at Taobao. InPro- ceedings of the 48th International ACM SIGIR Conference on Research and Develop- ment in Information Retrieval, SIGIR 2025, Padua, Italy, July 13-18, 2025, Nicola Ferro, Maria Maistro, Gabriella Pasi, Omar ...
arXiv 2025
-
[32]
Jie Ou, Yueming Chen, and Wenhong Tian. 2024. Lossless Acceleration of Large Language Model via Adaptive N-gram Parallel Decoding. InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Industry Track, NAACL 2024, Mexico City, Mexico, June 16-21, 2024, Yi Yang, Aida...
-
[33]
Liang Pang, Jun Xu, Qingyao Ai, Yanyan Lan, Xueqi Cheng, and Jirong Wen
-
[34]
Changhua Pei, Yi Zhang, Yongfeng Zhang, Fei Sun, Xiao Lin, Hanxiao Sun, Jian Wu, Peng Jiang, Junfeng Ge, Wenwu Ou, and Dan Pei. 2019. Personalized re-ranking for recommendation. InProceedings of the 13th ACM Conference on Recommender Systems, RecSys 2019, Copenhagen, Denmark, September 16-20, 2019, Toine Bogers, Alan Said, Peter Brusilovsky, and Domonkos ...
arXiv 2019
-
[35]
Tran, Jonah Samost, Maciej Kula, Ed H
Shashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan Hulikal Kesha- van, Trung Vu, Lukasz Heldt, Lichan Hong, Yi Tay, Vinh Q. Tran, Jonah Samost, Maciej Kula, Ed H. Chi, and Mahesh Sathiamoorthy. 2023. Rec- ommender Systems with Generative Retrieval. InAdvances in Neural Infor- mation Processing Systems 36: Annual Conference on Neural Information Pro- ...
2023
-
[36]
Yuxin Ren, Qiya Yang, Yichun Wu, Wei Xu, Yalong Wang, and Zhiqiang Zhang
-
[37]
Stephane Ross and Drew Bagnell. 2010. Efficient Reductions for Imitation Learn- ing. InProceedings of the Thirteenth International Conference on Artificial Intelli- gence and Statistics (Proceedings of Machine Learning Research, Vol. 9), Yee Whye Teh and Mike Titterington (Eds.). PMLR, Chia Laguna Resort, Sardinia, Italy, 661–668. https://proceedings.mlr....
2010
-
[38]
Robin Schmidt, Telmo Pires, Stephan Peitz, and Jonas Lööf. 2022. Non- Autoregressive Neural Machine Translation: A Call for Clarity. InProceed- ings of the 2022 Conference on Empirical Methods in Natural Language Pro- cessing, Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for Computational Linguistics, Abu Dhabi, United Arab Emirates...
-
[39]
Claude E. Shannon. 1948. A mathematical theory of communication.Bell Syst. Tech. J.27, 3 (1948), 379–423. https://doi.org/10.1002/J.1538-7305.1948.TB01338.X
arXiv 1948
-
[40]
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo. 2024. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.CoRRabs/2402.03300 (2024). https://doi.org/10.48550/ARXIV.2402.03300 arXiv:2402.03300
-
[41]
Shiqi Shen, Yong Cheng, Zhongjun He, Wei He, Hua Wu, Maosong Sun, and Yang Liu. 2016. Minimum Risk Training for Neural Machine Translation. InProceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Katrin Erk and Noah A. Smith (Eds.). Association for Computa- tional Linguistics, Berlin, Germany, 168...
doi:10.18653/v1/p16- 2016
-
[42]
Xiaowen Shi, Fan Yang, Ze Wang, Xiaoxu Wu, Muzhi Guan, Guogang Liao, Yongkang Wang, Xingxing Wang, and Dong Wang. 2023. PIER: Permutation-Level Interest-Based End-to-End Re-ranking Framework in E-commerce. InProceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2023, Long Beach, CA, USA, August 6-10, 2023, Ambuj K. Sing...
arXiv 2023
-
[43]
Dekai Sun, Yiming Liu, Jiafan Zhou, Xun Liu, Chenchen Yu, Yi Li, Jun Zhang, Huan Yu, and Jie Jiang. 2026. OneRanker: Unified Generation and Ranking with One Model in Industrial Advertising Recommendation.CoRRabs/2603.02999 (2026). https://doi.org/10.48550/ARXIV.2603.02999 arXiv:2603.02999
-
[44]
Cohen, and Donald Metzler
Yi Tay, Vinh Tran, Mostafa Dehghani, Jianmo Ni, Dara Bahri, Harsh Mehta, Zhen Qin, Kai Hui, Zhe Zhao, Jai Prakash Gupta, Tal Schuster, William W. Cohen, and Donald Metzler. 2022. Transformer Memory as a Differentiable Search Index. InAdvances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, Neur...
2022
-
[45]
OneRec Team. 2025. OpenOneRec Technical Report.CoRRabs/2512.24762 (2025). https://doi.org/10.48550/ARXIV.2512.24762 arXiv:2512.24762
-
[46]
Gomez, Łukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. InProceedings of the 31st International Conference on Neural Information Processing Systems(Long Beach, California, USA)(NIPS’17). Curran Associates Inc., Red Hook, NY, USA, 6000–6010
2017
-
[47]
Ruoxi Wang, Bin Fu, Gang Fu, and Mingliang Wang. 2017. Deep & Cross Network for Ad Click Predictions. InProceedings of the ADKDD’17, Halifax, NS, Canada, August 13 - 17, 2017. ACM, 12:1–12:7. https://doi.org/10.1145/3124749.3124754
arXiv 2017
-
[48]
Shuli Wang, Yinqiu Huang, Changhao Li, Yuan Zhou, Yonggang Liu, Yongqiang Zhang, Yinhua Zhu, Haitao Wang, and Xingxing Wang. 2025. You Only Evaluate Once: A Tree-based Rerank Method at Meituan. InProceedings of the 34th ACM International Conference on Information and Knowledge Management(Seoul, Re- public of Korea)(CIKM ’25). Association for Computing Mac...
arXiv 2025
-
[49]
Shuli Wang, Xue Wei, Senjie Kou, Chi Wang, Wenshuai Chen, Qi Tang, Yinhua Zhu, Xiong Xiao, and Xingxing Wang. 2025. NLGR: Utilizing Neighbor Lists for Generative Rerank in Personalized Recommendation Systems. InCompanion Proceedings of the ACM on Web Conference 2025, WWW 2025, Sydney, NSW, Aus- tralia, 28 April 2025 - 2 May 2025, Guodong Long, Michale Blu...
arXiv 2025
-
[50]
Yejing Wang, Shengyu Zhou, Jinyu Lu, Ziwei Liu, Langming Liu, Maolin Wang, Wenlin Zhang, Feng Li, Wenbo Su, Pengjie Wang, Jian Xu, and Xiangyu Zhao
-
[51]
Jianxiong Wei, Anxiang Zeng, Yueqiu Wu, Peng Guo, Qingsong Hua, and Qingpeng Cai. 2020. Generator and Critic: A Deep Reinforcement Learning Approach for Slate Re-ranking in E-commerce.CoRRabs/2005.12206 (2020). arXiv:2005.12206 https://arxiv.org/abs/2005.12206
Pith/arXiv arXiv 2020
-
[52]
Samuel Williams, Andrew Waterman, and David Patterson. 2009. Roofline: an insightful visual performance model for multicore architectures.Commun. ACM 52, 4 (April 2009), 65–76. https://doi.org/10.1145/1498765.1498785
arXiv 2009
-
[53]
Yunjia Xi, Weiwen Liu, Jieming Zhu, Xilong Zhao, Xinyi Dai, Ruiming Tang, Weinan Zhang, Rui Zhang, and Yong Yu. 2022. Multi-Level Interaction Reranking with User Behavior History. InSIGIR ’22: The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, Madrid, Spain, July 11 - 15, 2022, Enrique Amigó, Pablo Castells, ...
arXiv 2022
-
[54]
Heming Xia, Tao Ge, Peiyi Wang, Si-Qing Chen, Furu Wei, and Zhifang Sui
-
[55]
Junwei Xu, Zhibo Xiao, Chuxin Chen, Chengyu Lai, Qijie Shen, Jiuning Lin, Dimin Wang, Jialin Zhu, and Xiao-Ping Zhang. 2026. OMGRec: One-time Matching- based Generative Rerank with Permutation-level Modeling in E-commerce. In Proceedings of the ACM Web Conference 2026, WWW 2026, Dubai, United Arab Emirates, originally scheduled for April 13-17, 2026, resc...
arXiv 2026
-
[56]
Hailan Yang, Zhenyu Qi, Shuchang Liu, Xiaoyu Yang, Xiaobei Wang, Xiang Li, Lantao Hu, Han Li, and Kun Gai. 2025. Comprehensive List Generation for Multi-Generator Reranking. InProceedings of the 48th International ACM PSG: Pair-Space Generation for Efficient Generative Reranking Conference acronym ’XX, June 0305, 2018, Woodstock, NY SIGIR Conference on Re...
arXiv 2025
-
[57]
Lantao Yu, Weinan Zhang, Jun Wang, and Yong Yu. 2017. SeqGAN: sequence generative adversarial nets with policy gradient. InProceedings of the Thirty- First AAAI Conference on Artificial Intelligence(San Francisco, California, USA) (AAAI’17). AAAI Press, 2852–2858
2017
-
[58]
Jiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang, Rui Li, Xuan Cao, Leon Gao, Zhaojie Gong, Fangda Gu, Jiayuan He, Yinghai Lu, and Yu Shi. 2024. Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations. InForty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 (Pr...
2024
-
[59]
Kaike Zhang, Xiaobei Wang, Shuchang Liu, Hailan Yang, Xiang Li, Lantao Hu, Han Li, Qi Cao, Fei Sun, and Kun Gai. 2026. GoalRank: Group-Relative Optimiza- tion for a Large Ranking Model. InThe Fourteenth International Conference on Learning Representations. https://openreview.net/forum?id=gTMzRm8fb0
2026
-
[60]
Xin Zhao, Jiaxin Li, Zhiwei Fang, Yuchen Guo, Jinyuan Zhao, Jie He, Wenlong Chen, Changping Peng, and Guiguang Ding. 2024. JDRec: Practical Actor-Critic Framework for Online Combinatorial Recommender System. InProceedings of the 23rd International Conference on Autonomous Agents and Multiagent Sys- tems, AAMAS 2024, Auckland, New Zealand, May 6-10, 2024, ...
arXiv 2024
-
[61]
Guorui Zhou, Hengrui Hu, Hongtao Cheng, Huanjie Wang, Jiaxin Deng, Jing- hao Zhang, Kuo Cai, Lejian Ren, Lu Ren, Liao Yu, Pengfei Zheng, Qiang Luo, Qianqian Wang, Qigen Hu, Rui Huang, Ruiming Tang, Shiyao Wang, Shujie Yang, Tao Wu, Wuchao Li, Xinchen Luo, Xingmei Wang, Yi Su, Yunfan Wu, Zexuan Cheng, Zhanyu Liu, Zixing Zhang, Bin Zhang, Boxuan Wang, Chaoy...
-
[62]
Ruitao Zhu, Yangsu Liu, Dagui Chen, Zhenjia Ma, Chufeng Shi, Zhenzhe Zheng, Jie Zhang, Jian Xu, Bo Zheng, and Fan Wu. 2025. Contextual Generative Auction with Permutation-level Externalities for Online Advertising. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, V.1, KDD 2025, Toronto, ON, Canada, August 3-7, 2025, ...
arXiv 2025
-
[63]
pair space always yields4 × improvement
Tao Zhuang, Wenwu Ou, and Zhirong Wang. 2018. Globally Optimized Mutual Influence Aware Ranking in E-Commerce Search. InProceedings of the Twenty- Seventh International Joint Conference on Artificial Intelligence, IJCAI 2018, July 13-19, 2018, Stockholm, Sweden, Jérôme Lang (Ed.). ijcai.org, 3725–3731. https: //doi.org/10.24963/IJCAI.2018/518 A Odd L Exte...
-
[2020]
SetRank: Learning a Permutation-Invariant Ranking Model for Information Retrieval. InProceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval, SIGIR 2020, Virtual Event, China, July 25-30, 2020, Jimmy X. Huang, Yi Chang, Xueqi Cheng, Jaap Kamps, Vanessa Murdock, Ji-Rong Wen, and Yiqun Liu (Eds.). ACM,...
arXiv 2020
-
[2023]
Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation. InFindings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023 (Findings of ACL, Vol. EMNLP 2023), Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, 3909–3925. https://doi.org/10.18...
-
[2024]
Non-autoregressive Generative Models for Reranking Recommendation. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining(Barcelona, Spain)(KDD ’24). Association for Computing Machinery, New York, NY, USA, 5625–5634. https://doi.org/10.1145/3637528.3671645
-
[2026]
NEZHA: A Zero-sacrifice and Hyperspeed Decoding Architecture for Generative Recommendations. InProceedings of the ACM Web Conference 2026, WWW 2026, Dubai, United Arab Emirates, originally scheduled for April 13-17, 2026, rescheduled for June 29 - July 3, 2026, Hakim Hacid, Yoelle Maarek, Francesco Bonchi, Ido Guy, and Emine Yilmaz (Eds.). ACM, 8073–8082....
arXiv 2026
This paper was first reviewed by deepseek-v4-flash on August 1, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.