REVIEW 4 major objections 5 minor 71 references
Scaling Sequential Recommendation Models with Transformers
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This paper shows that transformer-based sequential recommenders exhibit language-model-style scaling laws, making NDCG@5 predictable from model size and training data.
desk verdict First scaling-law analysis for sequential recommendation, with a clean catalog-independent architecture, but the quantitative claims need matched evaluation and held-out validation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Scalable Recommendation Transformer (SRT), a modification of SASRec in which each catalog item is represented not by a trainable lookup embedding but by the output of a trainable feature extractor applied to the item's title and brand tokens; this decouples the parameter count from catalog size. The learning signal is a sampled-softmax contrastive loss over a popularity-sampled set of negatives with a logQ correction, which generalizes the single-negative losses used by SASRec and BERT4Rec. The scaling analysis rests on a risk-decomposition functional form $NDCG(N,T) = E - A/N^{\alpha} - B/T^{\beta}$, fitted to the envelope of best-performing runs across model sizes and dataset sizes, and on FLOP counts as the compute proxy. This combination is what lets the authors separate the effects of parameter count and data volume and extrapolate performance to untested configurations.
What would settle it
Re-run MINCE, ICLRec, CoSeRec, and CL4SRec on the same Beauty and Sports splits using the same 10,000-negative NDCG@5 protocol as the SRT models; if their scores rise above 0.0405 and 0.0206, the claimed superiority of fine-tuned SRT-1K does not hold.
Extended reading notes
Core claim
On its own terms, the paper establishes that a transformer-based sequential recommender trained on the full Amazon Product Data exhibits two scaling regularities. First, maximum NDCG@5 as a function of compute follows a saturating sigmoid with an upper limit of about 0.149, so gains from extra FLOPs diminish. Second, for fixed model complexity $N$ and number of seen interactions $T$, the achievable NDCG@5 is fit by $NDCG(N,T) = 0.163 - 18.56/N^{0.376} - 2.9/T^{0.364}$, with nearly equal exponents on $N$ and $T$, meaning parameters and data contribute similarly at the margin. The paper further claims that the scaling transfers: models pre-trained on the full data and then fine-tuned on small domains outperform both models trained from scratch and stronger published baselines, with relative NDCG@5 gains above 12% on Beauty and 7% on Sports.
Load-bearing premise
The comparison that supports transferability assumes NDCG@5 over 10,000 random negatives is directly comparable to published baseline numbers that were obtained with 100 negatives, even though the negative-set size for the cited baselines is never stated.
Editorial extensions
If this is right
- For a fixed FLOPs budget, the fitted sigmoid gives an estimated ceiling for achievable NDCG@5, so practitioners can stop spending compute once the budget passes the point of diminishing returns.
- Because the exponents on $N$ and $T$ are nearly equal, balanced increases in model size and training data should be roughly as effective as doubling either alone.
- SRT's parameter count is independent of catalog size, so adding or removing items does not require changing the model's vocabulary or retraining the embedding table.
- Fine-tuning larger pre-trained models on smaller domains improves NDCG@5 by more than 12% on Beauty and 7% on Sports relative to from-scratch training, suggesting pre-training at scale is a viable deployment strategy.
- The paper's scaling laws are derived on the full Amazon data, so they can guide data collection and model selection before training large models on a new catalog.
Reading between the lines
- Editorial inference: If the scaling laws hold beyond the Amazon catalog, the same SRT recipe could be applied to other marketplaces with different item text, making pre-training transferable across businesses; a test would be fine-tuning on an unseen catalog's small domain.
- Editorial inference: The saturating sigmoid suggests an irreducible ambiguity in next-item prediction from a long-tailed catalog, which no amount of compute alone can overcome; content-side signals or better item representations would be the next lever.
- Editorial inference: The paper's protocol uses 10,000 sampled negatives for its own models, whereas published baselines are cited from papers that typically use 100 negatives; re-evaluating baselines under the 10,000-negative protocol would test whether the reported margins are protocol artifacts.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SRT, a transformer-based sequential recommender that replaces the catalog-sized item embedding table with a trainable feature extractor over tokenized item text, making the model parameter count independent of catalog size. The authors train SRT on the full Amazon Product Data (APD) and report NDCG@5 for eight model sizes and dataset sizes up to 8.2M interactions. From these runs they fit two parametric laws: a sigmoid relating NDCG@5 to log FLOPs (Eq. 3) and a risk-decomposition form NDCG(N,T) = E - A/N^alpha - B/T^beta (Eq. 5). They also fine-tune large pre-trained SRT models on the Beauty and Sports subsets and report improvements over from-scratch training and over published baselines. The paper claims that these results reveal scaling behaviors similar to language modeling and provide a practical roadmap for compute-optimal training and transfer in sequential recommendation.
Significance. If the proposed scaling laws are genuinely predictive, the paper would be a useful step toward compute-aware design of sequential recommender systems, and the SRT architecture itself is a reusable contribution. The release of code and models is a strength, as is the use of the full APD rather than only 5-core subsets. However, the claim that SRT outperforms prior methods rests on an evaluation protocol that is not matched to the cited baselines, and the fitted laws lack validation on held-out configurations or uncertainty quantification. The scaling-law results are therefore currently better described as descriptive curve fitting than as established predictive laws, so the central claims need additional support.
major comments (4)
- [Sec. 4.1 / Tables 1 and 4] SRT is evaluated with 10,000 random negatives per positive (Sec. 4.1), while the baseline numbers in Tables 1 and 4 are taken from prior papers that standardly use 100 negatives; the paper never states the negative-pool size for those baselines. NDCG@5 computed over 10,001 candidates is not commensurable with NDCG@5 over 101 candidates, so the claims that SRT-1K 'outperforms all other alternatives' (Sec. 3) and that the fine-tuned variants outperform all alternatives by a margin (Sec. 4.5) are not supported as reported. The authors should re-run the baselines under the same negative-sampling protocol, or explicitly restrict the external comparison to NDCG values obtained with identical candidate sets.
- [Sec. 4.4.2 / Eq. (5)] The scaling law in Eq. (5) is fitted to the same runs it is then used to describe; the paper reports no held-out configurations, no cross-validation, and no bootstrap or other uncertainty estimates, and the fit is based on only eight model sizes and dataset sizes up to 8.2M interactions. A 'prediction' for a given (N,T) pair is therefore a direct evaluation of the fitted curve rather than a test of the law's validity. To support the claim that NDCG(N,T) is predictable, the authors should fit the functional form on a subset of configurations and evaluate it on held-out configurations, reporting residuals and prediction intervals.
- [Sec. 4.4.1 / Eq. (3)] The sigmoidal envelope in Eq. (3) is constructed by selecting, for each FLOP budget, the maximum NDCG among the runs in the experiment. With a small number of runs, this envelope is sensitive to the particular configurations included, and the fit parameters are again evaluated on the same points that generated them. The claimed diminishing-returns point at log(FLOPs)=30.7 is thus not validated. The authors should provide a validation procedure, such as leave-one-configuration-out fits, and report error bars on the sigmoid parameters.
- [Sec. 4.3 / Sec. 4.4.2] The paper defines T as 'seen interactions' equal to the dataset size multiplied by the number of epochs, and treats multi-epoch training as equivalent to seeing more data. This equivalence is nontrivial: repeated passes over the same sequences are not independent samples, and the improvement from multiple epochs may partly reflect optimization dynamics rather than additional data diversity. If T conflates these two effects, the interpretation of Eq. (5) as a scaling law in data size is weakened. The authors should test whether the fitted exponents and the quality of the fit change when T is defined as the number of unique interactions rather than seen interactions.
minor comments (5)
- [Sec. 4.4.1] The text states that log(FLOPs)=30.7 corresponds to approximately 2.15e-13 FLOPs; this should be 2.15e13 FLOPs, since e^30.7 is on the order of 10^13.
- [Sec. 4.4.2] Eq. (5) is introduced with 'N denotes the total parameter count,' but Sec. 4.3 and Figure 5 emphasize non-embedding parameters, and Figure 7's axis is labeled only 'N'. The authors should clarify which definition of N was used in the fit and whether token embeddings are included.
- [Figure 5] The caption of the left panel says 'number of seen iterations' but should say 'number of seen interactions' to match the terminology used in the text.
- [Sec. 4.4.2] The notation N is used both for the set of negatives in Eq. (2) and for the parameter count in Eq. (4); please use distinct symbols to avoid ambiguity.
- [Table 2] The column labels are easy to misread; please label each column explicitly as APD raw, APD 5-core, Beauty raw, Beauty 5-core, Sports raw, and Sports 5-core.
Circularity Check
The scaling-law 'predictions' in Eqs. (3) and (5) are fitted curves evaluated on the same data used to obtain them; the transfer-learning results are separate empirical claims and are not themselves circular.
-
fitted input called prediction
[Section 4.4.1, Eq. (3), Figure 6]
"In the experiments, we recorded the maximum NDCG achieved for each FLOP budget. Figure 6 shows such points together with linear and sigmoidal fits. ... a sigmoidal fit appears more appropriate, in which case it corresponds to: NDCG(FLOPS) ≜ 0.396/(1+e^{−0.18(log(FLOPS)−24.44)}) − 0.247."
The stated purpose is to 'get an estimate of the maximum achievable performance' for a fixed FLOP budget, but the estimate is obtained by evaluating this sigmoid at that FLOP count. The constants were fit to the same envelope points shown in Figure 6, so the predicted NDCG is the fitted curve by construction. The asymptotic plateau of 0.149 and the diminishing-return point are properties of the chosen sigmoidal form fitted to those points, not independent measurements. No held-out FLOP budget or bootstrap validation is used to test extrapolation.
-
fitted input called prediction
[Section 4.4.2, Eq. (5), Figure 7]
"To fit the model, we use a non-linear least squares approach and constrain the model coefficients to be non-negatives to avoid nonsensical solutions. We obtain the following solution: NDCG(N,T) ≜ 0.163 − 18.56/N^0.376 − 2.9/T^0.364. ... Figure 7 shows the predicted NDCG@5 score as a function of the number of model parameters, N, and number of seen interactions."
The 'predicted NDCG@5 score' in Figure 7 is obtained by evaluating Eq. (5) at the same N and T values used to fit its four constants. The coefficients and exponents were determined by nonlinear least squares on these observed maxima, so for any N/T pair the 'prediction' is the fitted curve by construction. The paper frames this as a scaling law and suggests extrapolation to novel regimes, but no held-out configurations or independent validation are provided, so the predictive content reduces to the fit itself.
full rationale
The central scaling-law claim is partially circular in the specific sense of fitted parameters being relabeled as predictions: Eq. (3) and Eq. (5) are explicitly fit to the experimental NDCG points they are then used to 'estimate' or 'predict.' Because these fitted curves are the entire basis for the claimed analytical laws and extrapolation to untested model/data sizes, the predictive step reduces by construction to the curve-fitting step. This is not a case of self-citation: the paper does not rely on the authors' prior work as load-bearing evidence. The SRT architecture and the fine-tuning comparison are independent, internally consistent empirical results, though the external baseline comparisons in Tables 1 and 4 are undermined by the 10,000-negative versus 100-negative evaluation mismatch, which is a correctness and comparability concern rather than a circularity concern. The paper is transparent that the equations are fits, so the flaw is partial: the scaling-law formulas are descriptive summaries presented with predictive framing, while the transfer-learning contribution retains independent empirical content. Overall circularity score is therefore 6 rather than higher.
Assumptions & free parameters
free parameters (12)
- sigmoid scale A =
0.396
- sigmoid growth rate k =
-0.18
- sigmoid midpoint logF0 =
24.44
- sigmoid offset B =
-0.247
- risk decomposition E =
0.163
- risk decomposition A =
18.56
- risk decomposition alpha =
0.376
- risk decomposition B =
2.9
- risk decomposition beta =
0.364
- number of negatives SRT-X =
10, 100, 300, 1000
- softmax temperature tau
- EWC lambda =
100
assumptions (6)
- domain assumption The sampled softmax loss with logQ correction approximates the full-catalog softmax
- domain assumption Item text (title and brand) tokenized with SentencePiece 30k vocab captures enough signal to learn user preferences
- ad hoc to paper The functional form NDCG(N,T) = E - A/N^alpha - B/T^beta is a valid risk decomposition
- ad hoc to paper The sigmoidal form for NDCG as a function of log(FLOPs) is the correct saturating model
- ad hoc to paper Multi-epoch training with the same sequences is equivalent to more data, and 'seen interactions' = dataset size times epochs is the right quantity
- domain assumption 10K random negatives for evaluation yield a fair, comparable approximation to full-catalog NDCG across methods
Cite this review
Pith. "Pith review of Scaling Sequential Recommendation Models with Transformers." pith.science (2026). https://pith.science/paper/5CCWAUO6
@misc{pith2026241207585,
author = {Pith},
title = {Pith review of: Scaling Sequential Recommendation Models with Transformers},
year = {2026},
howpublished = {\url{https://pith.science/paper/5CCWAUO6}},
note = {Machine review of arXiv:2412.07585}
}
read the original abstract
Modeling user preferences has been mainly addressed by looking at users' interaction history with the different elements available in the system. Tailoring content to individual preferences based on historical data is the main goal of sequential recommendation. The nature of the problem, as well as the good performance observed across various domains, has motivated the use of the transformer architecture, which has proven effective in leveraging increasingly larger amounts of training data when accompanied by an increase in the number of model parameters. This scaling behavior has brought a great deal of attention, as it provides valuable guidance in the design and training of even larger models. Taking inspiration from the scaling laws observed in training large language models, we explore similar principles for sequential recommendation. We use the full Amazon Product Data dataset, which has only been partially explored in other studies, and reveal scaling behaviors similar to those found in language models. Compute-optimal training is possible but requires a careful analysis of the compute-performance trade-offs specific to the application. We also show that performance scaling translates to downstream tasks by fine-tuning larger pre-trained models on smaller task-specific domains. Our approach and findings provide a strategic roadmap for model training and deployment in real high-dimensional preference spaces, facilitating better training and inference efficiency. We hope this paper bridges the gap between the potential of transformers and the intrinsic complexities of high-dimensional sequential recommendation in real-world recommender systems. Code and models can be found at https://github.com/mercadolibre/srt
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Alabdulmohsin, Behnam Neyshabur, and Xiaohua Zhai
Ibrahim M. Alabdulmohsin, Behnam Neyshabur, and Xiaohua Zhai. 2022. Revis- iting Neural Scaling Laws in Language and Vision. (2022)
work page 2022
-
[2]
Alabdulmohsin, Xiaohua Zhai, Alexander Kolesnikov, and Lucas Beyer
Ibrahim M. Alabdulmohsin, Xiaohua Zhai, Alexander Kolesnikov, and Lucas Beyer. 2023. Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design. (2023)
work page 2023
-
[3]
Newsha Ardalani, Carole-Jean Wu, Zeliang Chen, Bhargav Bhushanam, and Adnan Aziz. 2022. Understanding Scaling Laws for Recommendation Models. CoRR (2022). arXiv:2208.08489
arXiv 2022
-
[4]
Campos, Fernando Díez, and Iván Cantador
Pedro G. Campos, Fernando Díez, and Iván Cantador. 2014. Time-aware recom- mender systems: a comprehensive survey and analysis of existing evaluation protocols. User Model. User Adapt. Interact. 24, 1-2 (2014), 67–119
work page 2014
-
[5]
Yongjun Chen, Zhiwei Liu, Jia Li, Julian J. McAuley, and Caiming Xiong. 2022. Intent Contrastive Learning for Sequential Recommendation. In WWW ’22: The ACM Web Conference 2022, Virtual Event, Lyon, France, April 25 - 29, 2022, Frédérique Laforest, Raphaël Troncy, Elena Simperl, Deepak Agarwal, Aristides Gionis, Ivan Herman, and Lionel Médini (Eds.). ACM,...
work page 2022
-
[6]
Aidan Clark, Diego de Las Casas, Aurelia Guy, Arthur Mensch, Michela Pa- ganini, Jordan Hoffmann, Bogdan Damoc, Blake A. Hechtman, Trevor Cai, Se- bastian Borgeaud, George van den Driessche, Eliza Rutherford, Tom Henni- gan, Matthew J. Johnson, Albin Cassirer, Chris Jones, Elena Buchatskaya, David Budden, Laurent Sifre, Simon Osindero, Oriol Vinyals, Marc...
work page 2022
-
[7]
Paul Covington, Jay Adams, and Emre Sargin. 2016. Deep Neural Networks for YouTube Recommendations. In Proceedings of the 10th ACM Conference on Recommender Systems, Boston, MA, USA, September 15-19, 2016, Shilad Sen, Werner Geyer, Jill Freyne, and Pablo Castells (Eds.). ACM, 191–198
work page 2016
-
[8]
Wenqi Fan, Zihuai Zhao, Jiatong Li, Yunqing Liu, Xiaowei Mei, Yiqi Wang, Jiliang Tang, and Qing Li. 2023. Recommender Systems in the Era of Large Language Models (LLMs). CoRR (2023). arXiv:2307.02046
arXiv 2023
Show all 71 references
-
[9]
Ziwei Fan, Zhiwei Liu, Shelby Heinecke, Jianguo Zhang, Huan Wang, Caiming Xiong, and Philip S. Yu. 2023. Zero-shot Item-based Recommendation via Multi- task Product Knowledge Graph Pre-Training. (2023), 483–493
2023
-
[10]
Hui Fang, Danning Zhang, Yiheng Shu, and Guibing Guo. 2020. Deep Learning for Sequential Recommendation: Algorithms, Influential Factors, and Evaluations. ACM Trans. Inf. Syst. 39, 1 (2020), 10:1–10:42
2020
-
[11]
Jyotirmoy Gope and Sanjay Kumar Jain. 2017. A survey on solving cold start prob- lem in recommender systems. In 2017 International Conference on Computing, Communication and Automation (ICCCA). IEEE, 133–138
2017
-
[12]
Ruining He, Wang-Cheng Kang, and Julian J. McAuley. 2017. Translation- based Recommendation. In Proceedings of the Eleventh ACM Conference on Recommender Systems, RecSys 2017, Como, Italy, August 27-31, 2017, Paolo Cre- monesi, Francesco Ricci, Shlomo Berkovsky, and Alexander ...
2017
-
[13]
Ruining He and Julian J. McAuley. 2016. Fusing Similarity Models with Markov Chains for Sparse Sequential Recommendation. In IEEE 16th International Conference on Data Mining, ICDM 2016, December 12-15, 2016, Barcelona, Spain, Francesco Bonchi, Josep Domingo-Ferrer, Ricardo Ba...
2016
-
[14]
Ruining He and Julian J. McAuley. 2016. Ups and Downs: Modeling the Vi- sual Evolution of Fashion Trends with One-Class Collaborative Filtering. In Proceedings of the 25th International Conference on World Wide Web, WWW 2016, Montreal, Canada, April 11 - 15, 2016, Jacqueline B...
2016
-
[15]
Brown, Prafulla Dhariwal, Scott Gray, Chris Hallacy, Benjamin Mann, Alec Radford, Aditya Ramesh, Nick Ryder, Daniel M
Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom B. Brown, Prafulla Dhariwal, Scott Gray, Chris Hallacy, Benjamin Mann, Alec Radford, Aditya Ramesh, Nick Ryder, Daniel M. Ziegler, John Schulman, Dario Amodei, and Sam McCandlish...
2020 arXiv
-
[16]
Balázs Hidasi, Alexandros Karatzoglou, Linas Baltrunas, and Domonkos Tikk
-
[17]
Rae, Oriol Vinyals, and Laurent Sifre
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osin...
2022 arXiv
-
[18]
McAuley, and Wayne Xin Zhao
Yupeng Hou, Zhankui He, Julian J. McAuley, and Wayne Xin Zhao. 2023. Learn- ing Vector-Quantized Item Representation for Transferable Sequential Recom- menders. In Proceedings of the ACM WebConference 2023, WWW 2023, Austin, TX, USA, 30 April 2023 - 4 May 2023, Ying Ding, Jie ...
2023
-
[19]
Yupeng Hou, Shanlei Mu, Wayne Xin Zhao, Yaliang Li, Bolin Ding, and Ji-Rong Wen. 2022. Towards Universal Sequence Representation Learning for Recom- mender Systems. InKDD ’22: The 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Washington, DC, USA, August 14...
2022
-
[20]
Jin Huang, Wayne Xin Zhao, Hongjian Dou, Ji-Rong Wen, and Edward Y. Chang
-
[21]
Bowen Jin, Chen Gao, Xiangnan He, Depeng Jin, and Yong Li. 2020. Multi- behavior Recommendation with Graph Convolutional Networks. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval, SIGIR 2020, Virtual Event, Chi...
2020
-
[22]
Wang-Cheng Kang and Julian J. McAuley. 2018. Self-Attentive Sequential Rec- ommendation. In IEEE International Conference on Data Mining, ICDM 2018, Singapore, November 17-20, 2018. IEEE Computer Society, 197–206
2018
-
[23]
Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling Laws for Neural Language Models. CoRR (2020). arXiv:2001.08361
2020 arXiv
-
[24]
Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fa- had Shahbaz Khan, and Mubarak Shah
Salman H. Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fa- had Shahbaz Khan, and Mubarak Shah. 2022. Transformers in Vision: A Survey. ACM Comput. Surv. 54, 10s (2022), 200:1–200:41
2022
-
[25]
Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A
James Kirkpatrick, Razvan Pascanu, Neil C. Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell. 2016. Overcoming catastrophic for...
2016 arXiv
-
[26]
Taku Kudo and John Richardson. 2018. SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, EMNLP 2018: System Demonstrations, Bru...
2018
-
[27]
Siddique Latif, Aun Zaidi, Heriberto Cuayáhuitl, Fahad Shamshad, Moazzam Shoukat, and Junaid Qadir. 2023. Transformers in Speech Processing: A Survey. CoRR (2023). arXiv:2303.11607
2023 arXiv
-
[28]
Sara Latifi, Dietmar Jannach, and Andrés Ferraro. 2022. Sequential recommenda- tion: A study on transformers, nearest neighbors and sampled metrics. Inf. Sci. 609 (2022), 660–678
2022
-
[29]
Jing Li, Pengjie Ren, Zhumin Chen, Zhaochun Ren, Tao Lian, and Jun Ma. 2017. Neural Attentive Session-based Recommendation. In Proceedings of the 2017 ACM on Conference on Information and Knowledge Management, CIKM 2017, Singapore, November 06 - 10, 2017, Ee-Peng Lim, Marianne...
2017
-
[30]
Yang Li, Tong Chen, Peng-Fei Zhang, and Hongzhi Yin. 2021. Lightweight Self-Attentive Sequential Recommendation. In CIKM ’21: The 30th ACM International Conference on Information and Knowledge Management, Virtual Event, Queensland, Australia, November 1 - 5, 2021, Gianluca Dem...
2021
-
[31]
Blerina Lika, Kostas Kolomvatsos, and Stathes Hadjiefthymiades. 2014. Facing the cold start problem in recommender systems. Expert Syst. Appl. 41, 4 (2014), 2065–2073
2014
-
[32]
Bryan Lim, Sercan Ömer Arik, Nicolas Loeff, and Tomas Pfister. 2019. Temporal Fusion Transformers for Interpretable Multi-horizon Time Series Forecasting. CoRR (2019). arXiv:1912.09363
2019 arXiv
-
[33]
Yu, Julian J
Zhiwei Liu, Yongjun Chen, Jia Li, Philip S. Yu, Julian J. McAuley, and Caiming Xiong. 2021. Contrastive Self-supervised Sequential Recommendation with Robust Augmentation. CoRR (2021). arXiv:2108.06479
2021 arXiv
-
[34]
Jesús Lovón-Melgarejo, Laure Soulier, Karen Pinel-Sauvagnat, and Lynda Tamine. 2021. Studying Catastrophic Forgetting in Neural Ranking Mod- els. In Advances in Information Retrieval - 43rd European Conference on IR Research, ECIR 2021, Virtual Event, March 28 - April 1, 2021,...
2021
-
[35]
McAuley, Christopher Targett, Qinfeng Shi, and Anton van den Hengel
Julian J. McAuley, Christopher Targett, Qinfeng Shi, and Anton van den Hengel
-
[36]
Shanlei Mu, Yupeng Hou, Wayne Xin Zhao, Yaliang Li, and Bolin Ding. 2022. ID-Agnostic User Behavior Pre-training for Sequential Rec- ommendation. In Information Retrieval - 28th China Conference, CCIR 2022, Chongqing, China, September 16-18, 2022, Revised Selected Papers (Lect...
2022
-
[37]
Jianmo Ni, Jiacheng Li, and Julian J. McAuley. 2019. Justifying Recommenda- tions using Distantly-Labeled Reviews and Fine-Grained Aspects. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Na...
2019
-
[38]
Yu, Albert Y
Hao Peng, Renyu Yang, Zheng Wang, Jianxin Li, Lifang He, Philip S. Yu, Albert Y. Zomaya, and Rajiv Ranjan. 2022. Lime: Low-Cost and Incremental Learning for Dynamic Heterogeneous Information Networks. IEEE Trans. Computers 71, 3 (2022), 628–642
2022
-
[39]
Petrov and Craig Macdonald
Aleksandr V. Petrov and Craig Macdonald. 2022. A Systematic Review and Replicability Study of BERT4Rec for Sequential Recommendation. In RecSys ’22: Sixteenth ACM Conference on Recommender Systems, Seattle, WA, USA, September 18 - 23, 2022, Jennifer Golbeck, F. Maxwell Harper,...
2022
-
[40]
Ruihong Qiu, Zi Huang, and Hongzhi Yin. 2021. Memory Augmented Multi- Instance Contrastive Predictive Coding for Sequential Recommendation. In IEEE International Conference on Data Mining, ICDM 2021, Auckland, New Zealand, December 7-10, 2021, James Bailey, Pauli Miettinen, Yu...
2021
-
[41]
Dimitrios Rafailidis and Alexandros Nanopoulos. 2016. Modeling Users Pref- erence Dynamics and Side Information in Recommender Systems. IEEE Trans. Syst. Man Cybern. Syst. 46, 6 (2016), 782–792
2016
-
[42]
Steffen Rendle, Christoph Freudenthaler, and Lars Schmidt-Thieme. 2010. Fac- torizing personalized Markov chains for next-basket recommendation. In Proceedings of the 19th International Conference on World Wide Web, WWW 2010, Raleigh, North Carolina, USA, April 26-30, 2010, Mi...
2010
-
[43]
Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit
Jonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit. 2020. A Constructive Prediction of the Generalization Error Across Scales. (2020)
2020
-
[44]
Noveen Sachdeva, Mehak Preet Dhaliwal, Carole-Jean Wu, and Julian J. McAuley
-
[45]
Kyuyong Shin, Hanock Kwak, Kyung-Min Kim, Minkyu Kim, Young-Jin Park, Jisu Jeong, and Seungjae Jung. 2021. One4all User Representation for Recommender Systems in E-commerce. CoRR (2021). arXiv:2106.00573
2021 arXiv
-
[46]
Kyuyong Shin, Hanock Kwak, Su Young Kim, Max Nihlén Ramström, Jisu Jeong, Jung-Woo Ha, and Kyung-Min Kim. 2023. Scaling Law for Recommendation Models: Towards General-Purpose User Representations. In Thirty-Seventh AAAI Conference on Artificial Intelligence, AAAI 2023, Thirty-...
2023
-
[47]
Khoshgoftaar
Xiaoyuan Su and Taghi M. Khoshgoftaar. 2009. A Survey of Collaborative Filtering Techniques. Adv. Artif. Intell. 2009 (2009), 421425:1–421425:19
2009
-
[48]
Fei Sun, Jun Liu, Jian Wu, Changhua Pei, Xiao Lin, Wenwu Ou, and Peng Jiang
-
[49]
Ke Sun, Tieyun Qian, Tong Chen, Yile Liang, Quoc Viet Hung Nguyen, and Hongzhi Yin. 2020. Where to Go Next: Modeling Long- and Short-Term User Preferences for Point-of-Interest Recommendation. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thir...
2020
-
[50]
Yong Kiam Tan, Xinxing Xu, and Yong Liu. 2016. Improved Recurrent Neural Net- works for Session-based Recommendations. In Proceedings of the 1st Workshop on Deep Learning for Recommender Systems, DLRS@RecSys 2016, Boston, MA, USA, September 15, 2016, Alexandros Karatzoglou, Ba...
2016
-
[51]
Jiaxi Tang and Ke Wang. 2018. Personalized Top-N Sequential Recommendation via Convolutional Sequence Embedding. In Proceedings of the Eleventh ACM International Conference on WebSearch and Data Mining, WSDM 2018, Marina Del Rey, CA, USA, February 5-9, 2018, Yi Chang, Chengxia...
2018
-
[52]
Trinh Xuan Tuan and Tu Minh Phuong. 2017. 3D Convolutional Networks for Session-based Recommendation with Content Features. In Proceedings of the Eleventh ACM Conference on Recommender Systems, RecSys 2017, Como, Italy, August 27-31, 2017, Paolo Cremonesi, Francesco Ricci, Shl...
2017
-
[53]
Aäron van den Oord, Yazhe Li, and Oriol Vinyals. 2018. Representation Learning with Contrastive Predictive Coding. CoRR (2018). arXiv:1807.03748
2018 arXiv
-
[54]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. (2017), 5998–6008
2017
-
[55]
Chenyang Wang, Weizhi Ma, Chong Chen, Min Zhang, Yiqun Liu, and Shaoping Ma. 2023. Sequential Recommendation with Multiple Contrast Signals. ACM Trans. Inf. Syst. 41, 1 (2023), 11:1–11:27
2023
-
[56]
Shoujin Wang, Qi Zhang, Liang Hu, Xiuzhen Zhang, Yan Wang, and Charu Aggarwal. 2022. Sequential/Session-based Recommendations: Challenges, Ap- proaches, Applications and Opportunities. In SIGIR ’22: The 45th International ACM SIGIR Conference on Research and Development in Inf...
2022
-
[57]
Jiancan Wu, Xiang Wang, Xingyu Gao, Jiawei Chen, Hongcheng Fu, Tianyu Qiu, and Xiangnan He. 2022. On the Effectiveness of Sampled Softmax Loss for Item Recommendation. CoRR (2022). arXiv:2201.02327
2022 arXiv
-
[58]
Yiqing Wu, Ruobing Xie, Yongchun Zhu, Xiang Ao, Xin Chen, Xu Zhang, Fuzhen Zhuang, Leyu Lin, and Qing He. 2022. Multi-view Multi-behavior Contrastive Learning in Recommendation. In Database Systems for Advanced Applications - 27th International Conference, DASFAA 2022, Virtual...
2022
-
[59]
Xu Xie, Fei Sun, Zhaoyang Liu, Shiwen Wu, Jinyang Gao, Jiandong Zhang, Bolin Ding, and Bin Cui. 2022. Contrastive Learning for Sequential Recommendation. In 38th IEEE International Conference on Data Engineering, ICDE 2022, Kuala Lumpur, Malaysia, May 9-12, 2022. IEEE, 1259–1273
2022
-
[60]
Xinyang Yi, Ji Yang, Lichan Hong, Derek Zhiyuan Cheng, Lukasz Heldt, Aditee Kumthekar, Zhe Zhao, Li Wei, and Ed H. Chi. 2019. Sampling-bias-corrected neural modeling for large corpus item recommendations. In Proceedings of the 13th ACM Conference on Recommender Systems, RecSys...
2019
-
[61]
Zheng Yuan, Fajie Yuan, Yu Song, Youhua Li, Junchen Fu, Fei Yang, Yunzhu Pan, and Yongxin Ni. 2023. Where to Go Next for Recommender Systems? ID- vs. Modality-based Recommender Models Revisited. (2023), 2639–2649
2023
-
[62]
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer. 2022. Scaling Vision Transformers. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022. IEEE, 1204– 1213
2022
-
[63]
Shuai Zhang, Yi Tay, Lina Yao, Aixin Sun, and Jake An. 2019. Next item recom- mendation with self-attentive metric learning. InThirty-Third AAAI Conference on Artificial Intelligence, Vol. 9
2019
-
[64]
Sheng, Jiajie Xu, De- qing Wang, Guanfeng Liu, and Xiaofang Zhou
Tingting Zhang, Pengpeng Zhao, Yanchi Liu, Victor S. Sheng, Jiajie Xu, De- qing Wang, Guanfeng Liu, and Xiaofang Zhou. 2019. Feature-level Deeper Self-Attention Network for Sequential Recommendation. In Proceedings of the Twenty-Eighth International Joint Conference on Artific...
2019
-
[65]
Wayne Xin Zhao, Shanlei Mu, Yupeng Hou, Zihan Lin, Yushuo Chen, Xingyu Pan, Kaiyuan Li, Yujie Lu, Hui Wang, Changxin Tian, Yingqian Min, Zhichao Feng, Xinyan Fan, Xu Chen, Pengfei Wang, Wendi Ji, Yaliang Li, Xiaoling Wang, and Ji-Rong Wen. 2021. RecBole: Towards a Unified, Com...
2021
-
[66]
Kun Zhou, Hui Wang, Wayne Xin Zhao, Yutao Zhu, Sirui Wang, Fuzheng Zhang, Zhongyuan Wang, and Ji-Rong Wen. 2020. S3-Rec: Self-Supervised Learning for Sequential Recommendation with Mutual Information Maximization. In CIKM ’20: The 29th ACM International Conference on Informati...
2020
-
[2015]
Image-Based Recommendations on Styles and Substitutes. In Proceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval, Santiago, Chile, August 9-13, 2015, Ricardo Baeza-Yates, Mounia Lalmas, Alistair Moffat, and Berthier A. R...
2015
-
[2016]
Session-based Recommendations with Recurrent Neural Networks. (2016)
2016
-
[2018]
Improving Sequential Recommendation with Knowledge-Enhanced Mem- ory Networks. In The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, SIGIR 2018, Ann Arbor, MI, USA, July 08-12, 2018, Kevyn Collins-Thompson, Qiaozhu Mei, Brian D. Dav...
2018
-
[2019]
BERT4Rec: Sequential Recommendation with Bidirectional Encoder Rep- resentations from Transformer. In Proceedings of the 28th ACM International Conference on Information and Knowledge Management, CIKM 2019, Beijing, China, November 3-7, 2019, Wenwu Zhu, Dacheng Tao, Xueqi Chen...
2019
-
[2022]
Infinite Recommendation Networks: A Data-Centric Approach. (2022)
2022
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.