REVIEW 5 major objections 6 minor 13 cited by
Scaling New Frontiers: Insights into Large Recommendation Models
T0 review · 5 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper claims that HSTU, a transformer-based generative recommendation model, scales with depth in recall and ranking while GPT and SASRec do not, and that residual connections and relative attention bias are the source of the scaling…
desk verdict Useful but uneven empirical study: the HSTU ranking evaluation is new and worth having, but the central scaling-law contrast is under-evidenced. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the HSTU block, a transformer-style recommendation transducer that substitutes SiLU for Softmax in attention weighting, adds relative position and time-difference bucket biases to attention scores, and uses a residual connection placed around the pre-normalized sublayers plus a pointwise feature-interaction layer. The paper isolates the scaling law's origin by sweeping block count, ablating one component at a time, and surgically transferring the residual and relative-bias modules into SASRec, showing that these two modules, rather than raw parameter count, are what carry scalability.
What would settle it
Run the same 2-to-32 block depth sweep on a production-scale behavior log with sequence lengths above 1000 and embedding dimensions above 1000, comparing HSTU, Llama, GPT, and SASRec with and without the transplanted residual and relative-bias modules. If HSTU's recall or ranking curves plateau or decline while SASRec with the transplanted modules continues to improve, or if removing the relative attention bias no longer flattens the HSTU curve, the paper's claim about the origin and transferability of the scaling law would be falsified.
Extended reading notes
Core claim
The paper's central claim is that model depth produces a scaling law only for certain recommendation architectures. In controlled sweeps from 2 to 32 blocks on the public rating and review datasets, HSTU and Llama improve or hold their recall metrics, whereas GPT and SASRec degrade sharply, sometimes to near-random levels. Ablations that remove one component at a time show that the relative attention bias, especially the bucketed time-difference part, is what keeps the curve rising, and that the residual-connection pattern used by HSTU and Llama is more robust than SASRec's post-normalization residual; combining Llama-style residuals with the relative attention bias turns SASRec from non-scaling into scaling. The paper also presents the first evaluation of HSTU on ranking tasks, where it improves with depth while Llama does not, and where larger numbers of negative samples help rather than hurt, provided the model size and embedding dimension are matched to the dataset.
Load-bearing premise
The scaling trends measured on public datasets with embedding sizes between 50 and 400 and sequence lengths up to 200 transfer to trillion-parameter production systems with much longer user sequences; the paper never validates on a production-scale model or dataset.
Editorial extensions
If this is right
- Architecture choice determines whether adding depth helps at all: HSTU- and Llama-style backbones are the ones worth scaling vertically in recall, while GPT and SASRec backbones will not reward extra blocks.
- HSTU's scaling extends beyond recall: on ranking tasks it improves with depth, so the same backbone can serve both the recall and ranking stages of a recommendation pipeline.
- Legacy sequential recommenders can be made scalable by transplanting the residual connection pattern and relative attention bias, offering an upgrade path for existing systems without changing the whole architecture.
- Model size must be matched to dataset size: the optimal number of blocks shrinks as embedding size grows, and larger sequence lengths justify larger models, so a fixed large model is not universally best.
- Generative ranking models benefit from using more negative samples rather than aggressive subsampling, which suggests that curation for ranking should preserve negatives at scale.
Reading between the lines
- Going beyond the paper: if the depth-scaling curve is governed by the time-bucket attention bias rather than by parameter count, then models trained on temporally regular logs should scale faster than models on sparse, irregular logs; this can be tested by changing time-bucket granularity while holding depth fixed.
- The paper reports a near-constant product of optimal layers and embedding dimension but does not fit a power-law exponent; fitting the public-data curves to a two-exponent power law could give a practical rule for choosing depth and embedding size before training.
- Because GPT-style backbones fail to scale in recommendation despite scaling well in language, the missing ingredient may be recommendation-specific inductive biases; a direct test would be to add the same relative time-bias and residual pattern to a GPT-style backbone and rerun the block sweep.
- The negative-sampling result suggests that generative ranking models can exploit the full negative distribution, which would change how industrial click-through-rate datasets are curated if the effect transfers to production scale.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents an empirical study of scaling behavior in transformer-based recommendation models, with a focus on Meta's HSTU architecture. Using public datasets (ML-1M, ML-20M, AMZ-Books, CIKM, IJCAI, AMZ-MD), the authors compare HSTU, Llama, GPT, and SASRec as the number of attention blocks grows, concluding that HSTU and Llama exhibit better scalability while GPT and SASRec do not. They then ablate HSTU components (relative attention bias, SiLU activation, feature interaction), run parameter analyses over embedding dimension, number of heads, sequence length, and depth, and test HSTU on complex user behavior (side information, multi-behavior, multi-domain) and ranking tasks. The paper claims to identify residual connection patterns and relative attention bias as the origins of the scaling law and states that it is the first to evaluate HSTU on ranking tasks. The experiments are broad but rely on single-run evaluations and fixed hyperparameters for baselines, which weakens the central comparative claims.
Significance. If the conclusions were robustly supported, this would be a useful and timely contribution: the paper asks an important question about whether recommendation backbones scale with depth, provides a wide range of experiments across six datasets and multiple tasks, and includes ablations that isolate HSTU-specific components. The first evaluation of HSTU on ranking tasks is also practically relevant. A particular strength is the explicit attempt to attribute scaling behavior to observable architectural choices rather than treating HSTU as a black box. The paper states that supplementary code is available on GitHub, which is a valuable reproducibility commitment. However, the central comparative claim (HSTU and Llama scale, GPT and SASRec do not) is currently undermined by optimization confounds and by the absence of uncertainty quantification, so the significance is conditional on the core experiments being reworked.
major comments (5)
- [§5.2, Table 1] The conclusion that GPT and SASRec "show no scalability" conflates architectural scaling capacity with optimization stability under fixed hyperparameters. Section 5.1.4 states that original implementations' hyperparameters were kept fixed except for depth. Under this protocol, deeper GPT and SASRec runs collapse to near-random metrics (e.g., GPT on ML-1M drops from HR@10 0.2803 at 4 blocks to 0.0353 at 8 blocks; SASRec on ML-20M drops from 0.2781 at 2 blocks to 0.0599 at 4 blocks). Such a collapse is the classic signature of training instability under an unadjusted learning rate or warmup schedule, not evidence of an absent scaling law. The comparison should be repeated with per-depth hyperparameter tuning or with training-convergence diagnostics to separate architecture scaling from trainability.
- [§5.2, Table 1; §5.5.2, Table 14] The evidence that HSTU scales reliably is itself not statistically supported. Table 1 shows HSTU HR@10 on AMZ-Books falling from 0.0680 at 16 blocks to 0.0584 at 32 blocks, and on ML-1M the 16-to-32-block change is only 0.3322 to 0.3298. In the ranking task, Table 14 shows ML-20M AUC 0.7992 at 24 blocks dropping to 0.7914 at 32 blocks. Since no standard deviations, multiple seeds, or significance tests are reported anywhere in Section 5, the paper does not establish that HSTU scales more consistently than the baselines; several of the differences driving the qualitative conclusions are within a range that could easily be noise in single-run evaluations.
- [§5.3.2, Table 6] The claim that "the product of the optimal number of layers (L) and embedding dimension (D) remains constant, supporting our theoretical model that size is proportional to O(LD)" is not supported by the reported data. Reading the NDCG@10 columns for |S|=100, the apparent optimal depth is around 64 blocks for D=50 (LD=3200), 32 for D=100 (LD=3200), 24 for D=200 (LD=4800), and 32 for D=400 (LD=12800). The product is not constant, and the NDCG differences among neighboring depths are typically a few thousandths, so the notion of a well-defined optimum is fragile. This section should quantify the claimed relationship or reframe it as a qualitative observation.
- [§5.3.3, Table 8] The attribution of HSTU's scaling law to the residual connection pattern and relative attention bias is overreaching. The experiments modify SASRec by adopting the residual structure of HSTU or Llama and adding relative attention bias; these are precisely the architectural changes known to improve trainability of deep transformers, so the results are consistent with an optimization-stability explanation rather than a fundamental scaling-law origin. Moreover, even the best modified SASRec reaches only HR@10 0.3182 at 32 blocks on ML-20M, well below HSTU's 0.3569 (Table 1), so the modified model does not reproduce HSTU's scaling behavior. The wording should be softened to "factors associated with improved scaling" unless additional experiments single out the causal mechanism.
- [§5.5.3, Table 15] The text states that increasing the negative sampling ratio "leads to a continuous improvement in the model's performance," but the table is non-monotonic. For HSTU on ML-20M, AUC at ratio 0.4 (0.7952) is higher than at ratio 0.8 (0.7899) and at ratio 1.0 (0.7916); on AMZ-Books, AUC drops from 0.7338 at ratio 0.8 to 0.7037 at ratio 1.0. The claim should be revised to describe the observed pattern, or the experiment should be rerun with multiple seeds to determine whether the fluctuations are within noise.
minor comments (6)
- [§5.1.1, §5.1.4] The KuaiRand-27k dataset is listed in Section 5.1.1 but never appears in any reported experiment; it should either be used or removed from the list.
- [§5.1.4] The paper says that original implementations' hyperparameters were kept fixed, but it does not list the actual learning rate, schedule, batch size, warmup, or dropout for each backbone; a reproducibility appendix should report these details.
- [§5.3.4, Figure 4] The t-SNE visualization is used to infer that "a well-normalized model may enhance both scalability and recall performance," but no quantitative measure of clustering or normalization is provided, so the conclusion is not supported by the figure alone.
- [§5.5.1] The notation for the ranking task is confusing: the input sequence uses b′_i for labels, but the loss uses both b″_k and b′_k; the text should clearly define the ground-truth label and the prediction (e.g., y_i and ŷ_i).
- [Table 2] The table caption says "Impact of various HSTU components on scaling law" but the row labels use the abbreviation "r.a.b." without defining it in the caption or text; the abbreviation should be expanded at first use.
- [§6.3 and References] There is a duplicated in-text citation "[2, 40, 48, 48, 66, 68]" for efficiency optimizations, and reference [99] has missing venue and year details; these should be corrected.
Circularity Check
No significant circularity: the central claims are empirical measurements on public benchmarks, with no parameter fitted and renamed as a prediction.
full rationale
The paper's central claims concern how recall and ranking performance change as the number of transformer blocks increases for HSTU, Llama, GPT, and SASRec. These claims are supported by reported measurements in Tables 1, 8, and 14 on public datasets, rather than by a derivation in which one quantity is defined in terms of another. Section 5.1.4 states that original model hyperparameters were kept fixed except for depth; this makes the depth-collapse of GPT and SASRec an experimental outcome, not a quantity forced by construction. The scaling-law origin analysis in Section 5.3.3 modifies SASRec by adding relative attention bias and HSTU/Llama-style residual patterns, then measures the effect; the conclusion is induced from Table 8 rather than presupposed. The reliance on HSTU [126] is an external citation to prior work by other authors, and no uniqueness theorem or ansatz is imported through self-citation. Any concern that fixed hyperparameters make the GPT/SASRec comparison an optimization-collapse artifact is a validity or correctness question, not a circularity in the sense defined here. The paper is therefore self-contained against external benchmarks, and no circular step can be exhibited.
Assumptions & free parameters
free parameters (2)
- Per-dataset embedding dimension =
ML-1M: 50, ML-20M: 256, AMZ-Books: 64; reduced to 4 in Table 17
- Maximum sequence length =
up to 200 for parameter analyses (|S| = 100 and 200 in Table 6)
assumptions (4)
- domain assumption Public benchmark datasets such as MovieLens, Amazon, CIKM, and IJCAI are representative of large-scale industrial recommendation data.
- domain assumption The reproduced HSTU model faithfully corresponds to the original HSTU of [126] and the ranking head does not distort the base model.
- domain assumption Varying only the number of attention blocks while fixing all other hyperparameters isolates the effect of model depth on scaling.
- ad hoc to paper The optimal model size is proportional to O(LD), where L is the number of layers and D is the embedding dimension.
Cite this review
Pith. "Pith review of Scaling New Frontiers: Insights into Large Recommendation Models." pith.science (2026). https://pith.science/paper/SXLQVYXJ
@misc{pith2026241200714,
author = {Pith},
title = {Pith review of: Scaling New Frontiers: Insights into Large Recommendation Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/SXLQVYXJ}},
note = {Machine review of arXiv:2412.00714}
}
read the original abstract
Recommendation systems are essential for filtering data and retrieving relevant information across various applications. Recent advancements have seen these systems incorporate increasingly large embedding tables, scaling up to tens of terabytes for industrial use. However, the expansion of network parameters in traditional recommendation models has plateaued at tens of millions, limiting further benefits from increased embedding parameters. Inspired by the success of large language models (LLMs), a new approach has emerged that scales network parameters using innovative structures, enabling continued performance improvements. A significant development in this area is Meta's generative recommendation model HSTU, which illustrates the scaling laws of recommendation systems by expanding parameters to thousands of billions. This new paradigm has achieved substantial performance gains in online experiments. In this paper, we aim to enhance the understanding of scaling laws by conducting comprehensive evaluations of large recommendation models. Firstly, we investigate the scaling laws across different backbone architectures of the large recommendation models. Secondly, we conduct comprehensive ablation studies to explore the origins of these scaling laws. We then further assess the performance of HSTU, as the representative of large recommendation models, on complex user behavior modeling tasks to evaluate its applicability. Notably, we also analyze its effectiveness in ranking tasks for the first time. Finally, we offer insights into future directions for large recommendation models. Supplementary materials for our research are available on GitHub at https://github.com/USTC-StarTeam/Large-Recommendation-Models.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 13 Pith papers
-
From Scaling to Structured Expressivity: Rethinking Transformers for CTR Prediction
FAT specializes attention by semantic field and reports +0.51% AUC over baselines on Taobao data, but its power-law scaling law is an empirical fit, not a derived prediction.
-
FuXi-\beta: Towards a Lightweight and Fast Large-Scale Generative Recommendation Model
FuXi-β shows that removing query-key attention and using a functional relative time bias makes generative recommendation Transformers faster and, on industrial datasets, more accurate.
-
RankMixer: Scaling Up Ranking Models in Industrial Recommenders
RankMixer scales an industrial ranking model to 1B dense parameters with 10x MFU improvement and unchanged latency, gaining 1.08% in app duration in Douyin A/B tests.
-
Thought-Augmented Planning for LLM-Powered Interactive Recommender Agent
TAIRA, a thought-pattern-augmented multi-agent recommender, outperforms prior LLM agents in simulated interactive recommendation, with the largest gains on complex user intents.
-
TD3: Tucker Decomposition Based Dataset Distillation Method for Sequential Recommendation
TD3 factorizes a synthetic sequence summary into user, time, item, and core factors via Tucker decomposition, and trains recommenders on this summary with a feature-alignment meta-objective.
-
Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens
The middle layers of LVLMs process visual information in two stages, and amplifying image attention in the first 'enrichment' stage reduces object hallucinations.
-
Number it: Temporal Grounding Videos like Flipping Manga
Number-Prompt overlays frame numbers on video frames, improving temporal grounding in video LLMs and setting new state-of-the-art results on moment retrieval and highlight detection.
-
Why Thinking Hurts: Diagnosing and Rectifying Linguistic Inertia in Large Language Models for Recommendation
Chain-of-thought reasoning degrades semantic-ID recommendation accuracy through 'linguistic inertia,' and a training-free compression-plus-contrastive decoding fix restores and often improves accuracy.
-
DLF: Enhancing Explicit-Implicit Interaction via Dynamic Low-Order-Aware Fusion for CTR Prediction
DLF is a CTR prediction architecture that combines low-rank, high-rank, and implicit interaction blocks with layer-wise attention fusion, reporting state-of-the-art results on Criteo, Avazu, Movielens, and Frappe.
-
Towards Large-scale Generative Ranking
Generative ranking with action-oriented sequences and lightweight position and time biases improves user metrics at comparable inference cost to a production recommender.
-
Killing Two Birds with One Stone: Unifying Retrieval and Ranking with a Single Generative Recommendation Model
UniGRF trains one generative model to output both next-item predictions and click probabilities, and reports consistent gains over separate retrieval and ranking models on MovieLens and Amazon-Books.
-
FuXi-$\alpha$: Scaling Recommendation Model with Feature Interaction Enhanced Transformer
FuXi-alpha, a sequential recommender with decoupled temporal, positional, and semantic attention channels plus a two-stage FFN, reports gains over HSTU and positive online engagement results.
-
Climber: Toward Efficient Scaling Laws for Large Recommendation Models
Climber reports that splitting user sequences by behavior type, adding adaptive temperature, and co-designed batching enable more efficient Transformer scaling in recommender systems.
Reference graph
Works this paper leans on
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Flo- rencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shya- mal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)
arXiv 2023
-
[2]
Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwa- tra, Bhargav S Gulavani, Alexey Tumanov, and Ramachandran Ramjee. 2024. Taming throughput-latency tradeoff in llm inference with sarathi-serve. arXiv preprint arXiv:2403.02310 (2024)
arXiv 2024
-
[3]
Peter F Brown, Vincent J Della Pietra, Peter V Desouza, Jennifer C Lai, and Robert L Mercer. 1992. Class-based n-gram models of natural language. Com- putational linguistics 18, 4 (1992), 467–480
1992
-
[4]
Jie Cao, Xinyu Cong, Jianxin Sheng, et al . 2022. Contrastive Cross-Domain Sequential Recommendation. In Proceedings of the 31st ACM International Con- ference on Information & Knowledge Management . 138–147
2022
-
[5]
Rose Catherine and William Cohen. 2017. Transnets: Learning to transform for recommendation. In Proceedings of the 11th ACM Conference on Recommender Systems. 288–296
2017
-
[6]
Jianxin Chang, Chenbin Zhang, Yiqun Hui, Dewei Leng, Yanan Niu, Yang Song, and Kun Gai. 2023. Pepnet: Parameter and embedding personalized network for infusing with personalized prior information. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 3795–3804
2023
-
[7]
Chong Chen, Min Zhang, Yiqun Liu, and Shaoping Ma. 2018. Neural attentional rating regression with review-level explanations. In Proceedings of the 2018 World Wide Web conference. 1583–1592
2018
-
[8]
Qiwei Chen, Changhua Pei, Shanshan Lv, Chao Li, Junfeng Ge, and Wenwu Ou
Show all 146 references
-
[9]
2018.{TVM}: An automated{End-to-End} optimizing compiler for deep learning
Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, et al. 2018.{TVM}: An automated{End-to-End} optimizing compiler for deep learning. In 13th USENIX Symposium on Operating Systems Design and Implem...
2018
-
[10]
Xu Chen, Hongteng Xu, Yongfeng Zhang, Jiaxi Tang, Yixin Cao, Zheng Qin, and Hongyuan Zha. 2018. Sequential recommendation with user memory networks. In Proceedings of the eleventh ACM international conference on web search and data mining. 108–116
2018
-
[11]
Jin Yao Chin, Yile Chen, and Gao Cong. 2022. The Datasets Dilemma: How Much Do We Really Know About Recommendation Datasets?. In Proceedings of the Fifteenth ACM International Conference on Web Search and Data Mining . 141–149
2022
-
[12]
Jin Yao Chin, Kaiqi Zhao, Shafiq Joty, and Gao Cong. 2018. ANR: Aspect-based Neural Recommender. In Proceedings of the 27th ACM International Conference on Information and Knowledge Management . 147–156
2018
-
[13]
Sung Min Cho, Eunhyeok Park, and Sungjoo Yoo. 2020. MEANTIME: Mixture of Attention Mechanisms with Multi-temporal Embeddings for Sequential Recommendation. In Proceedings of the 14th ACM Conference on Recommender Systems. 515–520
2020
-
[14]
Tri Dao. 2023. Flashattention-2: Faster attention with better parallelism and work partitioning. arXiv preprint arXiv:2307.08691 (2023)
2023 arXiv
-
[15]
Chen Gao, Xiangnan He, Dahua Gan, Xiangning Chen, Fuli Feng, Yong Li, Tat-Seng Chua, and Depeng Jin. 2019. Neural Multi-task Recommendation from Multi-behavior Data. In 2019 IEEE 35th International Conference on Data Engineering. IEEE, 1554–1557
2019
-
[16]
Chongming Gao, Shijun Li, Yuan Zhang, Jiawei Chen, Biao Li, Wenqiang Lei, Peng Jiang, and Xiangnan He. 2022. Kuairand: an unbiased sequential rec- ommendation dataset with randomly exposed videos. In Proceedings of the 31st ACM International Conference on Information & Knowled...
2022
-
[17]
Aditya Grover and Jure Leskovec. 2016. node2vec: Scalable feature learning for networks. In Proceedings of the 22nd ACM SIGKDD international conference on Knowledge discovery and data mining . 855–864
2016
-
[18]
Yulong Gu, Zhuoye Ding, Shuaiqiang Wang, Lixin Zou, Yiding Liu, and Dawei Yin. 2020. Deep multifaceted transformers for multi-objective ranking in large- scale e-commerce recommender systems. In Proceedings of the 29th ACM Inter- national Conference on Information and Knowledg...
2020
-
[19]
Huifeng Guo, Wei Guo, Yong Gao, Ruiming Tang, Xiuqiang He, and Wenzhi Liu
-
[20]
Huifeng Guo, Ruiming Tang, Yunming Ye, Zhenguo Li, and Xiuqiang He. 2017. DeepFM: a factorization-machine based neural network for CTR prediction. arXiv preprint arXiv:1703.04247 (2017)
2017 arXiv
-
[21]
In Proceedings of the 44th international ACM SIGIR conference on research and development in information retrieval
Scalefreectr: Mixcache-based distributed training system for ctr models with huge embedding table. In Proceedings of the 44th international ACM SIGIR conference on research and development in information retrieval . 1269–1278
-
[22]
Wei Guo, Chang Meng, Enming Yuan, Zhicheng He, Huifeng Guo, Yingxue Zhang, Bo Chen, Yaochen Hu, Ruiming Tang, Xiu Li, et al. 2023. Compressed interaction graph based framework for multi-behavior recommendation. In Proceedings of the ACM Web Conference 2023 . 960–970
2023
-
[23]
Long Guo, Lifeng Hua, Rongfei Jia, Binqiang Zhao, Xiaobo Wang, and Bin Cui. 2019. Buying or browsing?: Predicting real-time purchasing intent using attention-based deep network with multiple behavior. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge ...
2019
-
[24]
Xingzhuo Guo, Junwei Pan, Ximei Wang, Baixu Chen, Jie Jiang, and Mingsheng Long. 2023. On the Embedding Collapse when Scaling up Recommendation Models. arXiv preprint arXiv:2310.04400 (2023)
2023 arXiv
-
[25]
Wei Guo, Rong Su, Renhao Tan, Huifeng Guo, Yingxue Zhang, Zhirong Liu, Ruiming Tang, and Xiuqiang He. 2021. Dual graph enhanced embedding neural network for CTR prediction. In Proceedings of the 27th ACM SIGKDD conference on knowledge discovery & data mining . 496–504
2021
-
[26]
Yongqiang Han, Likang Wu, Hao Wang, Guifeng Wang, Mengdi Zhang, Zhi Li, Defu Lian, and Enhong Chen. 2023. Guesr: A global unsupervised data- enhancement with bucket-cluster sampling for sequential recommendation. In International Conference on Database Systems for Advanced App...
2023
-
[27]
Yongqiang Han, Hao Wang, Kefan Wang, Likang Wu, Zhi Li, Wei Guo, Yong Liu, Defu Lian, and Enhong Chen. 2024. Efficient Noise-Decoupling for Multi- Behavior Sequential Recommendation. In Proceedings of the ACM on Web Con- ference 2024. 3297–3306
2024
-
[28]
Xiangnan He, Kuan Deng, Xiang Wang, Yan Li, Yongdong Zhang, and Meng Wang. 2020. LightGCN: Simplifying and Powering Graph Convolution Net- work for Recommendation. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieva...
2020
-
[29]
F Maxwell Harper and Joseph A Konstan. 2015. The movielens datasets: History and context. ACM Transactions on Interactive Intelligent Systems 5, 4 (2015), 1–19
2015
-
[30]
Zhicheng He, Weiwen Liu, Wei Guo, Jiarui Qin, Yingxue Zhang, Yaochen Hu, and Ruiming Tang. 2023. A survey on user behavior modeling in recommender systems. arXiv preprint arXiv:2302.11087 (2023)
2023 arXiv
-
[31]
Xiangnan He, Lizi Liao, Hanwang Zhang, Liqiang Nie, Xia Hu, and Tat-Seng Chua. 2017. Neural Collaborative Filtering. In Proceedings of the 26th Interna- tional Conference on World Wide Web. 173–182
2017
-
[32]
Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami. 2024. Kvquant: Towards 10 million context length llm inference with kv cache quantization. arXiv preprint arXiv:2401.18079 (2024)
2024 arXiv
-
[33]
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. 2022. An empirical analysis of compute-optimal large language model training. Advances in Neural In...
2022
-
[34]
Yupeng Hou, Shanlei Mu, Wayne Xin Zhao, Yaliang Li, Bolin Ding, and Ji- Rong Wen. 2022. Towards universal sequence representation learning for recommender systems. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 585–593
2022
-
[35]
Yupeng Hou, Zhankui He, Julian McAuley, and Wayne Xin Zhao. 2023. Learning vector-quantized item representation for transferable sequential recommenders. In Proceedings of the ACM Web Conference 2023 . 1162–1171
2023
-
[36]
Yukun Huang, Yanda Chen, Zhou Yu, and Kathleen McKeown. 2022. In-context learning distillation: Transferring few-shot learning ability of pre-trained lan- guage models. arXiv preprint arXiv:2212.10670 (2022)
2022 arXiv
-
[37]
Tinglin Huang, Yuxiao Dong, Ming Ding, Zhen Yang, Wenzheng Feng, Xinyu Wang, and Jie Tang. 2021. Mixgcf: An improved training method for graph neural network-based recommender systems. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining . 665–674
2021
-
[38]
Herve Jegou, Matthijs Douze, and Cordelia Schmid. 2010. Product quantization for nearest neighbor search. IEEE transactions on pattern analysis and machine intelligence 33, 1 (2010), 117–128
2010
-
[39]
Dmytro Ivchenko, Dennis Van Der Staay, Colin Taylor, Xing Liu, Will Feng, Rahul Kindi, Anirudh Sudarshan, and Shahin Sefati. 2022. Torchrec: a pytorch domain library for recommendation systems. In Proceedings of the 16th ACM Conference on Recommender Systems . 482–483
2022
-
[40]
Yunho Jin, Chun-Feng Wu, David Brooks, and Gu-Yeon Wei. 2023. S3: Increasing GPU Utilization during Generative Inference for Higher Throughput. Advances in Neural Information Processing Systems 36 (2023), 18015–18027
2023
-
[41]
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling up visual and vision- language representation learning with noisy text supervision. In International conference on machine learning . PMLR, 4904–4916
2021
-
[42]
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models.arXiv preprint arXiv:2001.08361 (2020)
2020 arXiv
-
[43]
Kang and J
W. Kang and J. McAuley. 2018. Self-Attentive Sequential Recommendation. In 2018 IEEE International Conference on Data Mining . IEEE Computer Society, 197–206
2018
-
[44]
Thomas N Kipf and Max Welling. 2016. Variational graph auto-encoders. arXiv preprint arXiv:1611.07308 (2016)
2016 arXiv
-
[45]
Donghyun Kim, Chanyoung Park, Jinoh Oh, Sungyoung Lee, and Hwanjo Yu. 2016. Convolutional matrix factorization for document context-aware recommendation. In Proceedings of the 10th ACM Conference on Recommender Systems. 233–240
2016
-
[46]
Garima Koushik, K Rajeswari, and Suresh Kannan Muthusamy. 2019. Auto- mated hate speech detection on Twitter. In 2019 5th International Conference Guo and Wang, et al. On Computing, Communication, Control And Automation . IEEE, 1–4
2019
-
[47]
Yehuda Koren. 2008. Factorization meets the neighborhood: a multifaceted col- laborative filtering model. In Proceedings of the 14th ACM SIGKDD international conference on Knowledge discovery and data mining . 426–434
2008
-
[48]
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th Symposium on Operating Systems Princip...
2023
-
[49]
Taku Kudo. 2018. Subword regularization: Improving neural network translation models with multiple subword candidates. arXiv preprint arXiv:1804.10959 (2018)
2018 arXiv
-
[50]
Doyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho, and Wook-Shin Han. 2022. Autoregressive image generation using residual quantization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 11523– 11532
2022
-
[51]
Riwei Lai, Li Chen, Rui Chen, and Chi Zhang. 2024. A Survey on Data-Centric Recommender Systems. arXiv preprint arXiv:2401.17878 (2024)
2024 arXiv
-
[52]
Jiacheng Li, Ming Wang, Jin Li, Jinmiao Fu, Xin Shen, Jingbo Shang, and Julian McAuley. 2023. Text is all you need: Learning language representations for sequential recommendation. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 1258–1267
2023
-
[53]
Cheng Li, Ming Zhao, Haoyan Zhang, et al. 2022. RecGURU: Adversarial Learn- ing of Generalized User Representations for Cross-domain Recommendation. In Proceedings of the Fifteenth ACM International Conference on Web Search and Data Mining (WSDM). 571–581
2022
-
[54]
Mu Li, David G Andersen, Jun Woo Park, Alexander J Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J Shekita, and Bor-Yiing Su. 2014. Scal- ing distributed machine learning with the parameter server. In 11th USENIX Symposium on operating systems design and implementatio...
2014
-
[55]
Jiacheng Li, Yujie Wang, and Julian McAuley. 2020. Time Interval Aware Self-Attention for Sequential Recommendation. In Proceedings of the 13th Inter- national Conference on Web Search and Data Mining . 322–330
2020
-
[56]
Chen Liang, Simiao Zuo, Qingru Zhang, Pengcheng He, Weizhu Chen, and Tuo Zhao. 2023. Less is more: Task-aware layer-wise distillation for language model compression. In International Conference on Machine Learning . PMLR, 20852–20867
2023
-
[57]
Xiucheng Li, Jin Yao Chin, Yile Chen, and Gao Cong. 2021. Sinkhorn Collabora- tive Filtering. In Proceedings of the Web Conference 2021 . 582–592
2021
-
[58]
Weiwen Liu, Wei Guo, Yong Liu, Ruiming Tang, and Hao Wang. 2023. User Behavior Modeling with Deep Learning for Recommendation: Recent Advances. In Proceedings of the 17th ACM Conference on Recommender Systems. 1286–1287
2023
-
[59]
Weilin Lin, Xiangyu Zhao, Yejing Wang, Yuanshao Zhu, and Wanyu Wang
-
[60]
H. Ma, R. Xie, L. Meng, et al. 2024. Triple Sequence Learning for Cross-Domain Recommendation. ACM Transactions on Information Systems 42, 4 (2024), 1–29. To appear
2024
-
[61]
Xinyin Ma, Gongfan Fang, and Xinchao Wang. 2023. Llm-pruner: On the structural pruning of large language models. Advances in neural information processing systems 36 (2023), 21702–21720
2023
-
[62]
Z Liu, Y Hou, and J McAuley. 2024. Multi-Behavior Generative Recommendation. arXiv preprint arXiv:2405.16871 (2024)
2024 arXiv
-
[63]
Brendan McMahan. 2011. Follow-the-regularized-leader and mirror descent: Equivalence theorems and l1 regularization. In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics . JMLR Workshop and Conference Proceedings, 525–533
2011
-
[64]
Chang Meng, Hengyu Zhang, Wei Guo, Huifeng Guo, Haotian Liu, Yingxue Zhang, Hongkun Zheng, Ruiming Tang, Xiu Li, and Rui Zhang. 2023. Hierar- chical projection enhanced multi-behavior recommendation. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and D...
2023
-
[65]
Julian McAuley and Jure Leskovec. 2013. Hidden factors and hidden topics: understanding rating dimensions with review text. In Proceedings of the 7th ACM conference on Recommender systems . 165–172
2013
-
[66]
Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Zhengxin Zhang, Rae Ying Yee Wong, Alan Zhu, Lijie Yang, Xiaoxiang Shi, et al. 2023. SpecInfer: Accelerating Generative Large Language Model Serv- ing with Tree-based Speculative Inference and Verification. a...
2023 arXiv
-
[67]
T Mikolov. 2013. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781 (2013)
2013 arXiv
-
[68]
Chang Meng, Ziqi Zhao, Wei Guo, Yingxue Zhang, Haolun Wu, Chen Gao, Dong Li, Xiu Li, and Ruiming Tang. 2023. Coarse-to-fine knowledge-enhanced multi-interest learning framework for multi-behavior recommendation. ACM Transactions on Information Systems 42, 1 (2023), 1–27
2023
-
[69]
Nikil Pancha, Andrew Zhai, Jure Leskovec, and Charles Rosenberg. 2022. Pinner- former: Sequence modeling for user representation at pinterest. In Proceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining . 3702–3712
2022
-
[70]
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gre- gory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing syst...
2019
-
[71]
ModelTC. 2024. Lightllm. https://github.com/ModelTC/lightllm. [Online]
2024
-
[72]
Qi Pi, Weijie Bian, Guorui Zhou, Xiaoqiang Zhu, and Kun Gai. 2019. Practice on long sequential user behavior modeling for click-through rate prediction. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . 2671–2679
2019
-
[73]
Qi Pi, Guorui Zhou, Yujing Zhang, Zhe Wang, Lejian Ren, Ying Fan, Xiaoqiang Zhu, and Kun Gai. 2020. Search-based user interest modeling with lifelong sequential behavior data for click-through rate prediction. In Proceedings of the 29th ACM International Conference on Informat...
2020
-
[74]
Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014. Glove: Global vectors for word representation. In Proceedings of the 2014 conference on empirical methods in natural language processing . 1532–1543
2014
-
[75]
Yuhan Quan, Jingtao Ding, Chen Gao, Lingling Yi, Depeng Jin, and Yong Li. 2023. Robust preference-guided denoising for graph based social recommendation. In Proceedings of the ACM Web Conference 2023 . 1097–1108
2023
-
[76]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sand- hini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al
-
[77]
Jiarui Qin, Weinan Zhang, Rong Su, Zhirong Liu, Weiwen Liu, Guangpeng Zhao, Hao Li, Ruiming Tang, Xiuqiang He, and Yong Yu. 2023. Learning to retrieve user behaviors for click-through rate estimation. ACM Transactions on Information Systems 41, 4 (2023), 1–31
2023
-
[78]
Shashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan Hulikal Keshavan, Trung Vu, Lukasz Heldt, Lichan Hong, Yi Tay, Vinh Tran, Jonah Samost, et al
-
[79]
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Rad- ford, Mark Chen, and Ilya Sutskever. 2021. Zero-shot text-to-image generation. In International conference on machine learning . Pmlr, 8821–8831
2021
-
[80]
In International conference on machine learning
Learning transferable visual models from natural language supervision. In International conference on machine learning . PMLR, 8748–8763
-
[81]
Nived Rajaraman, Jiantao Jiao, and Kannan Ramchandran. 2024. Toward a Theory of Tokenization in LLMs. arXiv preprint arXiv:2404.08335 (2024)
2024 arXiv
-
[82]
Mike Schuster and Kaisuke Nakajima. 2012. Japanese and korean voice search. In 2012 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 5149–5152
2012
-
[83]
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2015. Neural machine translation of rare words with subword units. arXiv preprint arXiv:1508.07909 (2015)
2015 arXiv
-
[84]
Alexander Sergeev and Mike Del Balso. 2018. Horovod: fast and easy distributed deep learning in TensorFlow. arXiv preprint arXiv:1802.05799 (2018)
2018 arXiv
-
[85]
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020. Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. InProceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . 3505–3506
2020
-
[86]
Kan Ren, Jiarui Qin, Yuchen Fang, Weinan Zhang, Lei Zheng, Weijie Bian, Guorui Zhou, Jian Xu, Yong Yu, Xiaoqiang Zhu, et al. 2019. Lifelong sequential modeling with personalized memorization for user response prediction. In Proceedings of the 42nd International ACM SIGIR Confe...
2019
-
[87]
Fei Sun, Jun Liu, Jian Wu, Changhua Pei, Xiao Lin, Wenwu Ou, and Peng Jiang
-
[88]
Mingjie Sun, Zhuang Liu, Anna Bair, and J Zico Kolter. 2023. A simple and effec- tive pruning approach for large language models.arXiv preprint arXiv:2306.11695 (2023)
2023 arXiv
-
[89]
Zhen Tian, Ting Bai, Zibin Zhang, Zhiyuan Xu, Kangyi Lin, Ji-Rong Wen, and Wayne Xin Zhao. 2023. Directed acyclic graph factorization machines for CTR prediction via knowledge distillation. In Proceedings of the Sixteenth ACM International Conference on Web Search and Data Min...
2023
-
[90]
Tingjia Shen, Hao Wang, Jiaqing Zhang, Sirui Zhao, Liangyue Li, Zulong Chen, Defu Lian, and Enhong Chen. 2024. Exploring User Retrieval Integration towards Large Language Models for Cross-Domain Sequential Recommendation. arXiv preprint arXiv:2406.03085 (2024)
2024 arXiv
-
[91]
Jiajie Su, Chaochao Chen, Zibin Lin, Xi Li, Weiming Liu, and Xiaolin Zheng
-
[92]
In Proceedings of the 31st ACM International Conference on Multimedia
Personalized Behavior-Aware Transformer for Multi-Behavior Sequential Recommendation. In Proceedings of the 31st ACM International Conference on Multimedia
-
[93]
Aaron Van Den Oord, Oriol Vinyals, et al. 2017. Neural discrete representation learning. Advances in neural information processing systems 30 (2017)
2017
-
[94]
Hao Wang, Yongqiang Han, Kefan Wang, Kai Cheng, Zhen Wang, Wei Guo, Yong Liu, Defu Lian, and Enhong Chen. 2024. Denoising Pre-Training and Customized Scaling New Frontiers: Insights into Large Recommendation Models Prompt Learning for Efficient Multi-Behavior Sequential Recomm...
2024 arXiv
-
[95]
Hao Wang, Defu Lian, Hanghang Tong, Qi Liu, Zhenya Huang, and Enhong Chen. 2021. Decoupled representation learning for attributed networks. IEEE Transactions on Knowledge and Data Engineering 35, 3 (2021), 2430–2444
2021
-
[96]
Hao Wang, Defu Lian, Hanghang Tong, Qi Liu, Zhenya Huang, and Enhong Chen. 2021. Hypersorec: Exploiting hyperbolic user and item representations with multiple aspects for social-aware recommendation. ACM Transactions on Information Systems 40, 2 (2021), 1–28
2021
-
[97]
Tianchi. 2018. IJCAI-15 Repeat Buyers Prediction Dataset. https://tianchi. aliyun.com/dataset/dataDetail?dataId=42
2018
-
[98]
Junxiong Tong, Mingjia Yin, Hao Wang, Qiushi Pan, Defu Lian, and Enhong Chen. 2024. MDAP: A Multi-view Disentangled and Adaptive Preference Learning Framework for Cross-Domain Recommendation. arXiv preprint arXiv:2410.05877 (2024)
2024 arXiv
-
[99]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023)
2023 arXiv
-
[100]
Qinyong Wang, Hongzhi Yin, Hao Wang, Quoc Viet Hung Nguyen, Zi Huang, and Lizhen Cui. 2019. Enhancing collaborative filtering with generative aug- mentation. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . 548–556
2019
-
[101]
Ruoxi Wang, Bin Fu, Gang Fu, and Mingliang Wang. 2017. Deep & cross network for ad click predictions. In Proceedings of the ADKDD’17 . 1–7
2017
-
[102]
Wenjie Wang, Fuli Feng, Xiangnan He, Liqiang Nie, and Tat-Seng Chua. 2021. Denoising implicit feedback for recommendation. In Proceedings of the 14th ACM international conference on web search and data mining . 373–381
2021
-
[103]
Zehuan Wang, Yingcan Wei, Minseok Lee, Matthias Langer, Fan Yu, Jie Liu, Shijie Liu, Daniel G Abel, Xu Guo, Jianbing Dong, et al. 2022. Merlin hugeCTR: GPU-accelerated recommender system training and inference. In Proceedings of the 16th ACM Conference on Recommender Systems . 534–537
2022
-
[104]
Huanyu Wang, Junjie Liu, Xin Ma, Yang Yong, Zhenhua Chai, and Jianxin Wu. 2022. Compressing models with few samples: Mimicking then replacing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 701–710
2022
-
[105]
Hao Wang, Tong Xu, Qi Liu, Defu Lian, Enhong Chen, Dongfang Du, Han Wu, and Wen Su. 2019. MCNE: An end-to-end framework for learning multiple conditional network representations of social network. InProceedings of the 25th ACM SIGKDD international conference on knowledge disco...
2019
-
[106]
Hao Wang, Mingjia Yin, Luankang Zhang, Sirui Zhao, and Enhong Chen. [n. d.]. MF-GSLAE: A Multi-Factor User Representation Pre-training Framework for Dual-Target Cross-Domain Recommendation. ACM Transactions on Information Systems ([n. d.])
-
[107]
Likang Wu, Zhi Zheng, Zhaopeng Qiu, Hao Wang, Hongchao Gu, Tingjia Shen, Chuan Qin, Chen Zhu, Hengshu Zhu, Qi Liu, et al. 2024. A Survey on Large Language Models for Recommendation. World Wide Web 27, 5 (2024), 60
2024
-
[108]
Yongji Wu, Defu Lian, Neil Zhenqiang Gong, Lu Yin, Mingyang Yin, Jingren Zhou, and Hongxia Yang. 2021. Linear-time self attention with codeword histogram for efficient recommendation. In Proceedings of the Web Conference
2021
-
[109]
Yiqing Wu, Ruobing Xie, Yongchun Zhu, Xiang Ao, Xin Chen, Xu Zhang, Fuzhen Zhuang, Leyu Lin, and Qing He. 2022. Multi-view multi-behavior contrastive learning in recommendation. In Database Systems for Advanced Applications: 27th International Conference. Springer, 166–182
2022
-
[110]
Lianghao Xia, Chao Huang, Yong Xu, Peng Dai, Bo Zhang, and Liefeng Bo
-
[111]
Zige Wang, Wanjun Zhong, Yufei Wang, Qi Zhu, Fei Mi, Baojun Wang, Lifeng Shang, Xin Jiang, and Qun Liu. 2024. Data Management For Training Large Language Models: A Survey. arXiv:cs.CL/2312.01700
2024 arXiv
-
[112]
Kilian Weinberger, Anirban Dasgupta, John Langford, Alex Smola, and Josh Attenberg. 2009. Feature Hashing for Large Scale Multitask Learning. InProceed- ings of the 26th annual international conference on machine learning . 1113–1120
2009
-
[113]
Le Wu, Yonghui Yang, Kun Zhang, Richang Hong, Yanjie Fu, and Meng Wang
-
[114]
Fadi Yamout and Rachad Lakkis. 2018. Improved TFIDF weighting techniques in document Retrieval. In 2018 Thirteenth International Conference on Digital Information Management. 69–73
2018
-
[115]
Bencheng Yan, Pengjie Wang, Jinquan Liu, Wei Chao Lin, Kuang Chih Lee, Jian Xu, and Bo Zheng. 2021. Binary Code based Hash Embedding for Web- scale Applications. Proceedings of the 30th ACM International Conference on Information & Knowledge Management (2021)
2021
-
[116]
Xuanhua Yang, Xiaoyu Peng, Penghui Wei, Shaoguo Liu, Liang Wang, and Bo Zheng. 2022. Adasparse: Learning adaptively sparse structures for multi-domain click-through rate prediction. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management ....
2022
-
[117]
Wenwen Ye, Shuaiqiang Wang, Xu Chen, Xuepeng Wang, Zheng Qin, and Dawei Yin. 2020. Time Matters: Sequential Recommendation with Complex Temporal Information. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval . 1459–1468
2020
-
[118]
Mingjia Yin, Hao Wang, Wei Guo, Yong Liu, Zhi Li, Sirui Zhao, Zhen Wang, Defu Lian, and Enhong Chen. 2024. Learning Partially Aligned Item Representation for Cross-Domain Sequential Recommendation. arXiv preprint arXiv:2405.12473 (2024)
2024 arXiv
-
[119]
In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval
Multiplex behavioral relation learning for recommendation via memory augmented transformer network. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval . 2397– 2406
-
[120]
Wenjia Xie, Hao Wang, Luankang Zhang, Rui Zhou, Defu Lian, and Enhong Chen. 2024. Breaking Determinism: Fuzzy Modeling of Sequential Recommenda- tion Using Discrete State Space Diffusion Model.arXiv preprint arXiv:2410.23994 (2024)
2024 arXiv
-
[121]
Wenjia Xie, Rui Zhou, Hao Wang, Tingjia Shen, and Enhong Chen. 2024. Bridging User Dynamics: Transforming Sequential Recommendations with Schrödinger Bridge and Diffusion Models. In Proceedings of the 33rd ACM Inter- national Conference on Information and Knowledge Management ...
2024
-
[122]
Xiang Xu, Hao Wang, Wei Guo, Luankang Zhang, Wanshan Yang, Runlong Yu, Yong Liu, Defu Lian, and Enhong Chen. 2024. Multi-granularity Interest Retrieval and Refinement Network for Long-Term User Behavior Modeling in CTR Prediction. arXiv preprint arXiv:2411.15005 (2024)
2024 arXiv
-
[123]
Wenhui Yu, Chao Feng, Yanze Zhang, Lantao Hu, Peng Jiang, and Han Li. 2024. IFA: Interaction Fidelity Attention for Entire Lifelong Behaviour Sequence Modeling. arXiv preprint arXiv:2406.09742 (2024)
2024 arXiv
-
[124]
Enming Yuan, Wei Guo, Zhicheng He, Huifeng Guo, Chengkai Liu, and Ruiming Tang. 2022. Multi-behavior sequential transformer recommender. InProceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval. 1642–1652
2022
-
[125]
Huanhuan Yuan, Yongli Wang, Xia Feng, and Shurong Sun. 2018. Sentiment analysis based on weighted word2vec and att-lstm. InProceedings of the 2018 2nd international conference on computer science and artificial intelligence . 420–424
2018
-
[126]
Jiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang, Rui Li, Xuan Cao, Leon Gao, Zhaojie Gong, Fangda Gu, Michael He, et al. [n. d.]. Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommen- dations. In Forty-first International Conference ...
-
[127]
Yujia Zhai, Chengquan Jiang, Leyuan Wang, Xiaoying Jia, Shang Zhang, Zizhong Chen, Xin Liu, and Yibo Zhu. 2023. ByteTransformer: A high-performance trans- former boosted for variable-length inputs. In 2023 IEEE International Parallel and Distributed Processing Symposium . 344–355
2023
-
[128]
Mingjia Yin, Hao Wang, Wei Guo, Yong Liu, Suojuan Zhang, Sirui Zhao, Defu Lian, and Enhong Chen. 2024. Dataset Regeneration for Sequential Recom- mendation. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 3954–3965
2024
-
[129]
Mingjia Yin, Hao Wang, Xiang Xu, Likang Wu, Sirui Zhao, Wei Guo, Yong Liu, Ruiming Tang, Defu Lian, and Enhong Chen. 2023. APGL4SR: A Generic Framework with Adaptive and Personalized Global Collaborative Information in Sequential Recommendation. In Proceedings of the 32nd ACM ...
2023
-
[130]
Mingjia Yin, Chuhan Wu, Yufei Wang, Hao Wang, Wei Guo, Yasheng Wang, Yong Liu, Ruiming Tang, Defu Lian, and Enhong Chen. 2024. Entropy Law: The Story behind Data Compression and LLM Performance. arXiv preprint arXiv:2407.06645 (2024)
2024 arXiv
-
[131]
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung- Gon Chun. 2022. Orca: A distributed serving system for{Transformer-Based} generative models. In 16th USENIX Symposium on Operating Systems Design and Implementation. 521–538
2022
-
[132]
Weijie Zhao, Deping Xie, Ronglai Jia, Yulei Qian, Ruiquan Ding, Mingming Sun, and Ping Li. 2020. Distributed hierarchical gpu parameter server for massive scale deep learning ads systems. Proceedings of Machine Learning and Systems 2 (2020), 412–428
2020
-
[133]
Yilong Zhao, Chien-Yu Lin, Kan Zhu, Zihao Ye, Lequn Chen, Size Zheng, Luis Ceze, Arvind Krishnamurthy, Tianqi Chen, and Baris Kasikci. 2024. Atom: Low- bit quantization for efficient and accurate llm serving. Proceedings of Machine Learning and Systems 6 (2024), 196–209
2024
-
[134]
Lei Zheng, Vahid Noroozi, and Philip S Yu. 2017. Joint deep modeling of users and items using reviews for recommendation. In Proceedings of the 10th ACM International Conference on Web Search and Data Mining . 425–434
2017
-
[135]
Guorui Zhou, Weijie Bian, Kailun Wu, Lejian Ren, Qi Pi, Yujing Zhang, Can Xiao, Xiang-Rong Sheng, Na Mou, Xinchen Luo, et al. 2020. CAN: revisiting feature co-action for click-through rate prediction. arXiv preprint arXiv:2011.05625 (2020)
2020 arXiv
-
[136]
Kun Zhou, Hui Wang, Wayne Xin Zhao, Yutao Zhu, Sirui Wang, Fuzheng Zhang, Zhongyuan Wang, and Ji-Rong Wen. 2020. S3-Rec: Self-Supervised Learning for Sequential Recommendation with Mutual Information Maximization. In Pro- ceedings of the 29th ACM International Conference on In...
2020
-
[137]
Buyun Zhang, Liang Luo, Yuxin Chen, Jade Nie, Xi Liu, Daifeng Guo, Yanli Zhao, Shen Li, Yuchen Hao, Yantao Yao, et al. 2024. Wukong: Towards a Scaling Law for Large-Scale Recommendation. arXiv preprint arXiv:2403.02545 (2024)
2024 arXiv
-
[138]
Chi Zhang, Yantong Du, Xiangyu Zhao, Qilong Han, Rui Chen, and Li Li
-
[140]
Luankang Zhang, Hao Wang, Suojuan Zhang, Mingjia Yin, Yongqiang Han, Jiaqing Zhang, Defu Lian, and Enhong Chen. 2024. A Unified Framework for Adaptive Representation Enhancement and Inversed Learning in Cross-Domain Recommendation. arXiv preprint arXiv:2404.00268 (2024)
2024 arXiv
-
[141]
Weinan Zhang, Jiarui Qin, Wei Guo, Ruiming Tang, and Xiuqiang He. 2021. Deep learning for click-through rate estimation. arXiv preprint arXiv:2104.10584 (2021)
2021 arXiv
-
[2019]
In Proceedings of the 28th ACM International Conference on Information and Knowledge Management
BERT4Rec: Sequential Recommendation with Bidirectional Encoder Rep- resentations from Transformer. In Proceedings of the 28th ACM International Conference on Information and Knowledge Management . Association for Com- puting Machinery, 1441–1450
-
[2020]
In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval
Joint item recommendation and attribute inference: An adaptive graph convolutional network approach. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval . 679–688
-
[2021]
arXiv preprint arXiv:2108.04468 (2021)
End-to-end user behavior retrieval in click-through rateprediction model. arXiv preprint arXiv:2108.04468 (2021)
2021 arXiv
-
[2022]
In Proceedings of the 31st ACM International Conference on Information & Knowledge Management
Hierarchical item inconsistency signal learning for sequence denoising in sequential recommendation. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management . 2508–2518
-
[2023]
In Proceedings of the ACM Web Conference 2023
Autodenoise: Automatic data instance denoising for recommendations. In Proceedings of the ACM Web Conference 2023 . 1003–1011
2023
-
[2024]
Advances in Neural Information Processing Systems 36 (2024)
Recommender systems with generative retrieval. Advances in Neural Information Processing Systems 36 (2024)
2024
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.