Pith. sign in

REVIEW 4 major objections 6 minor 1 cited by

Pre-train, Align, and Disentangle: Empowering Sequential Recommendation with Large Language Models

T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read PAD aligns frozen LLM text embeddings to collaborative space with a characteristic multi-kernel MMD loss, then fuses three experts by item frequency; the result is state-of-the-art nDCG@10 on three datasets, with the largest gains on cold…

desk verdict Solid three-phase recipe with a real cold-start win, but the claimed alignment mechanism is untested because no ablation removes the MMD term. read the letter →

arxiv 2412.04107 v2 pith:56DKATJ7 submitted 2024-12-05 cs.IR cs.AI

classification cs.IRcs.AI
keywords SequentialRecommendationLargeLanguageModelMaximumMeanDiscrepancyCold-startCharacteristicKernelMixtureofExpertsEmbeddingAlignmentReproducingHilbertSpace
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Sequential recommenders learn users' tastes from click histories, but they struggle when an item has few interactions because they rely on collaborative IDs alone. This paper argues that the missing signal is the semantic knowledge already encoded in a frozen large language model, and it proposes a three-phase recipe—Pre-train, Align, Disentangle (PAD)—for injecting that knowledge without paying LLM inference costs at serving time. The core move is a recommendation-anchored alignment loss using multi-kernel maximum mean discrepancy (MK-MMD) with Gaussian kernels, which is claimed to capture all distribution differences between textual and collaborative embeddings while a recommendation-label anchor keeps the collaborative embeddings intact. On MIND, Amazon Electronics, and Prime Pantry, PAD beats the best baseline by 1.5% to 9.5% in nDCG@10, with the largest relative gains on cold items.

What carries the argument

The load-bearing object is the characteristic recommendation-anchored alignment loss $L = L_{\mathrm{REC}} + \gamma \cdot D_k^2$, where $D_k$ is multi-kernel maximum mean discrepancy built from five Gaussian kernels and $L_{\mathrm{REC}}$ is the binary cross-entropy recommendation loss. In theory, a characteristic kernel makes the kernel mean embedding $P \mapsto \mu_P$ injective, so MMD between the two embeddings is sensitive to all distribution differences; the BCE anchor prevents the collaborative embeddings from drifting away from their original predictive structure. A second mechanism is the triple-expert decoder: a recommendation-specific expert, an alignment expert, and an LLM-specific expert, each fed through its own embedding table and fused by a frequency-aware gating network that assigns more weight to text-derived experts for low-frequency target items. This architecture is what lets the model keep the pre-trained collaborative space intact while still exploiting text for cold items.

What would settle it

Replace the MMD term in Eq. (4) with a fixed random Gaussian feature projection of the same output cost, keeping the BCE anchor and the triple experts; if the cold-item nDCG@10 gains over SMEM do not degrade, the characteristic-kernel alignment is not the source of the improvement.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that LLM knowledge helps sequential recommendation only when alignment is anchored to the recommendation signal and when the model keeps dedicated experts per modality. In Phase 1, a standard sequential model is pre-trained on item IDs and item text is encoded once by a frozen LLM. In Phase 2, the frozen text embeddings are projected to the collaborative space under $L = L_{\mathrm{REC}} + \gamma \cdot D_k^2(\{h^s_i\}_a,\{h^c_i\}_a)$, where $D_k^2$ is multi-kernel MMD over five Gaussian kernels and $L_{\mathrm{REC}}$ is binary cross-entropy on the recommendation label; the MMD term is meant to match all distribution statistics, and the BCE term is meant to prevent catastrophic forgetting of collaborative embeddings. In Phase 3, three experts—recommendation-specific, alignment, and LLM-specific—are fused by a frequency-aware gate that leans on text-derived signals for rare items. The paper's headline result is that this yields the best nDCG@10 on all three datasets, and that the textual distances between pairs of items now re-order to follow collaborative distances, which is what drives the cold-item gains.

Load-bearing premise

The load-bearing premise is that a finite MMD computed from five Gaussian kernels on 128-dimensional mini-batch embeddings captures the distribution differences between text and collaborative spaces that actually matter for recommendation.

Editorial extensions

If this is right

  • Text embeddings are computed once by the frozen LLM, so serving-time inference adds only a small MLP plus three experts; PAD fits inside the latency budget of ID-based recommenders.
  • PAD transfers to other sequence backbones: on GRU4Rec and Caser it raises both HR@10 and nDCG@10 across all three datasets, making it a model-agnostic enhancement.
  • Characteristic kernels (Gaussian and Laplacian) beat linear, cosine, and InfoNCE losses in the alignment phase, so kernel choice is part of the method, not an implementation detail.
  • The Kendall-tau discrepancy metric gives future work a direct way to measure catastrophic forgetting in aligned embeddings by comparing distance orderings before and after alignment.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • One extension the paper leaves untested: hold the BCE anchor and triple experts fixed and replace the five-Gaussian MK-MMD with a fixed random feature map of the same cost. If cold-start gains survive, the characteristic-kernel property is not the active ingredient, and the theoretical story would need to change.
  • The frequency-aware gating suggests a general design rule for multi-modal recommenders: trust auxiliary modalities more in data-sparse regions and always keep a dedicated expert for the original modality, which could be tested with image or review embeddings.
  • The Kendall-tau tool could serve as a general diagnostic for embedding-space drift in any multi-modal alignment pipeline, flagging which item-frequency buckets suffer the most reordering.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes PAD (Pre-train, Align, Disentangle), a three-phase framework for sequential recommendation that augments ID-based collaborative embeddings with frozen LLM-generated text embeddings. Phase 1 pre-trains a sequential recommender (SASRec) and extracts text embeddings via LLM2Vec. Phase 2 introduces a ``characteristic recommendation-anchored alignment'' loss: a multi-kernel MMD (MK-MMD) with Gaussian kernels, combined with a BCE recommendation loss, to align text embeddings toward the collaborative space. Phase 3 fine-tunes a triple-expert architecture (alignment expert, LLM-specific expert, ID-specific expert) with frequency-aware gating based on the target item's frequency bucket. Experiments on MIND, Amazon Electronics, and Prime Pantry report consistent HR@10 and nDCG@10 gains over several baselines, with larger relative gains on cold items, and compatibility with GRU4Rec and Caser backbones. The code and datasets are released.

Significance. If the mechanistic claims hold, PAD provides a practical, low-latency recipe for incorporating LLM knowledge into sequential recommender systems, addressing the cold-start problem while avoiding the inference cost of LLM-as-recommender approaches. The empirical results on three public datasets, compatibility with multiple backbones, and the released code are tangible strengths. However, the central contribution—the characteristic MK-MMD alignment—is not isolated in the ablation study; every variant that keeps the alignment module also keeps the BCE anchor, so the paper's causal attribution of the gains to the characteristic kernel is not yet supported. The theoretical motivation based on characteristic kernels is also presented as a population-level property, while the implementation uses a finite set of five Gaussian kernels on 128-dimensional embeddings with mini-batch training, a gap that is not discussed.

major comments (4)
  1. [Sec. 4.5 / Fig. 5(a)] The ablation that claims characteristic kernels outperform non-characteristic ones does not include a condition that removes the MMD term entirely (i.e., setting γ=0 and keeping only the BCE anchor). All variants in Fig. 5(a) include the same BCE component from Eq. (6), so the observed ordering could be driven by the kernel regularizer's interaction with the BCE loss, or even by the BCE anchor alone, rather than by the characteristic property of the kernel. Please report PAD (or the Phase-2 alignment model) with γ=0, and ideally with a simple per-item cosine or MSE alignment term in place of MMD, to determine whether the characteristic-kernel claim is load-bearing.
  2. [Sec. 4.4 / Fig. 4] The comparison of anchored vs. non-anchored alignment also conflates the presence of the BCE anchor with the presence of the MMD term. The Non-Anchored condition uses only the MMD loss, while the Rec-Anchored condition uses BCE plus MMD; there is no BCE-only condition. Consequently, the conclusion that ``recommendation anchoring avoids catastrophic forgetting'' cannot distinguish the effect of the recommendation label from the effect of keeping the MMD. A BCE-only (γ=0) condition is needed to isolate the MMD's contribution to the anchored result.
  3. [Tab. 2 and Sec. 4.2] The paper reports that results are averaged over 3 runs (Sec. 4.1.4) but does not provide error bars or standard deviations in any table or figure, and the t-test is reported only against the best baseline. This makes it difficult to assess the stability of the claimed improvements, particularly on the smaller Prime Pantry dataset where the reported gains are large but the underlying HR@10 values are around 3.8. Please report mean±std (or per-run values) for the main results, and preferably also for the ablation comparisons.
  4. [Sec. 2.2 and Sec. 4.1.4] The theoretical justification for using MMD with characteristic kernels is that the kernel mean embedding is injective in an infinite-dimensional RKHS, preserving ``all information about the distribution.'' The implementation, however, uses a finite multi-kernel MMD with five fixed Gaussian bandwidths on 128-dimensional embeddings and mini-batches. The paper does not discuss how closely this finite approximation preserves the characteristic property, nor how the bandwidth set (σ={-3,-2,-1,0,1}) was chosen or what scale of distances it covers. If the characteristic property is not preserved under this approximation, the claimed theoretical advantage of the Gaussian/Laplacian kernels in Fig. 5(a) would not follow directly. Please either provide a finite-sample justification or temper the theoretical claims to match the implemented estimator.
minor comments (6)
  1. [Eq. (3)] The constraint in Eq. (3) writes sum of β_u = d, but d is not defined; in the MK-MMD literature the coefficients usually sum to 1. Please clarify the intended normalization.
  2. [Eq. (12)] The Gaussian kernel in Eq. (12) uses a bandwidth σ, but Sec. 4.1.4 states σ = {-3,-2,-1,0,1}, which includes 0 and would make the denominator zero. It appears the listed values are log-scale bandwidths (e.g., 2^σ); please state this explicitly for reproducibility.
  3. [Sec. 3.3 / Eq. (11)] The frequency-aware gating network is described as taking the frequency bucket ID and the expert embedding, but the paper does not specify how the gating probabilities are normalized (e.g., softmax), nor the exact input concatenation. Please detail the gating architecture.
  4. [Fig. 3 / Fig. 7 captions] The captions refer to PID Top-10% and PID Bottom-10% without defining these symbols; the definitions appear only in the main text of Sec. 4.2. The captions should be self-contained.
  5. [References] References [22] and [23] cite the same SASRec paper (Kang and McAuley 2018); one duplicate should be removed.
  6. [Sec. 4.6.1] The description of the 'with align' and 'w/o align' lines in Fig. 6(b) refers to color (violet) that is not explained in the caption; please add a legend or describe the colors in the text.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation; the characteristic-MMD claim is empirical and the missing-MMD ablation is an attribution concern, not a self-referential reduction.

full rationale

PAD is an empirical training-and-evaluation paper. Its headline results (Tab. 2, Fig. 4, Fig. 5) are measured on external held-out test splits against external and author-proposed baselines; they are not derived from the fitted alignment parameters by construction. The theoretical invocation of characteristic kernels is supported by external RKHS literature (Gretton, Fukumizu, Muandet, Sriperumbudur), not by a self-citation chain, and the paper does not import a uniqueness theorem from its own prior work. The self-citations that do appear (multi-embedding paradigms, STEM-like structures) are contextual inspiration rather than load-bearing justifications for the reported gains. The skeptical reading that the BCE anchor in Eq. (6) may do the work attributed to the MK-MMD term is a missing-ablation and causal-attribution concern: Eq. (4) is a sum of two losses, not a definitional identity, and removing the MMD term would be an empirical ablation rather than a logical contradiction. Likewise, the observation that the author-proposed SMEM baseline is used as the strongest comparator is an evaluation-design point, not a circular reduction: PAD's numbers are still measured, not forced. No step of the paper's derivation reduces, by its own equations or by self-citation, to its inputs, so significant circularity is not present.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The paper's central claims rest on established RKHS theory, the assumption that frozen LLM text embeddings carry recommendation-relevant semantics, and the ad hoc assumption that item frequency indicates which expert to trust. The only hand-set constants that materially affect the method are the alignment weight gamma, the kernel bandwidth list, and the frequency bucket count. No new physical or mathematical entities are introduced.

free parameters (3)
  • gamma (alignment loss weight) = 0.2
    Hyperparameter in L = L_REC + gamma * L_MK-MMD (Eq. 4); searched and set to 0.2 (Appendix A). The central claim depends on this trade-off.
  • Gaussian kernel bandwidth set = sigma = {-3, -2, -1, 0, 1}, m = 5
    Chosen in Sec. 4.1.4; the MK-MMD behavior depends on these scales, and no sensitivity analysis is reported.
  • Frequency bucket count B = not reported; figures use 10 buckets
    Frequency-aware gating divides items into buckets; B determines the granularity of the gating signal and is not specified in the paper.
assumptions (3)
  • standard math Characteristic kernels in RKHS preserve all information of the distribution P (injective mean embedding), so MK-MMD with Gaussian kernels captures all statistical aspects (Sec. 2.2, Sec. 3.2).
    Theoretical result from Fukumizu et al. [12,13]; accepted but population-level, with finite-sample caveats.
  • domain assumption LLM2Vec (Llama3-8B, frozen) produces text embeddings that are semantically aligned with collaborative preference space after a learned MLP projection (Sec. 3.2, Eq. 7).
    This is the premise of the whole alignment phase; not proven beyond downstream task performance.
  • ad hoc to paper The frequency of an item in the training set is a good proxy for the credibility of collaborative vs. textual evidence (Sec. 3.3, Eq. 11).
    Introduced for the gating mechanism; no independent evidence or theory, only an ablation on MIND.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Pre-train, Align, and Disentangle: Empowering Sequential Recommendation with Large Language Models." pith.science (2026). https://pith.science/paper/56DKATJ7

@misc{pith2026241204107,
  author       = {Pith},
  title        = {Pith review of: Pre-train, Align, and Disentangle: Empowering Sequential Recommendation with Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/56DKATJ7}},
  note         = {Machine review of arXiv:2412.04107}
}
read the original abstract

Sequential Recommendation (SR) aims to leverage the sequential patterns in users' historical interactions to accurately track their preferences. However, the primary reliance of existing SR methods on collaborative data results in challenges such as the cold-start problem and sub-optimal performance. Concurrently, despite the proven effectiveness of large language models (LLMs), their integration into commercial recommender systems is impeded by issues such as high inference latency, incomplete capture of all distribution statistics, and catastrophic forgetting. To address these issues, we introduce a novel Pre-train, Align, and Disentangle (PAD) framework to enhance SR models with LLMs. In particular, we initially pre-train both the SR and LLM models to obtain collaborative and textual embeddings. Subsequently, we propose a characteristic recommendation-anchored alignment loss using multi-kernel maximum mean discrepancy with Gaussian kernels. Lastly, a triple-experts architecture, comprising aligned and modality-specific experts with disentangled embeddings, is fine-tuned in a frequency-aware manner. Experimental results on three public datasets validate the efficacy of PAD, indicating substantial enhancements and compatibility with various SR backbone models, particularly for cold items. The code and datasets are accessible for reproduction at https://github.com/Applied-Machine-Learning-Lab/PAD.

Figures

Figures reproduced from arXiv: 2412.04107 by the authors.

Figure 1
Figure 1. Overall framework of PAD. The number in parentheses (128 and 4096) denotes the embedding dimension. The [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Comparison of the original SASRec and PAD on [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. P ID Top-10% and P ID Bottom-10% under the distance distribution regarding the collaborative and textual embeddings in original SASRec, SMEM, CTRL, and our proposed PAD on cold items. (a) Anchored Loss: MIND 18.4 18.5 18.6 18.7 (b) Anchored Loss: Electronics 2.3 2.4 2.5 (c) Anchored Loss: Prime Pantry 3.1 3.3 3.5 3.7 3.9 Rec-Anchored Rec-Anchored & Frozen Non-Anchored SMEM [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (2 more)
Figure 6
Figure 6. Figure 6: Kendall’s tau between the Euclidean distance (same as L2 distance) of (a) collaborative embedding of different anchored [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: P ID Top-10% and P ID Bottom-10% under the distance distribution regarding the collaborative and text embeddings in the original SASRec, SMEM, CTRL, and PAD on warm items. ACKNOWLEDGEMENT This research was partially supported by Tencent Rhino-Bird Fo￾cused Research Pro…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Large Language Model Enhanced Recommender Systems: A Survey

    cs.IR 2024-12 unverdicted novelty 4.0 of 10

    A survey organizing LLM-enhanced recommender systems into knowledge, interaction, and model enhancement, and tracing a shift from explicit text to implicit embeddings and fine-tuned open-source LLMs.

Reference graph

Works this paper leans on

81 extracted references · 20 canonical work pages · cited by 1 Pith paper

  1. [1]

    Keqin Bao, Jizhi Zhang, Yang Zhang, Wenjie Wang, Fuli Feng, and Xiangnan He. 2023. Tallrec: An effective and efficient tuning framework to align large language model with recommendation. InProceedings of the 17th ACM Conference on Recommender Systems. 1007–1014

  2. [2]

    Parishad BehnamGhader, Vaibhav Adlakha, Marius Mosbach, Dzmitry Bahdanau, Nicolas Chapados, and Siva Reddy. 2024. Llm2vec: Large language models are secretly powerful text encoders. arXiv preprint arXiv:2404.05961 (2024)

  3. [3]

    Shuqing Bian, Xingyu Pan, Wayne Xin Zhao, Jinpeng Wang, Chuyuan Wang, and Ji-Rong Wen. 2023. Multi-modal mixture of experts represetation learning for sequential recommendation. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management . 110–119

  4. [4]

    Jianxin Chang, Chenbin Zhang, Zhiyi Fu, Xiaoxue Zang, Lin Guan, Jing Lu, Yiqun Hui, Dewei Leng, Yanan Niu, Yang Song, et al. 2023. TWIN: TWo-stage interest network for lifelong user behavior modeling in CTR prediction at kuaishou. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 3785–3794

  5. [5]

    Gaode Chen, Ruina Sun, Yuezihan Jiang, Jiangxia Cao, Qi Zhang, Jingjian Lin, Han Li, Kun Gai, and Xinghua Zhang. 2024. A Multi-modal Modeling Framework for Cold-start Short-video Recommendation. In Proceedings of the 18th ACM Conference on Recommender Systems . 391–400

  6. [6]

    Jingyuan Chen, Hanwang Zhang, Xiangnan He, Liqiang Nie, Wei Liu, and Tat- Seng Chua. 2017. Attentive collaborative filtering: Multimedia recommendation with item-and component-level attention. In Proceedings of the 40th International ACM SIGIR conference on Research and Development in Information Retrieval . 335–344

  7. [7]

    Xu Chen, Hanxiong Chen, Hongteng Xu, Yongfeng Zhang, Yixin Cao, Zheng Qin, and Hongyuan Zha. 2019. Personalized fashion recommendation with visual explanations based on multimodal attention network: Towards visually explainable recommendation. In Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieva...

  8. [8]

    Kounianhua Du, Jizheng Chen, Jianghao Lin, Yunjia Xi, Hangyu Wang, Xinyi Dai, Bo Chen, Ruiming Tang, and Weinan Zhang. 2024. DisCo: Towards Har- monious Disentanglement and Collaboration between Tabular and Semantic Space for Recommendation. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 666–676

Show all 81 references
  1. [9]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024)

  2. [10]

    Ningya Feng, Junwei Pan, Jialong Wu, Baixu Chen, Ximei Wang, Qian Li, Xian Hu, Jie Jiang, and Mingsheng Long. 2024. Long-Sequence Recommendation Models Need Decoupled Embeddings. arXiv preprint arXiv:2410.02604 (2024)

  3. [11]

    Yufei Feng, Fuyu Lv, Weichen Shen, Menghan Wang, Fei Sun, Yu Zhu, and Keping Yang. 2019. Deep session interest network for click-through rate prediction. In IJCAI

  4. [12]

    Kenji Fukumizu, Francis R Bach, and Michael I Jordan. 2004. Dimensionality reduction for supervised learning with reproducing kernel Hilbert spaces.Journal of Machine Learning Research 5, Jan (2004), 73–99

  5. [13]

    Kenji Fukumizu, Arthur Gretton, Bernhard Schölkopf, and Bharath K Sriperum- budur. 2008. Characteristic kernels on groups and semigroups. Advances in neural information processing systems 21 (2008)

  6. [14]

    Jingtong Gao, Xiangyu Zhao, Muyang Li, Minghao Zhao, Runze Wu, Ruocheng Guo, Yiding Liu, and Dawei Yin. 2024. SMLP4Rec: an Efficient all-MLP architec- ture for sequential recommendations. ACM Transactions on Information Systems 42, 3 (2024), 1–23

  7. [15]

    Arthur Gretton, Karsten M Borgwardt, Malte J Rasch, Bernhard Schölkopf, and Alexander Smola. 2012. A kernel two-sample test. The Journal of Machine Learning Research 13, 1 (2012), 723–773

  8. [16]

    Arthur Gretton, Dino Sejdinovic, Heiko Strathmann, Sivaraman Balakrishnan, Massimiliano Pontil, Kenji Fukumizu, and Bharath K Sriperumbudur. 2012. Opti- mal kernel choice for large-scale two-sample tests.Advances in neural information processing systems 25 (2012)

  9. [17]

    Xingzhuo Guo, Junwei Pan, Ximei Wang, Baixu Chen, Jie Jiang, and Mingsheng Long. 2024. On the Embedding Collapse when Scaling up Recommendation Models. ICML (2024)

  10. [18]

    Ruining He and Julian McAuley. 2016. VBPR: visual bayesian personalized ranking from implicit feedback. In Proceedings of the AAAI conference on artificial intelligence, Vol. 30

  11. [19]

    Xinran He, Junfeng Pan, Ou Jin, Tianbing Xu, Bo Liu, Tao Xu, Yanxin Shi, Antoine Atallah, Ralf Herbrich, Stuart Bowers, et al. 2014. Practical lessons from predicting clicks on ads at facebook. In International Workshop on Data Mining for Online Advertising (ADKDD). 1–9

  12. [21]

    Balázs Hidasi, Alexandros Karatzoglou, Linas Baltrunas, and Domonkos Tikk

  13. [22]

    Wang-Cheng Kang and Julian McAuley. 2018. Self-attentive sequential recom- mendation. In 2018 IEEE International Conference on Data Mining (ICDM) . IEEE, 197–206

  14. [23]

    Wang-Cheng Kang and Julian McAuley. 2018. Self-attentive sequential recom- mendation. In 2018 IEEE international conference on data mining (ICDM) . IEEE, 197–206

  15. [24]

    Maurice G Kendall. 1938. A new measure of rank correlation. Biometrika 30, 1-2 (1938), 81–93

  16. [25]

    Chengxi Li, Yejing Wang, Qidong Liu, Xiangyu Zhao, Wanyu Wang, Yiqi Wang, Lixin Zou, Wenqi Fan, and Qing Li. 2023. STRec: Sparse transformer for sequential recommendations. In Proceedings of the 17th ACM conference on recommender systems. 101–111

  17. [26]

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning . PMLR, 19730–19742

  18. [27]

    Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International conference on machine learning . PMLR, 12888–12900

  19. [28]

    Xiangyang Li, Bo Chen, Lu Hou, and Ruiming Tang. 2023. CTRL: Connect Collab- orative and Language Model for CTR Prediction. arXiv preprint arXiv:2306.02841 (2023)

  20. [29]

    Jiahao Liang, Xiangyu Zhao, Muyang Li, Zijian Zhang, Wanyu Wang, Haochen Liu, and Zitao Liu. 2023. Mmmlp: Multi-modal multilayer perceptron for sequen- tial recommendations. InProceedings of the ACM Web Conference 2023. 1109–1117

  21. [30]

    Jiayi Liao, Sihang Li, Zhengyi Yang, Jiancan Wu, Yancheng Yuan, Xiang Wang, and Xiangnan He. 2023. Llara: Aligning large language models with sequential recommenders. arXiv preprint arXiv:2312.02445 (2023)

  22. [31]

    Zhutian Lin, Junwei Pan, Haibin Yu, Xi Xiao, Ximei Wang, Zhixiang Feng, Shifeng Wen, Shudong Huang, Lei Xiao, and Jie Jiang. 2024. Crocodile: Cross Experts Covariance for Disentangled Learning in Multi-Domain Recommendation. arXiv preprint arXiv:2405.12706 (2024)

  23. [32]

    Fan Liu, Zhiyong Cheng, Changchang Sun, Yinglong Wang, Liqiang Nie, and Mohan Kankanhalli. 2019. User diverse preference modeling by multimodal attentive metric learning. In Proceedings of the 27th ACM international conference on multimedia. 1526–1534

  24. [33]

    Qidong Liu, Xian Wu, Yejing Wang, Zijian Zhang, Feng Tian, Yefeng Zheng, and Xiangyu Zhao. 2024. Llm-esr: Large language models enhancement for long- tailed sequential recommendation. Advances in Neural Information Processing Systems 37 (2024), 26701–26727

  25. [34]

    Yifan Liu, Kangning Zhang, Xiangyuan Ren, Yanhua Huang, Jiarui Jin, Yingjie Qin, Ruilong Su, Ruiwen Xu, and Weinan Zhang. 2024. An Aligning and Training Framework for Multimodal Recommendations. arXiv preprint arXiv:2403.12384 (2024)

  26. [35]

    Ziru Liu, Shuchang Liu, Zijian Zhang, Qingpeng Cai, Xiangyu Zhao, Kesen Zhao, Lantao Hu, Peng Jiang, and Kun Gai. 2024. Sequential recommendation for optimizing both immediate feedback and long-term retention. In Proceedings of the 47th International ACM SIGIR Conference on Re...

  27. [36]

    Ziru Liu, Jiejie Tian, Qingpeng Cai, Xiangyu Zhao, Jingtong Gao, Shuchang Liu, Dayou Chen, Tonghao He, Dong Zheng, Peng Jiang, et al. 2023. Multi-task recommendations with reinforcement learning. In Proceedings of the ACM web conference 2023. 1273–1282

  28. [37]

    Ilya Loshchilov, Frank Hutter, et al. 2017. Fixing weight decay regularization in adam. arXiv preprint arXiv:1711.05101 (2017)

  29. [38]

    H Brendan McMahan, Gary Holt, David Sculley, Michael Young, Dietmar Ebner, Julian Grady, Lan Nie, Todd Phillips, Eugene Davydov, Daniel Golovin, et al. 2013. Ad click prediction: a view from the trenches. In ACM SIGKDD International conference on Knowledge Discovery & Data Min...

  30. [39]

    Krikamol Muandet, Kenji Fukumizu, Bharath Sriperumbudur, Bernhard Schölkopf, et al. 2017. Kernel mean embedding of distributions: A review and beyond. Foundations and Trends® in Machine Learning 10, 1-2 (2017), 1–141

  31. [40]

    Jianmo Ni, Jiacheng Li, and Julian McAuley. 2019. Justifying recommendations using distantly-labeled reviews and fine-grained aspects. In Proc. of EMNLP

  32. [41]

    Junwei Pan, Wei Xue, Ximei Wang, Haibin Yu, Xun Liu, Shijie Quan, Xueming Qiu, Dapeng Liu, Lei Xiao, and Jie Jiang. 2024. Ads Recommendation in a Collapsed and Entangled World. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 5566–5577

  33. [42]

    Qi Pi, Guorui Zhou, Yujing Zhang, Zhe Wang, Lejian Ren, Ying Fan, Xiaoqiang Zhu, and Kun Gai. 2020. Search-based user interest modeling with lifelong sequential behavior data for click-through rate prediction. In ACM International Conference on Information & Knowledge Manageme...

  34. [43]

    Shashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan Hulikal Keshavan, Trung Vu, Lukasz Heldt, Lichan Hong, Yi Tay, Vinh Tran, Jonah Samost, et al

  35. [44]

    Dino Sejdinovic, Bharath Sriperumbudur, Arthur Gretton, and Kenji Fukumizu

  36. [45]

    Xiang-Rong Sheng, Feifan Yang, Litong Gong, Biao Wang, Zhangming Chan, Yujing Zhang, Yueyao Cheng, Yong-Nan Zhu, Tiezheng Ge, Han Zhu, et al

  37. [46]

    Zihua Si, Lin Guan, ZhongXiang Sun, Xiaoxue Zang, Jing Lu, Yiqun Hui, Xingchao Cao, Zeyu Yang, Yichen Zheng, Dewei Leng, et al . 2024. TWIN V2: Scaling Ultra-Long User Behavior Sequence Modeling for Enhanced CTR Prediction at Kuaishou. arXiv preprint arXiv:2407.16357 (2024)

  38. [47]

    Anima Singh, Trung Vu, Nikhil Mehta, Raghunandan Keshavan, Maheswaran Sathiamoorthy, Yilin Zheng, Lichan Hong, Lukasz Heldt, Li Wei, Devansh Tandon, et al. 2024. Better Generalization with Semantic IDs: A Case Study in Ranking for Recommendations. In Proceedings of the 18th AC...

  39. [48]

    Liangcai Su, Junwei Pan, Ximei Wang, Xi Xiao, Shijie Quan, Xihua Chen, and Jie Jiang. 2024. STEM: Unleashing the Power of Embeddings for Multi-task Recommendation. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 9002–9010

  40. [49]

    Fei Sun, Jun Liu, Jian Wu, Changhua Pei, Xiao Lin, Wenwu Ou, and Peng Jiang

  41. [50]

    InProceedings of the 33rd ACM International Conference on Information and Knowledge Management

    Enhancing Taobao Display Advertising with Multimodal Representations: Challenges, Approaches and Insights. InProceedings of the 33rd ACM International Conference on Information and Knowledge Management . 4858–4865

  42. [51]

    A Vaswani. 2017. Attention is all you need. Advances in Neural Information Processing Systems (2017)

  43. [52]

    Hangyu Wang, Jianghao Lin, Xiangyang Li, Bo Chen, Chenxu Zhu, Ruiming Tang, Weinan Zhang, and Yong Yu. 2023. FLIP: Towards Fine-grained Alignment between ID-based Models and Pretrained Language Models for CTR Prediction. arXiv e-prints (2023), arXiv–2310

  44. [53]

    Hanbing Wang, Xiaorui Liu, Wenqi Fan, Xiangyu Zhao, Venkataramana Kini, Devendra Yadav, Fei Wang, Zhen Wen, Jiliang Tang, and Hui Liu. 2024. Rethinking large language model architectures for sequential recommendations. arXiv preprint arXiv:2402.09543 (2024)

  45. [54]

    Yuhao Wang, Ha Tsz Lam, Yi Wong, Ziru Liu, Xiangyu Zhao, Yichao Wang, Bo Chen, Huifeng Guo, and Ruiming Tang. 2023. Multi-task deep recommender systems: A survey. arXiv preprint arXiv:2302.03525 (2023)

  46. [55]

    Yuhao Wang, Ziru Liu, Yichao Wang, Xiangyu Zhao, Bo Chen, Huifeng Guo, and Ruiming Tang. 2024. Diff-MSR: A diffusion model enhanced paradigm for cold-start multi-scenario recommendation. In Proceedings of the 17th ACM International Conference on Web Search and Data Mining . 779–787

  47. [56]

    Jiaxi Tang and Ke Wang. 2018. Personalized top-n sequential recommenda- tion via convolutional sequence embedding. In Proceedings of the eleventh ACM international conference on web search and data mining . 565–573

  48. [57]

    Yuhao Wang, Xiangyu Zhao, Bo Chen, Qidong Liu, Huifeng Guo, Huanshuo Liu, Yichao Wang, Rui Zhang, and Ruiming Tang. 2023. PLATE: A prompt-enhanced paradigm for multi-scenario recommendations. In Proceedings of the 46th In- ternational ACM SIGIR Conference on Research and Devel...

  49. [58]

    Yinwei Wei, Xiang Wang, Liqiang Nie, Xiangnan He, Richang Hong, and Tat-Seng Chua. 2019. MMGCN: Multi-modal graph convolution network for personalized recommendation of micro-video. In Proceedings of the 27th ACM international conference on multimedia. 1437–1445

  50. [59]

    Fangzhao Wu, Ying Qiao, Jiun-Hung Chen, Chuhan Wu, Tao Qi, Jianxun Lian, Danyang Liu, Xing Xie, Jianfeng Gao, Winnie Wu, et al. 2020. Mind: A large-scale dataset for news recommendation. In Proceedings of the 58th annual meeting of the association for computational linguistics...

  51. [60]

    Derong Xu, Ziheng Zhang, Zhenxi Lin, Xian Wu, Zhihong Zhu, Tong Xu, Xiangyu Zhao, Yefeng Zheng, and Enhong Chen. 2024. Multi-perspective Improvement of Knowledge Graph Completion with Large Language Models. In LREC/COLING

  52. [61]

    Xihong Yang, Heming Jing, Zixing Zhang, Jindong Wang, Huakang Niu, Shuaiqiang Wang, Yu Lu, Junfeng Wang, Dawei Yin, Xinwang Liu, et al. 2024. DaRec: A Disentangled Alignment Framework for Large Language Model and Recommender System. arXiv preprint arXiv:2408.08231 (2024)

  53. [62]

    Yuhao Wang, Yichao Wang, Zichuan Fu, Xiangyang Li, Wanyu Wang, Yuyang Ye, Xiangyu Zhao, Huifeng Guo, and Ruiming Tang. 2024. Llm4msr: An llm- enhanced paradigm for multi-scenario recommendation. In Proceedings of the 33rd ACM International Conference on Information and Knowled...

  54. [63]

    Chi Zhang, Yantong Du, Xiangyu Zhao, Qilong Han, Rui Chen, and Li Li. 2022. Hierarchical item inconsistency signal learning for sequence denoising in se- quential recommendation. In Proceedings of the 31st ACM international conference on information & knowledge management . 2508–2518

  55. [64]

    Chi Zhang, Qilong Han, Rui Chen, Xiangyu Zhao, Peng Tang, and Hongtao Song

  56. [65]

    Taolin Zhang, Junwei Pan, Jinpeng Wang, Yaohua Zha, Bin Chen, Shengshui Luo, Yuan Wang, Ming Yue, Jie Jiang, and Shu-Tao Xia. 2024. Towards Scalable Semantic Representation for Recommendation. arXiv preprint arXiv:2410.09560 (2024)

  57. [66]

    Yang Zhang, Fuli Feng, Jizhi Zhang, Keqin Bao, Qifan Wang, and Xiangnan He

  58. [67]

    Kesen Zhao, Xiangyu Zhao, Zijian Zhang, and Muyang Li. 2022. Mae4rec: Storage- saving transformer for sequential recommendations. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management . 2681– 2690

  59. [68]

    Zheng Yuan, Fajie Yuan, Yu Song, Youhua Li, Junchen Fu, Fei Yang, Yunzhu Pan, and Yongxin Ni. 2023. Where to go next for recommender systems? id- vs. modality-based recommender models revisited. In Proceedings of the 46th International ACM SIGIR Conference on Research and Deve...

  60. [69]

    Xiangyu Zhao, Liang Zhang, Zhuoye Ding, Long Xia, Jiliang Tang, and Dawei Yin

  61. [70]

    Bowen Zheng, Yupeng Hou, Hongyu Lu, Yu Chen, Wayne Xin Zhao, Ming Chen, and Ji-Rong Wen. 2024. Adapting large language models by integrating collaborative semantics for recommendation. In 2024 IEEE 40th International Conference on Data Engineering (ICDE) . IEEE, 1435–1448

  62. [71]

    In 2024 IEEE 40th International Conference on Data Engineering (ICDE)

    Ssdrec: self-augmented sequence denoising for sequential recommendation. In 2024 IEEE 40th International Conference on Data Engineering (ICDE) . IEEE, 803–815

  63. [72]

    Guorui Zhou, Xiaoqiang Zhu, Chenru Song, Ying Fan, Han Zhu, Xiao Ma, Yanghui Yan, Junqi Jin, Han Li, and Kun Gai. 2018. Deep interest network for click-through rate prediction. In SIGKDD. 1059–1068

  64. [73]

    Haolin Zhou, Junwei Pan, Xinyi Zhou, Xihua Chen, Jie Jiang, Xiaofeng Gao, and Guihai Chen. 2024. Temporal Interest Network for User Response Prediction. In Companion Proceedings of the ACM on Web Conference 2024 . 413–422

  65. [76]

    Xiangyu Zhao, Long Xia, Liang Zhang, Zhuoye Ding, Dawei Yin, and Jiliang Tang. 2018. Deep reinforcement learning for page-wise recommendations. In Proceedings of the 12th ACM conference on recommender systems . 95–103

  66. [80]

    Guorui Zhou, Na Mou, Ying Fan, Qi Pi, Weijie Bian, Chang Zhou, Xiaoqiang Zhu, and Kun Gai. 2019. Deep interest evolution network for click-through rate prediction. In AAAI. 5941–5948

  67. [2013]

    The annals of statistics (2013), 2263–2291

    Equivalence of distance-based and RKHS-based statistics in hypothesis testing. The annals of statistics (2013), 2263–2291. Pre-train, Align, and Disentangle: Empowering Sequential Recommendation with Large Language Models SIGIR ’25, July 13–18, 2025, Padua, Italy

  68. [2015]

    arXiv preprint arXiv:1511.06939 (2015)

    Session-based recommendations with recurrent neural networks. arXiv preprint arXiv:1511.06939 (2015)

  69. [2016]

    SESSION-BASED RECOMMENDATIONS WITH RECURRENT NEURAL NETWORKS. In ICLR

  70. [2018]

    In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining

    Recommendations with negative feedback via pairwise deep reinforcement learning. In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining . 1040–1048

  71. [2019]

    In Proceedings of the 28th ACM international conference on information and knowledge management

    BERT4Rec: Sequential recommendation with bidirectional encoder rep- resentations from transformer. In Proceedings of the 28th ACM international conference on information and knowledge management . 1441–1450

  72. [2023]

    arXiv preprint arXiv:2310.19488 (2023)

    Collm: Integrating collaborative embeddings into large language models for recommendation. arXiv preprint arXiv:2310.19488 (2023)

  73. [2024]

    Advances in Neural Information Processing Systems 36 (2024)

    Recommender systems with generative retrieval. Advances in Neural Information Processing Systems 36 (2024)

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.