Pith. sign in

REVIEW 2 major objections 6 minor 59 references

Stream-aware side adapters can use every layer of a frozen multimodal embedder without the usual deep-adaptation collapse.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 08:21 UTC pith:R2UNXKMU

load-bearing objection Solid empirical side-adaptation paper: SHAF+ReSA stabilize full-depth frozen multimodal embedding adaptation and beat standard side adapters, with modest gains and a still-under-isolated mechanism story. the 2 major comments →

arxiv 2607.10909 v1 pith:R2UNXKMU submitted 2026-07-12 cs.IR

Stream-aware Side Adaptation for Large Pre-trained Multimodal Embedding Models in Sequential Recommendation

classification cs.IR
keywords Multimodal RecommendationSequential RecommendationLarge Embedding ModelsSide AdapterStream-aware FusionFrozen Backbone AdaptationParameter-efficient Tuning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Large multimodal embedding models give sequential recommenders rich item representations, but domain mismatch means they need adaptation, and full fine-tuning is too expensive. Side adapters keep the backbone frozen and train only a light branch, yet standard residual-plus-sigmoid designs get worse as more layers are adapted, so practitioners drop layers and discard useful hidden states. This paper argues the collapse comes from two design gaps: residual addition never models what to keep or suppress in already-fused states, and progressive sigmoid fusion overwrites earlier side memory. Stresa fixes both with stream-aware fusion that preserves historical memory and a residual stream adapter that applies selective updates. On public micro-video and product datasets, with two scales of a frozen vision-language embedder, Stresa beats standard side adapters and strong prior side-adaptation methods, and it stays stable when every backbone layer is used.

Core claim

Deep side adaptation of frozen large multimodal embedding models fails not because deeper hidden states are useless, but because conventional residual addition and sigmoid fusion lack selection and memory preservation; redesigning those two operators with stream-aware routing and residual stream mixing unlocks full-depth adaptation and improves sequential recommendation.

What carries the argument

Stresa: at each adapted layer, Stream-aware Hidden-Adapter Fusion (SHAF) reshapes hidden and side states into streams, applies identity-preserving Sinkhorn routing to transport earlier side memory, then fuses with the current backbone state; Residual Stream Adapter (ReSA) then applies stream-wise pre-, post-, and residual mappings around a bottleneck MLP to produce the next side state.

Load-bearing premise

The paper assumes that side adapters degrade with depth mainly because residual addition and sigmoid fusion fail to select and preserve information, rather than because deeper layers are noisy, capacity-limited, or poorly suited to the data.

What would settle it

On the same frozen backbones and datasets, adapt all layers with a standard residual/sigmoid side adapter that is given comparable capacity or a different fusion rule; if it matches or beats full-depth Stresa without stream-aware identity preservation, the claimed causal diagnosis of depth degradation fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. The paper proposes Stresa, a stream-aware side-adaptation framework for frozen large multimodal embedding backbones (Qwen3-VL-Embedding 2B/8B) in sequential recommendation. Motivated by the observation that conventional residual+sigmoid side adapters degrade when more backbone layers are adapted (Fig. 1, Fig. 3), the authors attribute this to unselective residual addition and progressive overwriting of earlier side memory under sigmoid fusion. Stresa replaces these with SHAF (stream reshape, Sinkhorn routing with identity-preserving gate β, and gated fusion with λ; Eqs. 10–15) and ReSA (stream-wise pre/post/residual mixers around a bottleneck MLP; Eqs. 16–20). Structural Observations 1–3 algebraically relate SHAF to plain sigmoid fusion plus a stream correction and ReSA to a residual bottleneck. Empirically, Stresa outperforms Two-Stage frozen embeddings, a standard Side Adapter, and side-adaptation baselines (IISAN, IISAN-Versa, CROSSAN, XSMoE) on MicroLens-50K/100K and Scientific (Tables 2–3), with module ablations (Tables 4–6), layer-wise similarity analysis (Fig. 4), efficiency numbers (Table 7), and public code.

Significance. If the empirical findings hold, the work is a useful systems-level contribution for multimodal sequential recommendation: it shows that full-depth side adaptation of large frozen unified embedding models is viable without end-to-end fine-tuning, and that a structured side branch can improve over the residual/sigmoid design used in recent IISAN-style work. Strengths include multi-dataset and multi-backbone evaluation, leave-one-out full ranking, t-tests over five seeds, progressive module ablations, a depth sweep (Fig. 3), representation-trajectory analysis (Fig. 4), algebraic characterizations of the operators, and released code. The practical deployment story (frozen backbone, offline-cacheable item embeddings, trainable side branch only) is well aligned with large-catalog recommendation constraints. The novelty is architectural rather than theoretical; significance rests on reliable gains under deep adaptation and clearer design principles for side memory transport.

major comments (2)
  1. [Introduction; Secs. 3.3–3.4; Fig. 3; Tables 4–5; Observation 1] Intro and Secs. 3.3–3.4 present residual non-selectivity and progressive sigmoid overwriting as the primary causes of deep side-adapter degradation, with SHAF/ReSA as the structural fixes. Fig. 3 is necessary support (Side Adapter peaks near ~9 layers then declines; Stresa improves to full depth), but it does not isolate those causes from capacity, optimization, or better use of noisy deep states. Stresa jointly introduces stream reshape, Sinkhorn routing, dual gates β/λ, and learnable Hpre/Hpost/Hres mixers (Eqs. 10–20). Tables 4–5 show additive gains for ReSA and SHAF, and Table 5 shows identity preservation matters, yet there is no capacity-matched residual/sigmoid control, no full-depth sweep with β forced to 0 (or R=I), and no identity-residual ReSA depth curve. Observation 1 only rewrites SHAF algebraically; it does not establish overwriting as Side Adapter’s failure mode. Please e
  2. [Fig. 3; Sec. 5.1; Table 7; Sec. 4 Implementation Details] The depth comparison that underpins the central narrative (Fig. 3; also Fig. 1) should report, for each adapted-layer count, the trainable parameter count, bottleneck rank, and whether Side Adapter and Stresa use the same adapted-layer set L and same rank search. Table 7 reports Stresa with fewer parameters (78.06M vs 155.73M) but higher time/FLOPs than Side Adapter under one setting; without linking this to the depth sweep, it is hard to rule out that Side Adapter’s late-depth decline is partly a capacity/optimization artifact of a particular hyperparameter choice rather than residual/sigmoid structure alone. Please make the depth-sweep protocol fully capacity- and hyperparameter-transparent (and, if needed, re-plot under matched budgets).
minor comments (6)
  1. [Sec. 7 Conclusion] Conclusion: typo “earlier represtation preservation fusion” → “representation”.
  2. [Fig. 1] Fig. 1 annotates a “+0.73” performance gap when adapting deeper layers, but Table 2 absolute HR@10 gaps are on the order of ~0.003–0.01. Clarify units (absolute, relative %, or another metric) and ensure the figure matches the tabulated protocol.
  3. [Sec. 3.3; Sec. 4] State default values for number of streams s, stream width d_s, Sinkhorn temperature τ, and Sinkhorn iterations; these free parameters are central to SHAF/ReSA but only partially specified relative to rank and learning-rate search.
  4. [Headers / footer] Venue placeholders (“Conference’17, July 2017, Washington, DC, USA”) should be removed or replaced for the camera-ready version.
  5. [Sec. 5.2; Table 3] When comparing to multi-tower baselines (Table 3), a short explicit statement of which backbone towers/modalities each baseline used under “best original settings” would improve reproducibility of the cross-architecture comparison.
  6. [Secs. 3.3–3.4] Notation: H_k is used both for reshaped backbone streams (Eq. 10) and for ReSA mixing operators Hpre/Hpost/Hres; the text notes the distinction, but a different symbol for mixers would reduce confusion.

Circularity Check

0 steps flagged

No significant circularity: empirical side-adapter design evaluated on held-out ranking metrics; Observations 1–3 are algebraic rewrites of the authors’ own operators, not forced predictions.

full rationale

Stresa is an architecture-and-ablation paper. Its load-bearing claims (full-depth side adaptation works better with SHAF+ReSA; gains over Side Adapter, IISAN, CROSSAN, etc.) are supported by leave-one-out HR/NDCG on MicroLens and Scientific under frozen Qwen3-VL backbones (Tables 2–3, Fig. 3), not by identities that force the metrics. SHAF and ReSA are proposed designs (Eqs. 9–20); Observations 1–3 only rewrite those operators (plain sigmoid fusion plus a stream correction; ReSA reduces to residual bottleneck when mixers are identity) and are explicitly labeled as forward-mapping characterizations, not optimization guarantees or data-derived predictions. Self-citations to IISAN/IISAN-Versa/CROSSAN establish the side-adaptation baseline program and experimental protocol; they are not used as uniqueness theorems or as proof that Stresa must outperform. Inspiration from Hyper-Connections and Sinkhorn is design citation, not smuggled ansatz that equates input to claimed result. No fitted parameter is renamed as a first-principles prediction. Mechanism isolation (capacity vs. residual selection/memory preservation) is a correctness/causal concern, not circularity. Derivation chain is self-contained against external benchmarks.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 3 invented entities

The central claim is architectural and empirical: better deep side adaptation via stream-aware fusion and residual stream adapters. It inherits standard recsys evaluation assumptions and frozen multimodal backbone utility, introduces a few design hyperparameters (streams, bottleneck rank, Sinkhorn settings), and invents SHAF/ReSA as modules rather than physical entities. No parameter-free theoretical prediction is claimed.

free parameters (5)
  • bottleneck adapter rank r
    Selected from {128, 256, 512, 1024}; capacity of the side MLP that directly affects performance.
  • number of streams s (and stream width d_s)
    Stream decomposition is a design choice that structures SHAF routing and ReSA mixers; not uniquely determined by theory.
  • adapter learning rate
    Searched over {1e-5, 1e-4, 1e-3}; set to 1e-4 for reported runs.
  • adapted layer set L / number of adapted layers
    Which backbone hidden states enter the side branch is a free experimental choice; depth curves depend on it.
  • Sinkhorn temperature τ and iteration count
    Controls softness/balance of stream routing matrix R_k; implementation detail that affects SHAF behavior.
axioms (5)
  • domain assumption Frozen large multimodal embedding hidden states are useful offline-cacheable features for sequential recommendation once lightly adapted.
    Stated throughout Introduction and Problem Formulation; inherited from IISAN-style side adaptation.
  • domain assumption Next-item ranking under leave-one-out full-item evaluation with HR/NDCG is the right success criterion for the method.
    Evaluation Protocol §4; standard in sequential recommendation but not forced by the architecture.
  • ad hoc to paper Deep side-adapter degradation is mainly due to unselective residual addition and progressive sigmoid overwriting of earlier side memory.
    Core causal diagnosis in Introduction and Fig. 1; motivates SHAF/ReSA specifically.
  • ad hoc to paper Stream reshape of a d-dimensional vector into s streams is a valid structured coordinate system for fusion/residual operators.
    Sec. 3.3 explicitly calls this a lightweight structured coordinate system rather than a semantic claim.
  • standard math Linear algebra of residual bottlenecks, sigmoid gates, and Sinkhorn-normalized soft routing is valid.
    Used in SHAF/ReSA definitions and Observations 1–3.
invented entities (3)
  • Stresa framework no independent evidence
    purpose: Overall stream-aware side-adaptation pipeline separating memory transport (SHAF) from state refinement (ReSA).
    Paper-specific architecture name and composition; evaluated only in this work.
  • SHAF (Stream-aware Hidden-Adapter Fusion) no independent evidence
    purpose: Identity-preserving stream routing of historical side memory plus gated fusion with current backbone hidden states.
    New module introduced in Sec. 3.3; no external independent validation beyond this paper’s ablations.
  • ReSA (Residual Stream Adapter) no independent evidence
    purpose: Stream-wise pre/post/residual mixers around a bottleneck MLP for selective residual updates.
    New module in Sec. 3.4 inspired by Hyper-Connections but specialized to side adapters.

pith-pipeline@v1.1.0-grok45 · 21387 in / 3504 out tokens · 36547 ms · 2026-07-14T08:21:35.630939+00:00 · methodology

0 comments
read the original abstract

Recently, large pretrained multimodal embedding models such as Qwen3-VL Embedding have shown strong promise for sequential recommendation, as they provide reusable semantic item representations across modalities and domains. However, directly using these embeddings often leads to suboptimal performance because of domain misalignment. Efficient side adaptation is therefore an attractive solution. Although adapting all backbone layers should help, existing side adapters often degrade with depth, prompting layer dropping despite the loss of useful hidden states. This is due to two major challenges: (1) the lack of modeling in selecting fused representations during residual addition, and (2) the insufficient preservation of earlier representations during progressive sigmoid fusion. This paper therefore asks a practical question: How can we design a side adaptation approach that effectively unlocks the potential of large pre-trained multimodal embedding models? To address this question, we propose Stresa, a stream-aware side-adaptation framework for frozen large pre-trained multimodal embedding models in sequential recommendation. Stresa introduces Stream-aware Hidden-Adapter Fusion (SHAF) to preserve historical side memory during fusion and Residual Stream Adapter (ReSA) to produce selective residual updates across layers. Empirically, Stresa consistently outperforms standard side adapters and state-of-the-art baselines on public datasets across multiple backbone embedding models. These results highlight the promise of adapting large embedding models for sequential recommendation. Our code is publicly available at https://github.com/GAIR-Lab/Stresa.

Figures

Figures reproduced from arXiv: 2607.10909 by Ioannis Arapakis, Joemon M. Jose, Junchen Fu, Kaiwen Zheng, Wenhao Deng, Xin Xin, Xuri Ge.

Figure 1
Figure 1. Figure 1: Conventional Side Adapter vs. Stresa. Using the [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of the proposed stream-aware fusion framework. (1) Stream-aware Hidden-Adapter Fusion (SHAF) at layer [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: NDCG@10 and NDCG@20 on MicroLens-100K (Qwen3-VL-8B) vs. the number of adapted layers of hidden states, for Side Adapter and Stresa. Metrics are taken at the epoch with the best validation HR@10. verify our motivation of adapting more of the backbone’s potential hidden representations, we further vary the number of adapted hidden-state layers and compare the resulting trends between Side Adapter and Stresa … view at source ↗
Figure 4
Figure 4. Figure 4: Layer-wise cosine similarity on MicroLens-100K for [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

59 extracted references · 10 linked inside Pith

  1. [1]

    Ryan Prescott Adams and Richard S Zemel. 2011. Ranking via sinkhorn propaga- tion.arXiv preprint arXiv:1106.1925(2011)

  2. [2]

    Keqin Bao, Jizhi Zhang, Yang Zhang, Wenjie Wang, Fuli Feng, and Xiangnan He. 2023. Tallrec: An effective and efficient tuning framework to align large language model with recommendation. InProceedings of the 17th ACM conference on recommender systems. 1007–1014

  3. [3]

    Xiaoxi Cui, Weihai Lu, Yu Tong, Yiheng Li, and Zhejun Zhao. 2025. Multi-modal multi-behavior sequential recommendation with conditional diffusion-based feature denoising. InProceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval. 1593–1602

  4. [4]

    Ji Dai, Quan Fang, Jun Hu, Desheng Cai, Yang Yang, and Can Zhao. 2026. Cross- Modal Attention Network with Dual Graph Learning in Multimodal Recom- mendation.ACM Transactions on Multimedia Computing, Communications and Applications(2026)

  5. [5]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers). 4171–4186

  6. [6]

    Junchen Fu, Xuri Ge, Alexandros Karatzoglou, Ioannis Arapakis, Suzan Ver- berne, Joemon M Jose, and Zhaochun Ren. 2026. Differentiable Semantic ID for Generative Recommendation.arXiv preprint arXiv:2601.19711(2026)

  7. [7]

    Junchen Fu, Xuri Ge, Xin Xin, Alexandros Karatzoglou, Ioannis Arapakis, Jie Wang, and Joemon M Jose. 2024. IISAN: Efficiently adapting multimodal repre- sentation for sequential recommendation with decoupled PEFT. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. 687–697

  8. [8]

    Junchen Fu, Xuri Ge, Xin Xin, Alexandros Karatzoglou, Ioannis Arapakis, Kaiwen Zheng, Yongxin Ni, and Joemon M Jose Joemon. 2025. Efficient and effective adaptation of multimodal foundation models in sequential recommendation.IEEE Transactions on Knowledge and Data Engineering(2025)

  9. [9]

    Junchen Fu, Yongxin Ni, Joemon M Jose, Ioannis Arapakis, Kaiwen Zheng, Youhua Li, and Xuri Ge. 2025. Crossan: Towards efficient and effective adaptation of multiple multimodal foundation models for sequential recommendation.arXiv preprint arXiv:2504.10307(2025)

  10. [10]

    Junchen Fu, Fajie Yuan, Yu Song, Zheng Yuan, Mingyue Cheng, Shenghui Cheng, Jiaqi Zhang, Jie Wang, and Yunzhu Pan. 2024. Exploring adapter-based transfer learning for recommender systems: Empirical studies and practical insights. In Proceedings of the 17th ACM international conference on web search and data mining. 208–217

  11. [11]

    Shijie Geng, Juntao Tan, Shuchang Liu, Zuohui Fu, and Yongfeng Zhang. 2023. Vip5: Towards multimodal foundation models for recommendation. InFindings of the Association for Computational Linguistics: EMNLP 2023. 9606–9620

  12. [12]

    Xu Guo, Tong Zhang, Fuyun Wang, Xudong Wang, Xiaoya Zhang, Xin Liu, and Zhen Cui. 2025. Mmhcl: Multi-modal hypergraph contrastive learning for recommendation.ACM Transactions on Multimedia Computing, Communications and Applications21, 10 (2025), 1–23

  13. [13]

    Xu Guo, Tong Zhang, Yufei Xue, Chenxu Wang, Fuyun Wang, and Zhen Cui

  14. [14]

    InICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

    M 3 rec: Selective state space models with mixture-of-modality experts for multi-modal sequential recommendation. InICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1–5

  15. [15]

    Yaoqin He, Junchen Fu, Kaiwen Zheng, Songpei Xu, Fuhai Chen, Jie Li, Joe- mon M Jose, and Xuri Ge. 2025. Double-filter: Efficient fine-tuning of pre-trained vision-language models via patch&layer filtering. InForty-second International Conference on Machine Learning

  16. [16]

    Yupeng Hou, Shanlei Mu, Wayne Xin Zhao, Yaliang Li, Bolin Ding, and Ji-Rong Wen. 2022. Towards universal sequence representation learning for recommender systems. InProceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining. 585–593

  17. [17]

    Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. InInternational conference on machine learning. PMLR, 2790–2799

  18. [18]

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Liang Wang, Weizhu Chen, et al. 2022. Lora: Low-rank adaptation of large language models.Iclr1, 2 (2022), 3

  19. [19]

    Jiaxi Hu, Jingtong Gao, Xiangyu Zhao, Yuehong Hu, Yuxuan Liang, Yiqi Wang, Ming He, Zitao Liu, and Hongzhi Yin. 2024. BiVRec: Bidirectional view-based multimodal sequential recommendation.arXiv preprint arXiv:2402.17334(2024)

  20. [20]

    Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling up visual and vision- language representation learning with noisy text supervision. InInternational conference on machine learning. PMLR, 4904–4916

  21. [21]

    Wang-Cheng Kang and Julian McAuley. 2018. Self-attentive sequential recom- mendation. In2018 IEEE international conference on data mining (ICDM). IEEE, 197–206

  22. [22]

    Sein Kim, Hongseok Kang, Kibum Kim, Jiwan Kim, Donghyun Kim, Minchul Yang, Kwangjin Oh, Julian McAuley, and Chanyoung Park. 2025. Lost in Sequence: Do Large Language Models Understand Sequential Recommendation?. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V

  23. [23]

    Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. InProceedings of the 2021 conference on empirical methods in natural language processing. 3045–3059

  24. [24]

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. InInternational conference on machine learning. PMLR, 19730–19742

  25. [25]

    Mingxin Li, Yanzhao Zhang, Dingkun Long, Keqin Chen, Sibo Song, Shuai Bai, Zhibo Yang, Pengjun Xie, An Yang, Dayiheng Liu, et al . 2026. Qwen3-VL- Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking.arXiv preprint arXiv:2601.04720(2026)

  26. [26]

    Ruyu Li, Wenhao Deng, Yu Cheng, Zheng Yuan, Jiaqi Zhang, and Fajie Yuan

  27. [27]

    InProceedings of the 34th ACM International Conference on Information and Knowledge Management

    Exploring the upper limits of text-based collaborative filtering using large language models: Discoveries and insights. InProceedings of the 34th ACM International Conference on Information and Knowledge Management. 1643–1653

  28. [28]

    Zilong Li, Jia Zhu, Chenglei Huang, Zhangze Chen, Hanghui Guo, Guoqing Ma, and Jianxia Ling. 2026. Capturing Dynamic User Interests Under Modality Imbalance for Multimodal Sequential Recommendation. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 15198–15206

  29. [29]

    Jiahao Liang, Xiangyu Zhao, Muyang Li, Zijian Zhang, Wanyu Wang, Haochen Liu, and Zitao Liu. 2023. Mmmlp: Multi-modal multilayer perceptron for sequen- tial recommendations. InProceedings of the ACM Web Conference 2023. 1109–1117

  30. [30]

    Qijiong Liu, Jieming Zhu, Yanting Yang, Quanyu Dai, Zhaocheng Du, Xiao-Ming Wu, Zhou Zhao, Rui Zhang, and Zhenhua Dong. 2024. Multimodal pretraining, adaptation, and generation for recommendation: A survey. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 6566–6576

  31. [31]

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach.arXiv preprint arXiv:1907.11692 (2019)

  32. [32]

    Yifan Liu, Kangning Zhang, Xiangyuan Ren, Yanhua Huang, Jiarui Jin, Yingjie Qin, Ruilong Su, Ruiwen Xu, Yong Yu, and Weinan Zhang. 2024. Alignrec: Aligning and training in multimodal recommendations. InProceedings of the 33rd ACM International Conference on Information and Knowledge Management. 1503–1512

  33. [33]

    Weihai Lu and Li Yin. 2025. Dmmd4sr: Diffusion model-based multi-level multi- modal denoising for sequential recommendation. InProceedings of the 33rd ACM International Conference on Multimedia. 6363–6372

  34. [34]

    Fanshen Meng, Zhenhua Meng, Ru Jin, Yuli Chen, Rongheng Lin, and Budan Wu. 2025. TAMER: Interest Tree Augmented Modality Graph Recommender for Multimodal Recommendation. InProceedings of the 33rd ACM International Conference on Multimedia. 5998–6006

  35. [35]

    Yongxin Ni, Yu Cheng, Xiangyan Liu, Junchen Fu, Youhua Li, Xiangnan He, Yongfeng Zhang, and Fajie Yuan. 2025. A content-driven micro-video recommen- dation dataset at scale. InProceedings of the 34th ACM International Conference on Information and Knowledge Management. 6486–6491

  36. [36]

    Yunke Qu, Liang Qu, Tong Chen, Quoc Viet Hung Nguyen, and Hongzhi Yin. 2025. Efficient multimodal streaming recommendation via expandable side mixture-of- experts. InProceedings of the 34th ACM International Conference on Information and Knowledge Management. 2460–2470

  37. [37]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. InInternational conference on machine learning. PmLR, 8748–8763

  38. [38]

    Te Song, Lianyong Qi, Weiming Liu, Fan Wang, Xiaolong Xu, Hongsheng Hu, Yang Cao, Xuyun Zhang, and Amin Beheshti. 2025. Boosting Guided Diffusion with Large Language Models for Multimodal Sequential Recommendation. In Proceedings of the 33rd ACM International Conference on Multimedia. 6203–6212

  39. [39]

    Weiwei Sun, Keyi Kong, Xinyu Ma, Shuaiqiang Wang, Dawei Yin, Maarten de Rijke, Zhaochun Ren, and Yiming Yang. 2026. ZeroGR: A Generalizable and Scalable Framework for Zero-Shot Generative Retrieval. InThe Fourteenth Inter- national Conference on Learning Representations. https://openreview.net/forum? id=RBoAwiQl5L

  40. [40]

    Yi-Lin Sung, Jaemin Cho, and Mohit Bansal. 2022. Lst: Ladder side-tuning for parameter and memory efficient transfer learning.Advances in Neural Information Processing Systems35 (2022), 12991–13005

  41. [41]

    Hanbing Wang, Xiaorui Liu, Wenqi Fan, Xiangyu Zhao, Venkataramana Kini, Devendra Pratap Yadav, Fei Wang, Zhen Wen, and Hui Liu. 2025. Rethinking large language model architectures for sequential recommendations. InProceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the...

  42. [42]

    Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2022. Text embeddings by weakly-supervised Conference’17, July 2017, Washington, DC, USA Fu et al. contrastive pre-training.arXiv preprint arXiv:2212.03533(2022)

  43. [43]

    Yuhao Wang, Junwei Pan, Xinhang Li, Maolin Wang, Yuan Wang, Yue Liu, Dapeng Liu, Jie Jiang, and Xiangyu Zhao. 2025. Empowering large language model for sequential recommendation via multimodal embeddings and semantic ids. InProceedings of the 34th ACM International Conference on Information and Knowledge Management. 3209–3219

  44. [44]

    Zhaoliang Wang, Baisong Liu, Weiming Huang, Tingting Hao, Huiqian Zhou, and Yuxin Guo. 2025. Leveraging multimodal large language model for multimodal sequential recommendation.Scientific Reports15, 1 (2025), 28960

  45. [45]

    Zhenda Xie, Yixuan Wei, Huanqi Cao, Chenggang Zhao, Chengqi Deng, Jiashi Li, Damai Dai, Huazuo Gao, Jiang Chang, Kuai Yu, et al. 2025. mhc: Manifold- constrained hyper-connections.arXiv preprint arXiv:2512.24880(2025)

  46. [46]

    Jinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li, Hewei Wang, and Edith CH Ngai

  47. [47]

    InProceedings of the AAAI Conference on Artificial Intelligence, Vol

    Mentor: multi-level self-supervised learning for multimodal recommen- dation. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 12908–12917

  48. [48]

    Mengde Xu, Zheng Zhang, Fangyun Wei, Han Hu, and Xiang Bai. 2023. Side adapter network for open-vocabulary semantic segmentation. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2945–2954

  49. [49]

    Wei Yang, Rui Zhong, Yiqun Chen, Shixuan Li, Heng Ping, Chi Lu, and Peng Jiang. 2025. FITMM: Adaptive Frequency-Aware Multimodal Recommendation via Information-Theoretic Representation Learning. InProceedings of the 33rd ACM International Conference on Multimedia. 6193–6202

  50. [50]

    Yu Ye, Junchen Fu, Yu Song, Kaiwen Zheng, and Joemon M Jose. 2026. Are multimodal embeddings truly beneficial for recommendation? A deep dive into whole vs. individual modalities. InEuropean Conference on Information Retrieval. Springer, 66–81

  51. [51]

    Yuyang Ye, Zhi Zheng, Yishan Shen, Tianshu Wang, Hengruo Zhang, Peijun Zhu, Runlong Yu, Kai Zhang, and Hui Xiong. 2025. Harnessing multimodal large language models for multimodal sequential recommendation. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 13069–13077

  52. [52]

    Zheng Yuan, Fajie Yuan, Yu Song, Youhua Li, Junchen Fu, Fei Yang, Yunzhu Pan, and Yongxin Ni. 2023. Where to go next for recommender systems? id- vs. modality-based recommender models revisited. InProceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2639–2649

  53. [53]

    Lingzi Zhang, Xin Zhou, Zhiwei Zeng, and Zhiqi Shen. 2024. Multimodal pre- training for sequential recommendation via contrastive learning.ACM Transac- tions on Recommender Systems3, 1 (2024), 1–23

  54. [54]

    Shengzhe Zhang, Liyi Chen, Dazhong Shen, Chao Wang, and Hui Xiong. 2025. Hierarchical time-aware mixture of experts for multi-modal sequential recom- mendation. InProceedings of the ACM on Web Conference 2025. 3672–3682

  55. [55]

    Zhengxin Zhang, Dan Zhao, Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Qing Li, Yong Jiang, and Zhihao Jia. 2024. Quantized side tuning: Fast and memory- efficient tuning of quantized large language models. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 1–17

  56. [56]

    Zhicheng Zhou, Xiangwu Meng, and Yujie Zhang. 2025. Rethinking Convo- lutional Neural Network in Multimodal Sequential Recommendation.ACM Transactions on Information Systems44, 2 (2025), 1–35

  57. [57]

    Defa Zhu, Hongzhi Huang, Zihao Huang, Yutao Zeng, Yunyao Mao, Banggu Wu, Qiyang Min, and Xun Zhou. 2024. Hyper-connections.arXiv preprint arXiv:2409.19606(2024)

  58. [58]

    Jing Zhu, Mingxuan Ju, Yozen Liu, Danai Koutra, Neil Shah, and Tong Zhao. 2025. Beyond unimodal boundaries: Generative recommendation with multimodal semantics.arXiv preprint arXiv:2503.23333(2025)

  59. [59]

    Ziyi Zhuang, Hongji Li, Junchen Fu, Jiacheng Liu, Joemon M Jose, Youhua Li, and Yongxin Ni. 2025. Frequency-Decoupled distillation for efficient multimodal recommendation. InProceedings of the 34th ACM International Conference on Information and Knowledge Management. 4571–4581