Pith. sign in

REVIEW 2 cited by

CROSSAN: Towards Efficient and Effective Adaptation of Multiple Multimodal Foundation Models for Sequential Recommendation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.10307 v2 pith:G5EGWENB submitted 2025-04-14 cs.IR

classification cs.IR
keywords crossanefficientfoundationmodelsmultimodaladaptationcross-modaleffective
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

In this paper, we explore a less-studied yet practically important problem: how to efficiently and effectively adapt multiple ($>$2) multimodal foundation models (MFMs) for the sequential recommendation task. To this end, we propose a plug-and-play Cross-modal Side Adapter Network (CROSSAN), which leverages a fully decoupled side adapter-based paradigm to achieve efficient and scalable adaptation. Compared to the state-of-the-art efficient approaches, CROSSAN reduces training time by over 30%, GPU memory consumption by 20%, and trainable parameters by over 57%, while enabling effective cross-modal learning across diverse modalities. To further enhance multimodal fusion, we introduce the Mixture of Modality Expert Fusion (MOMEF) mechanism. Extensive experiments on public benchmarks demonstrate that CROSSAN consistently outperforms existing methods, achieving 6.7%--8.1% performance improvements when adapting four foundation models with raw modalities. Moreover, the overall performance continues to improve as more MFMs are incorporated. We will release our code and datasets to faciliate future research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Stream-aware Side Adaptation for Large Pre-trained Multimodal Embedding Models in Sequential Recommendation

    cs.IR 2026-07 conditional novelty 5.0 of 10

    Stream-aware fusion (SHAF) and residual stream adapters (ReSA) stabilize deep side adaptation of frozen multimodal embedding models and improve sequential recommendation over standard side adapters.

  2. Role of Personality in Conversational Information Seeking

    cs.IR 2026-08 conditional novelty 4.0 of 10

    In conversational information seeking, assistant personality effects are task-dependent, and trust depends on both task style fit and user-assistant personality compatibility.

Pith tools