REVIEW 10 cited by
A Comprehensive Survey on Multimodal Recommender Systems: Taxonomy, Evaluation, and Future Directions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recommendation systems have become popular and effective tools to help users discover their interesting items by modeling the user preference and item property based on implicit interactions (e.g., purchasing and clicking). Humans perceive the world by processing the modality signals (e.g., audio, text and image), which inspired researchers to build a recommender system that can understand and interpret data from different modalities. Those models could capture the hidden relations between different modalities and possibly recover the complementary information which can not be captured by a uni-modal approach and implicit interactions. The goal of this survey is to provide a comprehensive review of the recent research efforts on the multimodal recommendation. Specifically, it shows a clear pipeline with commonly used techniques in each step and classifies the models by the methods used. Additionally, a code framework has been designed that helps researchers new in this area to understand the principles and techniques, and easily runs the SOTA models. Our framework is located at: https://github.com/enoche/MMRec
Forward citations
Cited by 10 Pith papers
-
One Graph, Multiple Gains: Single High-Quality Item-Item Graph for Multimodal Recommendation
A single NCER-refined item-item graph, reused via adaptive gating, UI expansion, and discounted soft-positive BPR, improves multimodal recommendation accuracy and efficiency.
-
FASH-iCNN: Making Editorial Fashion Identity Inspectable Through Multimodal CNN Probing
A multimodal CNN on 87,547 Vogue images classifies fashion houses at 78.2% top-1 accuracy, decades at 88.6%, and years at 58.3% with 2.2-year mean error, and shows texture and luminance carry most of the house-identit...
-
Joint Behavior-guided and Modality-coherence Conditional Graph Diffusion Denoising for Multi Modal Recommendation
JBM-Diff applies conditional graph diffusion to remove preference-irrelevant multimodal noise and false-positive/negative behaviors, then augments training data via partial-order credibility scoring.
-
TRU: Targeted Reverse Update for Efficient Multimodal Recommendation Unlearning
TRU is a plug-and-play unlearning method for multimodal recommenders that applies ranking fusion, modality scaling, and layer isolation to achieve better retain-forget trade-offs than uniform baselines.
-
Binge Watch: Reproducible Multimodal Benchmarks Datasets for Large-Scale Movie Recommendation on MovieLens-10M and 20M
M3L-10M and M3L-20M add plot, poster, audio, and video embeddings to MovieLens and release them publicly as reproducible multimodal benchmarks.
-
Hi-SAM: A Hierarchical Structure-Aware Multi-modal Framework for Large-Scale Recommendation
Hi-SAM improves semantic-ID multimodal recommendation by disentangling shared versus modality-specific item codes and by letting transformers access history only through compressed anchor tokens.
-
URecJPQ: Memory-efficient Multimodal Recommendation Models through RecJPQ in Large-Scale Scenarios
URecJPQ compresses user and item embeddings via joint product quantization for multimodal top-k recommendation, cutting checkpoint size 86-98% and parameters 98-99% with average 8.5% recall drop across three datasets.
-
Modality-Aware Identity Construction and Counterfactual Structure Learning for ID-Free Multimodal Recommendation
MAIL constructs modality-aware ID-free identities via dynamic positional encoding modulation and applies counterfactual structure learning with popularity penalization, yielding 7.81% Recall@10 and 12.81% NDCG@10 gain...
-
Modality Alignment with Multi-scale Bilateral Attention for Multimodal Recommendation
MambaRec improves multimodal recommendation accuracy on Baby, Sports, and Clothing datasets through local dilated-attention alignment and global MMD/contrastive alignment.
-
A Scenario-Oriented Survey of Federated Recommender Systems: Techniques, Challenges, and Future Directions
A scenario-oriented taxonomy of federated recommender systems that argues research should be organized around recommendation use cases rather than federated-learning abstractions.
Discussion (0). Sign in to comment.