Pith. sign in

REVIEW 2 cited by

MM-Rec: Multimodal News Recommendation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.07407 v2 pith:A2ELE2N5 submitted 2021-04-15 cs.IR

classification cs.IR
keywords newsimagesmultimodalrecommendationinformationaccurateclickedcrossmodal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Accurate news representation is critical for news recommendation. Most of existing news representation methods learn news representations only from news texts while ignore the visual information in news like images. In fact, users may click news not only because of the interest in news titles but also due to the attraction of news images. Thus, images are useful for representing news and predicting user behaviors. In this paper, we propose a multimodal news recommendation method, which can incorporate both textual and visual information of news to learn multimodal news representations. We first extract region-of-interests (ROIs) from news images via object detection. Then we use a pre-trained visiolinguistic model to encode both news texts and news image ROIs and model their inherent relatedness using co-attentional Transformers. In addition, we propose a crossmodal candidate-aware attention network to select relevant historical clicked news for accurate user modeling by measuring the crossmodal relatedness between clicked news and candidate news. Experiments validate that incorporating multimodal news information can effectively improve news recommendation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GALA: Generative Aligned Learning for Adaptive Multimodal Representation in the Taobao Shangou Recommender System

    cs.IR 2026-07 conditional novelty 6.0 of 10

    A three-stage multimodal recommender pipeline with GRPO-based behavior alignment and adaptive ID-content fusion claims a 0.55% online order-volume increase and small offline AUC gains at Taobao Shangou.

  2. Privacy-Preserving Multimodal News Recommendation through Federated Learning

    cs.SI 2025-07 reject novelty 4.0 of 10

    A multimodal federated news recommender that fuses BERT text and ViT image features with long- and short-term user modeling, plus Shamir-secret-sharing secure aggregation, reports AUC 0.698 on MIND data.

Pith tools