Pith. sign in

REVIEW 2 cited by

N24News: A New Dataset for Multimodal News Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.13327 v4 pith:M7EV2M5V submitted 2021-08-30 cs.CL

classification cs.CL
keywords newsclassificationmultimodaln24newstextdatasetfeaturesaccuracy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current news datasets merely focus on text features on the news and rarely leverage the feature of images, excluding numerous essential features for news classification. In this paper, we propose a new dataset, N24News, which is generated from New York Times with 24 categories and contains both text and image information in each news. We use a multitask multimodal method and the experimental results show multimodal news classification performs better than text-only news classification. Depending on the length of the text, the classification accuracy can be increased by up to 8.11%. Our research reveals the relationship between the performance of a multimodal classifier and its sub-classifiers, and also the possible improvements when applying multimodal in news classification. N24News is shown to have great potential to prompt the multimodal news studies.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning

    cs.IR 2026-03 conditional novelty 5.0 of 10

    CoCoA forces an MLLM to reconstruct masked text through a single EOS token, improving multimodal embedding quality on MMEB-V1 and matching MoCa at 3B with far less pretraining data.

  2. Visual Semantic Description Generation with MLLMs for Image-Text Matching

    cs.MM 2025-07 conditional novelty 5.0 of 10

    Adding MLLM-generated visual semantic descriptions, fused at instance and prototype levels, improves image-text retrieval across multiple baselines and domains.

Pith tools