Pith. sign in

REVIEW 1 cited by

Visually-Aware Context Modeling for News Image Captioning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.08325 v2 pith:E4HUYFUC submitted 2023-08-16 cs.CV

classification cs.CV
keywords imagecontextnewscaptionsimagesarticlearticlescaptioning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

News Image Captioning aims to create captions from news articles and images, emphasizing the connection between textual context and visual elements. Recognizing the significance of human faces in news images and the face-name co-occurrence pattern in existing datasets, we propose a face-naming module for learning better name embeddings. Apart from names, which can be directly linked to an image area (faces), news image captions mostly contain context information that can only be found in the article. We design a retrieval strategy using CLIP to retrieve sentences that are semantically close to the image, mimicking human thought process of linking articles to images. Furthermore, to tackle the problem of the imbalanced proportion of article context and image context in captions, we introduce a simple yet effective method Contrasting with Language Model backbone (CoLaM) to the training pipeline. We conduct extensive experiments to demonstrate the efficacy of our framework. We out-perform the previous state-of-the-art (without external data) by 7.97/5.80 CIDEr scores on GoodNews/NYTimes800k. Our code is available at https://github.com/tingyu215/VACNIC.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ReCap: Event-Aware Image Captioning with Article Retrieval and Semantic Gaussian Normalization

    cs.CV 2025-09 conditional novelty 4.0 of 10

    A two-stage DINOv2 retrieval and Qwen3 LLM pipeline with a CIDEr-aware length normalizer achieved 2nd place in the EVENTA 2025 event-captioning challenge.

Pith tools