A pre-trained alignment model is adapted for multimodal coreference resolution via similarity aggregation and evidence theory, reporting gains on the CIN benchmark over prior dedicated methods and VLLMs.
Pixels to prose: Understanding the art of image captioning,
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
UNVERDICTED 2representative citing papers
An LLM-and-TTS pipeline produces art descriptions with higher lexical diversity, adjective density, and narrative detail than baseline captions on 50 artworks, at low cost and speed.
citing papers explorer
-
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model
A pre-trained alignment model is adapted for multimodal coreference resolution via similarity aggregation and evidence theory, reporting gains on the CIN benchmark over prior dedicated methods and VLLMs.
-
CANVAS: Captioning Art with Narrative Visual-Audio AI Systems
An LLM-and-TTS pipeline produces art descriptions with higher lexical diversity, adjective density, and narrative detail than baseline captions on 50 artworks, at low cost and speed.