Pith. sign in

REVIEW 3 cited by

GroundCap: A Visually Grounded Image Captioning Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.13898 v3 pith:BLLYWUK4 submitted 2025-02-19 cs.CV cs.CL

classification cs.CVcs.CL
keywords objectactionsgroundcapgroundinglinkingobjectsaction-objectapproach
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Current image captioning systems lack the ability to link descriptive text to specific visual elements, making their outputs difficult to verify. While recent approaches offer some grounding capabilities, they cannot track object identities across multiple references or ground both actions and objects simultaneously. We propose a novel ID-based grounding system that enables consistent object reference tracking and action-object linking. We present GroundCap, a dataset containing 52,016 images from 77 movies, with 344 human-annotated and 52,016 automatically generated captions. Each caption is grounded on detected objects (132 classes) and actions (51 classes) using a tag system that maintains object identity while linking actions to the corresponding objects. Our approach features persistent object IDs for reference tracking, explicit action-object linking, and the segmentation of background elements through K-means clustering. We propose gMETEOR, a metric combining caption quality with grounding accuracy, and establish baseline performance by fine-tuning Pixtral-12B and Qwen2.5-VL 7B on GroundCap. Human evaluation demonstrates our approach's effectiveness in producing verifiable descriptions with coherent object references.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A unified visual grounding framework combining a broadcast cross-attention head, a JEPA auxiliary loss, and an MLLM-generated caption dataset preserves representation diversity and generalizes across RefCOCO/+/g.

  2. StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new multi-frame visual storytelling dataset with explicit entity grounding, plus a fine-tuned Qwen2.5-VL baseline that reduces measured hallucinations by 12.3%.

  3. Entity Re-identification in Visual Storytelling via Contrastive Reinforcement Learning

    cs.CV 2025-07 reject novelty 4.0 of 10

    A contrastive reinforcement learning method with synthetic negative stories improves cross-frame entity grounding and re-identification for a 7B visual storyteller, evaluated only on the authors' own dataset.

Pith tools