Pith. sign in

REVIEW 5 cited by

VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.04923 v3 pith:U73BM72Z submitted 2024-11-07 cs.CV

classification cs.CV
keywords groundingvideosfine-grainedlargemodelmultimodalpixel-levelspatial
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Fine-grained alignment between videos and text is challenging due to complex spatial and temporal dynamics in videos. Existing video-based Large Multimodal Models (LMMs) handle basic conversations but struggle with precise pixel-level grounding in videos. To address this, we introduce VideoGLaMM, a LMM designed for fine-grained pixel-level grounding in videos based on user-provided textual inputs. Our design seamlessly connects three key components: a Large Language Model, a dual vision encoder that emphasizes both spatial and temporal details, and a spatio-temporal decoder for accurate mask generation. This connection is facilitated via tunable V-L and L-V adapters that enable close Vision-Language (VL) alignment. The architecture is trained to synchronize both spatial and temporal elements of video content with textual instructions. To enable fine-grained grounding, we curate a multimodal dataset featuring detailed visually-grounded conversations using a semiautomatic annotation pipeline, resulting in a diverse set of 38k video-QA triplets along with 83k objects and 671k masks. We evaluate VideoGLaMM on three challenging tasks: Grounded Conversation Generation, Visual Grounding, and Referring Video Segmentation. Experimental results show that our model consistently outperforms existing approaches across all three tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Promptception: How Sensitive Are Large Multimodal Models to Prompts?

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A study of 61 prompt variants across 10 vision-language models and 3 benchmarks finds accuracy swings of up to 15 points, with proprietary models more sensitive than open-source ones.

  2. MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning

    cs.CV 2025-08 reject novelty 6.0 of 10

    A new marine wildlife video dataset with clip-level captions and grounded segmentation masks, benchmarked on captioning, grounding, and video generation.

  3. VideoMolmo: Spatio-Temporal Grounding Meets Pointing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A video language model that conditions each frame on earlier frames via a temporal attention module, predicts text-requested object points, and uses SAM2-based bidirectional mask fusion to outperform prior models on v...

  4. InterRVOS: Interaction-aware Referring Video Object Segmentation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    InterRVOS extends referring video object segmentation to segment both actor and target objects separately for interaction expressions, with a new dataset and MLLM-based method.

  5. PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?

    cs.CV 2025-09 conditional novelty 5.0 of 10

    Video MLLMs mostly ignore motion in pixel-level visual grounding; a new motion-centric benchmark shows large performance drops.

Pith tools