REVIEW 5 cited by
VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Fine-grained alignment between videos and text is challenging due to complex spatial and temporal dynamics in videos. Existing video-based Large Multimodal Models (LMMs) handle basic conversations but struggle with precise pixel-level grounding in videos. To address this, we introduce VideoGLaMM, a LMM designed for fine-grained pixel-level grounding in videos based on user-provided textual inputs. Our design seamlessly connects three key components: a Large Language Model, a dual vision encoder that emphasizes both spatial and temporal details, and a spatio-temporal decoder for accurate mask generation. This connection is facilitated via tunable V-L and L-V adapters that enable close Vision-Language (VL) alignment. The architecture is trained to synchronize both spatial and temporal elements of video content with textual instructions. To enable fine-grained grounding, we curate a multimodal dataset featuring detailed visually-grounded conversations using a semiautomatic annotation pipeline, resulting in a diverse set of 38k video-QA triplets along with 83k objects and 671k masks. We evaluate VideoGLaMM on three challenging tasks: Grounded Conversation Generation, Visual Grounding, and Referring Video Segmentation. Experimental results show that our model consistently outperforms existing approaches across all three tasks.
Forward citations
Cited by 5 Pith papers
-
Promptception: How Sensitive Are Large Multimodal Models to Prompts?
A study of 61 prompt variants across 10 vision-language models and 3 benchmarks finds accuracy swings of up to 15 points, with proprietary models more sensitive than open-source ones.
-
MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning
A new marine wildlife video dataset with clip-level captions and grounded segmentation masks, benchmarked on captioning, grounding, and video generation.
-
VideoMolmo: Spatio-Temporal Grounding Meets Pointing
A video language model that conditions each frame on earlier frames via a temporal attention module, predicts text-requested object points, and uses SAM2-based bidirectional mask fusion to outperform prior models on v...
-
InterRVOS: Interaction-aware Referring Video Object Segmentation
InterRVOS extends referring video object segmentation to segment both actor and target objects separately for interaction expressions, with a new dataset and MLLM-based method.
-
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?
Video MLLMs mostly ignore motion in pixel-level visual grounding; a new motion-centric benchmark shows large performance drops.
Discussion (0). Continue with ORCID to comment.