Molmo2 delivers state-of-the-art open-weight video VLMs with new grounding datasets and training methods that outperform prior open models and match or exceed some proprietary ones on pointing and tracking tasks.
hub Canonical reference
MOSEv2: A more challenging dataset for video object segmentation in complex scenes
Canonical reference. 83% of citing Pith papers cite this work as background.
hub tools
citation-role summary
citation-polarity summary
representative citing papers
Introduces the FeVOS task, a 968-clip dataset with foresight expressions, and an MLLM model FeVOS-R1 trained via SFT then RL that reports SOTA on the new task plus generalization to prior RVOS benchmarks.
HyRo decouples hierarchical and semantic alignment in hyperbolic space by adjusting Poincaré ball radius for hierarchy and applying radius-preserving orthogonal transformation for semantics, yielding state-of-the-art results on open-vocabulary semantic segmentation benchmarks.
Fully end-to-end training with a sentence-conditioned adapter outperforms frozen-backbone baselines for localizing video segments that match sentence queries.
3AM integrates MUSt3R 3D features into SAM2 via a Feature Merger and FOV-aware sampling to deliver geometry-consistent video object segmentation from RGB alone, with large gains on wide-baseline datasets.
SAM 3 introduces promptable concept segmentation that doubles accuracy of prior systems on images and videos while improving standard SAM segmentation performance.
A modified video segmentation architecture decouples processing latency from target count, enabling real-time (>36 FPS) tracking of 10+ objects simultaneously while preserving individual identities.
SAM2Matting decouples tracking from matting to deliver SOTA video matting performance using models trained only on images.
AMNet enables modality-agnostic low-light video enhancement via a Spatial-Spectral Dual-Gated Translator for implicit auxiliary representations and large-scale pretraining on RGB data with synthetic auxiliaries.
VISTA is a new ~12K-pair benchmark and taxonomy for open-set multi-entity spatio-temporal understanding in VLMs that decomposes videos into entities, actions, and relational dynamics for multi-axis diagnostics.
The 2026 PVUW Challenge introduces a new audio track and evaluates top multimodal methods on challenging video datasets for pixel-level understanding.
ASR-SaSaSa2VA turns audio into text via ASR then feeds it to pre-trained referring video segmentation models, achieving 80.7 and second place in the 5th PVUW MeViS-v2-Audio track.
An occlusion-aware extension to DAM4SAM adds a reliability state machine, branch-based recovery, delayed memory promotion, and selective native memory rules to improve robustness under long occlusions and reappearances without altering the backbone.
A staged pipeline using ASR transcription, visual existence verification, Sa2VA coarse segmentation, and agent-guided SAM3 refinement won first place in the PVUW MeViS-Audio track by decomposing audio-conditioned Ref-VOS into sequential verification and refinement steps.
An agent-augmented Sa2VA pipeline for referring video object segmentation placed third in the MeViS-Text track of the 5th PVUW Challenge by adding verification, search, and refinement stages.
A survey that categorizes and summarizes methods applying 3D Gaussian Splatting to segmentation, editing, generation, and related tasks, including datasets and evaluation protocols.
citing papers explorer
-
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
Molmo2 delivers state-of-the-art open-weight video VLMs with new grounding datasets and training methods that outperform prior open models and match or exceed some proprietary ones on pointing and tracking tasks.
-
FeVOS: Foresight Expression Video Object Segmentation
Introduces the FeVOS task, a 968-clip dataset with foresight expressions, and an MLLM model FeVOS-R1 trained via SFT then RL that reports SOTA on the new task plus generalization to prior RVOS benchmarks.
-
Semantic Alignment in Hyperbolic Space for Open-Vocabulary Semantic Segmentation
HyRo decouples hierarchical and semantic alignment in hyperbolic space by adjusting Poincaré ball radius for hierarchy and applying radius-preserving orthogonal transformation for semantics, yielding state-of-the-art results on open-vocabulary semantic segmentation benchmarks.
-
A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos
Fully end-to-end training with a sentence-conditioned adapter outperforms frozen-backbone baselines for localizing video segments that match sentence queries.
-
3AM: 3egment Anything with Geometric Consistency in Videos
3AM integrates MUSt3R 3D features into SAM2 via a Feature Merger and FOV-aware sampling to deliver geometry-consistent video object segmentation from RGB alone, with large gains on wide-baseline datasets.
-
SAM 3: Segment Anything with Concepts
SAM 3 introduces promptable concept segmentation that doubles accuracy of prior systems on images and videos while improving standard SAM segmentation performance.
-
SAM-MT: Real-Time Interactive Multi-Target Video Segmentation
A modified video segmentation architecture decouples processing latency from target count, enabling real-time (>36 FPS) tracking of 10+ objects simultaneously while preserving individual identities.
-
SAM2Matting: Generalized Image and Video Matting
SAM2Matting decouples tracking from matting to deliver SOTA video matting performance using models trained only on images.
-
AnyMod-LLVE: Low-Light Video Enhancement with Modality-Agnostic Inference
AMNet enables modality-agnostic low-light video enhancement via a Spatial-Spectral Dual-Gated Translator for implicit auxiliary representations and large-scale pretraining on RGB data with synthetic auxiliaries.
-
VISTA: Video Interaction Spatio-Temporal Analysis Benchmark
VISTA is a new ~12K-pair benchmark and taxonomy for open-set multi-entity spatio-temporal understanding in VLMs that decomposes videos into entities, actions, and relational dynamics for multi-axis diagnostics.
-
Report of the 5th PVUW Challenge: Towards More Diverse Modalities in Pixel-Level Understanding
The 2026 PVUW Challenge introduces a new audio track and evaluates top multimodal methods on challenging video datasets for pixel-level understanding.
-
2nd of the 5th PVUW MeViS-Audio Track: ASR-SaSaSa2VA
ASR-SaSaSa2VA turns audio into text via ASR then feeds it to pre-trained referring video segmentation models, achieving 80.7 and second place in the 5th PVUW MeViS-v2-Audio track.
-
OAMVOS:2nd Report for 5th PVUW MOSE Track
An occlusion-aware extension to DAM4SAM adds a reliability state machine, branch-based recovery, delayed memory promotion, and selective native memory rules to improve robustness under long occlusions and reappearances without altering the backbone.
-
APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track
A staged pipeline using ASR transcription, visual existence verification, Sa2VA coarse segmentation, and agent-guided SAM3 refinement won first place in the PVUW MeViS-Audio track by decomposing audio-conditioned Ref-VOS into sequential verification and refinement steps.
-
AgentRVOS for MeViS-Text Track of 5th PVUW Challenge: 3rd Method
An agent-augmented Sa2VA pipeline for referring video object segmentation placed third in the MeViS-Text track of the 5th PVUW Challenge by adding verification, search, and refinement stages.
-
A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation
A survey that categorizes and summarizes methods applying 3D Gaussian Splatting to segmentation, editing, generation, and related tasks, including datasets and evaluation protocols.
- Controllable Video Object Insertion via Multi-View Priors