Pith. sign in

hub Canonical reference

MOSEv2: A more challenging dataset for video object segmentation in complex scenes

Canonical reference. 83% of citing Pith papers cite this work as background.

17 Pith papers citing it
Background 83% of classified citations

hub tools

citation-role summary

background 5 dataset 1

citation-polarity summary

fields

cs.CV 16 cs.SD 1

years

2026 15 2025 2

representative citing papers

FeVOS: Foresight Expression Video Object Segmentation

cs.CV · 2026-06-24 · unverdicted · novelty 7.0

Introduces the FeVOS task, a 968-clip dataset with foresight expressions, and an MLLM model FeVOS-R1 trained via SFT then RL that reports SOTA on the new task plus generalization to prior RVOS benchmarks.

Semantic Alignment in Hyperbolic Space for Open-Vocabulary Semantic Segmentation

cs.CV · 2026-05-09 · unverdicted · novelty 7.0

HyRo decouples hierarchical and semantic alignment in hyperbolic space by adjusting Poincaré ball radius for hierarchy and applying radius-preserving orthogonal transformation for semantics, yielding state-of-the-art results on open-vocabulary semantic segmentation benchmarks.

3AM: 3egment Anything with Geometric Consistency in Videos

cs.CV · 2026-01-13 · unverdicted · novelty 7.0

3AM integrates MUSt3R 3D features into SAM2 via a Feature Merger and FOV-aware sampling to deliver geometry-consistent video object segmentation from RGB alone, with large gains on wide-baseline datasets.

SAM 3: Segment Anything with Concepts

cs.CV · 2025-11-20 · unverdicted · novelty 7.0

SAM 3 introduces promptable concept segmentation that doubles accuracy of prior systems on images and videos while improving standard SAM segmentation performance.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation

cs.CV · 2026-07-09 · conditional · novelty 6.0

A modified video segmentation architecture decouples processing latency from target count, enabling real-time (>36 FPS) tracking of 10+ objects simultaneously while preserving individual identities.

VISTA: Video Interaction Spatio-Temporal Analysis Benchmark

cs.CV · 2026-05-02 · unverdicted · novelty 6.0 · 2 refs

VISTA is a new ~12K-pair benchmark and taxonomy for open-set multi-entity spatio-temporal understanding in VLMs that decomposes videos into entities, actions, and relational dynamics for multi-axis diagnostics.

2nd of the 5th PVUW MeViS-Audio Track: ASR-SaSaSa2VA

cs.CV · 2026-04-27 · unverdicted · novelty 3.0

ASR-SaSaSa2VA turns audio into text via ASR then feeds it to pre-trained referring video segmentation models, achieving 80.7 and second place in the 5th PVUW MeViS-v2-Audio track.

OAMVOS:2nd Report for 5th PVUW MOSE Track

cs.CV · 2026-04-20 · unverdicted · novelty 3.0

An occlusion-aware extension to DAM4SAM adds a reliability state machine, branch-based recovery, delayed memory promotion, and selective native memory rules to improve robustness under long occlusions and reappearances without altering the backbone.

APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track

cs.SD · 2026-04-20 · unverdicted · novelty 3.0

A staged pipeline using ASR transcription, visual existence verification, Sa2VA coarse segmentation, and agent-guided SAM3 refinement won first place in the PVUW MeViS-Audio track by decomposing audio-conditioned Ref-VOS into sequential verification and refinement steps.

citing papers explorer

Showing 17 of 17 citing papers.