Pith. sign in

REVIEW 7 cited by

LVOS: A Benchmark for Large-scale Long-term Video Object Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.19326 v2 pith:WUWCZJ32 submitted 2024-04-30 cs.CV

classification cs.CV
keywords lvosvideomodelsvideosbenchmarksexistinglong-termobjects
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Video object segmentation (VOS) aims to distinguish and track target objects in a video. Despite the excellent performance achieved by off-the-shell VOS models, existing VOS benchmarks mainly focus on short-term videos lasting about 5 seconds, where objects remain visible most of the time. However, these benchmarks poorly represent practical applications, and the absence of long-term datasets restricts further investigation of VOS in realistic scenarios. Thus, we propose a novel benchmark named LVOS, comprising 720 videos with 296,401 frames and 407,945 high-quality annotations. Videos in LVOS last 1.14 minutes on average, approximately 5 times longer than videos in existing datasets. Each video includes various attributes, especially challenges deriving from the wild, such as long-term reappearing and cross-temporal similar objects. Compared to previous benchmarks, our LVOS better reflects VOS models' performance in real scenarios. Based on LVOS, we evaluate 20 existing VOS models under 4 different settings and conduct a comprehensive analysis. On LVOS, these models suffer a large performance drop, highlighting the challenge of achieving precise tracking and segmentation in real-world scenarios. Attribute-based analysis indicates that key factor to accuracy decline is the increased video length, emphasizing LVOS's crucial role. We hope our LVOS can advance development of VOS in real scenes. Data and code are available at https://lingyihongfd.github.io/lvos.github.io/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SAM 3: Segment Anything with Concepts

    cs.CV 2025-11 unverdicted novelty 7.0 of 10

    SAM 3 introduces promptable concept segmentation that doubles accuracy of prior systems on images and videos while improving standard SAM segmentation performance.

  2. SAM 2++: Tracking Anything at Any Granularity

    cs.CV 2025-10 conditional novelty 7.0 of 10

    SAM 2++ unifies video tracking across mask, box, and point granularities via task-specific prompts, a unified decoder, task-adaptive memory, and a new multi-granularity dataset, reporting state-of-the-art results.

  3. MOVE: Motion-Guided Few-Shot Video Object Segmentation

    cs.CV 2025-07 conditional novelty 7.0 of 10

    MOVE provides a new motion-guided few-shot video object segmentation benchmark, and the proposed DMA baseline outperforms six existing methods across all settings.

  4. SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning

    cs.CV 2025-06 conditional novelty 7.0 of 10

    SIV-Bench is a new video benchmark with 2,792 clips and 5,455 QA pairs that evaluates MLLMs on social scene understanding, state reasoning, and dynamics prediction using social relation theory.

  5. Object-centric Video Question Answering with Visual Grounding and Referring

    cs.CV 2025-07 conditional novelty 6.0 of 10

    RGA3 unifies visual referring (arbitrary prompts at any timestamp) and grounding (segmentation masks) for object-centric video QA, introducing the STOM prompt-propagation module and the VideoInfer dataset.

  6. SAM 2: Segment Anything in Images and Videos

    cs.CV 2024-08 conditional novelty 6.0 of 10

    SAM 2 delivers more accurate video segmentation with 3x fewer user interactions and 6x faster image segmentation than the original SAM by training a streaming-memory transformer on the largest video segmentation datas...

  7. SAMannot: A Memory-Efficient, Local, Open-source Framework for Interactive Video Instance Segmentation based on SAM2

    cs.CV 2026-01 conditional novelty 5.0 of 10

    SAMannot delivers a memory-efficient local framework for interactive video instance segmentation by optimizing SAM2 with persistent identity tracking, lock-and-refine workflows, and auto-prompting.

Pith tools