Pith. sign in

REVIEW 1 cited by

ReferEverything: Towards Segmenting Everything We Can Speak of in Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.23287 v2 pith:CR2EK73A submitted 2024-10-30 cs.CV

classification cs.CV
keywords videosegmentationconceptsdatasetsgenerativemodelpredictingreferring
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present REM, a framework for segmenting a wide range of concepts in video that can be described through natural language. Our method leverages the universal visual-language mapping learned by video diffusion models on Internet-scale data by fine-tuning them on small-scale Referring Object Segmentation datasets. Our key insight is to preserve the entirety of the generative model's architecture by shifting its objective from predicting noise to predicting mask latents. The resulting model can accurately segment rare and unseen objects, despite only being trained on a limited set of categories. Additionally, it can effortlessly generalize to non-object dynamic concepts, such as smoke or raindrops, as demonstrated in our new benchmark for Referring Video Process Segmentation (Ref-VPS). REM performs on par with the state-of-the-art on in-domain datasets, like Ref-DAVIS, while outperforming them by up to 12 IoU points out-of-domain, leveraging the power of generative pre-training. We also show that advancements in video generation directly improve segmentation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Temporal-Conditional Referring Video Object Segmentation with Noise-Free Text-to-Video Diffusion Model

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    Removing the noise-prediction branch from a text-to-video diffusion feature extractor, plus a temporal context mask refinement module, yields claimed state-of-the-art referring video object segmentation on four benchmarks.

Pith tools