Pith. sign in

REVIEW 2 cited by

SOC: Semantic-Assisted Object Cluster for Referring Video Object Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.17011 v1 pith:ZBUVYI2W submitted 2023-05-26 cs.CV

classification cs.CV
keywords objectvideosegmentationtemporalrvosalignmentbenchmarkscluster
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper studies referring video object segmentation (RVOS) by boosting video-level visual-linguistic alignment. Recent approaches model the RVOS task as a sequence prediction problem and perform multi-modal interaction as well as segmentation for each frame separately. However, the lack of a global view of video content leads to difficulties in effectively utilizing inter-frame relationships and understanding textual descriptions of object temporal variations. To address this issue, we propose Semantic-assisted Object Cluster (SOC), which aggregates video content and textual guidance for unified temporal modeling and cross-modal alignment. By associating a group of frame-level object embeddings with language tokens, SOC facilitates joint space learning across modalities and time steps. Moreover, we present multi-modal contrastive supervision to help construct well-aligned joint space at the video level. We conduct extensive experiments on popular RVOS benchmarks, and our method outperforms state-of-the-art competitors on all benchmarks by a remarkable margin. Besides, the emphasis on temporal coherence enhances the segmentation stability and adaptability of our method in processing text expressions with temporal variations. Code will be available.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EventRR: Event Referential Reasoning for Referring Video Object Segmentation

    cs.CV 2025-08 conditional novelty 7.0 of 10

    EventRR builds a Referential Event Graph from AMR parsing of the referring expression and uses graph-guided temporal reasoning over detector queries to select and segment the referent, reporting state-of-the-art resul...

  2. Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    OpenBench, a new benchmark with categories semantically far from the COCO training space, shows that fine-tuning CLIP hurts open-vocabulary segmentation, and the proposed OVSNet method achieves state-of-the-art on bot...

Pith tools