Pith. sign in

REVIEW 2 cited by

Towards Robust Referring Video Object Segmentation with Cyclic Relational Consensus

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.01203 v3 pith:65O6ZI6D submitted 2022-07-04 cs.CV

classification cs.CV
keywords r-vosobjectvideoconsensusconstraintmodelpairsrobust
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Referring Video Object Segmentation (R-VOS) is a challenging task that aims to segment an object in a video based on a linguistic expression. Most existing R-VOS methods have a critical assumption: the object referred to must appear in the video. This assumption, which we refer to as semantic consensus, is often violated in real-world scenarios, where the expression may be queried against false videos. In this work, we highlight the need for a robust R-VOS model that can handle semantic mismatches. Accordingly, we propose an extended task called Robust R-VOS, which accepts unpaired video-text inputs. We tackle this problem by jointly modeling the primary R-VOS problem and its dual (text reconstruction). A structural text-to-text cycle constraint is introduced to discriminate semantic consensus between video-text pairs and impose it in positive pairs, thereby achieving multi-modal alignment from both positive and negative pairs. Our structural constraint effectively addresses the challenge posed by linguistic diversity, overcoming the limitations of previous methods that relied on the point-wise constraint. A new evaluation dataset, R\textsuperscript{2}-Youtube-VOSis constructed to measure the model robustness. Our model achieves state-of-the-art performance on R-VOS benchmarks, Ref-DAVIS17 and Ref-Youtube-VOS, and also our R\textsuperscript{2}-Youtube-VOS~dataset.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EventRR: Event Referential Reasoning for Referring Video Object Segmentation

    cs.CV 2025-08 conditional novelty 7.0 of 10

    EventRR builds a Referential Event Graph from AMR parsing of the referring expression and uses graph-guided temporal reasoning over detector queries to select and segment the referent, reporting state-of-the-art resul...

  2. Object-centric Video Question Answering with Visual Grounding and Referring

    cs.CV 2025-07 conditional novelty 6.0 of 10

    RGA3 unifies visual referring (arbitrary prompts at any timestamp) and grounding (segmentation masks) for object-centric video QA, introducing the STOM prompt-propagation module and the VideoInfer dataset.

Pith tools