Pith. sign in

REVIEW 1 cited by

RefVOS: A Closer Look at Referring Expressions for Video Object Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.00263 v1 pith:SMN7LSC2 submitted 2020-10-01 cs.CV

classification cs.CV
keywords tasklanguage-guidedobjectresultssegmentationvideoexpressionsnon-trivial
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The task of video object segmentation with referring expressions (language-guided VOS) is to, given a linguistic phrase and a video, generate binary masks for the object to which the phrase refers. Our work argues that existing benchmarks used for this task are mainly composed of trivial cases, in which referents can be identified with simple phrases. Our analysis relies on a new categorization of the phrases in the DAVIS-2017 and Actor-Action datasets into trivial and non-trivial REs, with the non-trivial REs annotated with seven RE semantic categories. We leverage this data to analyze the results of RefVOS, a novel neural network that obtains competitive results for the task of language-guided image segmentation and state of the art results for language-guided VOS. Our study indicates that the major challenges for the task are related to understanding motion and static actions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 182 citations worldwide. Full citation record

  1. MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation

    cs.CV 2025-01 conditional novelty 4.0 of 10

    A SAM 2-based model with BEiT-3-derived mask priors and global-historical aggregation achieves state-of-the-art referring video object segmentation on Ref-YouTube-VOS, MeViS, and Ref-DAVIS17.

Pith tools