SetCon achieves state-of-the-art open-ended referring segmentation by using LVLM-generated set-level concepts for joint mask decoding, with gains increasing for multi-target cases on image and video benchmarks.
Image segmentation using text and image prompts
2 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CV 2years
2026 2representative citing papers
ReferEndoscopy plus attribute-retrieval and frequency-aware fusion yields open-vocabulary compositional referring segmentation that outperforms natural-image RIS baselines on endoscopic data and generalizes to an unseen robotic prostatectomy set.
citing papers explorer
-
SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction
SetCon achieves state-of-the-art open-ended referring segmentation by using LVLM-generated set-level concepts for joint mask decoding, with gains increasing for multi-target cases on image and video benchmarks.
-
Attribute Retrieving for Open-Vocabulary Endoscopic Compositional Referring Segmentation
ReferEndoscopy plus attribute-retrieval and frequency-aware fusion yields open-vocabulary compositional referring segmentation that outperforms natural-image RIS baselines on endoscopic data and generalizes to an unseen robotic prostatectomy set.