VASA is a vision-guided agent for open ad-hoc segmentation that creates and validates masks through planning, tool use, and error recovery, outperforming baselines on the new PARS benchmark and RefCOCOm.
Ref-diff: Zero-shot referring image segmentation with generative models
6 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CV 6representative citing papers
SAM 3 introduces promptable concept segmentation that doubles accuracy of prior systems on images and videos while improving standard SAM segmentation performance.
Prompt2Seg augments diffusion models with an explicit spatial prompt conditioning branch, enabling zero-shot instance segmentation that generalizes from limited synthetic category training to diverse unseen objects and visual domains.
Pretrained instruction-based image editing models exhibit early foreground-background separability that enables a training-free framework for zero-shot referring image segmentation using a single denoising step.
A reinforced self-evolving framework (L2L) for semi-supervised referring expression segmentation that jointly optimizes the segmentation model and pseudo-labels using multimodal priors and adaptive selection.
citing papers explorer
-
Vision Harnessing Agent for Open Ad-hoc Segmentation
VASA is a vision-guided agent for open ad-hoc segmentation that creates and validates masks through planning, tool use, and error recovery, outperforming baselines on the new PARS benchmark and RefCOCOm.
-
SAM 3: Segment Anything with Concepts
SAM 3 introduces promptable concept segmentation that doubles accuracy of prior systems on images and videos while improving standard SAM segmentation performance.
-
Prompting Diffusion Models for Zero-Shot Instance Segmentation
Prompt2Seg augments diffusion models with an explicit spatial prompt conditioning branch, enabling zero-shot instance segmentation that generalizes from limited synthetic category training to diverse unseen objects and visual domains.
-
Early Semantic Grounding in Image Editing Models for Zero-Shot Referring Image Segmentation
Pretrained instruction-based image editing models exhibit early foreground-background separability that enables a training-free framework for zero-shot referring image segmentation using a single denoising step.
-
Learning to Label: A Reinforced Self-Evolving Framework for Semi-supervised Referring Expression Segmentation
A reinforced self-evolving framework (L2L) for semi-supervised referring expression segmentation that jointly optimizes the segmentation model and pseudo-labels using multimodal priors and adaptive selection.
- Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation