Rea2Seg turns image segmentation into candidate mask discovery from MLLM attention followed by MLLM-based comparative scoring and selection, plus a new multi-dimensional reasoning benchmark ReasonSeg-SGDR.
Lens: Learning to segment anything with unified reinforced reasoning
5 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CV 5years
2026 5representative citing papers
VASA is a vision-guided agent for open ad-hoc segmentation that creates and validates masks through planning, tool use, and error recovery, outperforming baselines on the new PARS benchmark and RefCOCOm.
Topo-R1 fine-tunes a vision-language model using a topology-aware reward and GRPO to detect anomalies such as broken or spurious connections in tubular segmentation masks, outperforming standard VLMs.
PinPoint is a deterministic point selector that replaces naive sampling with stable interior points, raising cIoU by 12-18 points on RefCOCO/+/g and matching supervised/RL methods without any training.
citing papers explorer
-
Reason Twice: Segmentation via Candidate Discovery and Comparative Reasoning
Rea2Seg turns image segmentation into candidate mask discovery from MLLM attention followed by MLLM-based comparative scoring and selection, plus a new multi-dimensional reasoning benchmark ReasonSeg-SGDR.
-
Vision Harnessing Agent for Open Ad-hoc Segmentation
VASA is a vision-guided agent for open ad-hoc segmentation that creates and validates masks through planning, tool use, and error recovery, outperforming baselines on the new PARS benchmark and RefCOCOm.
-
Topo-R1: Detecting Topological Anomalies via Vision-Language Models
Topo-R1 fine-tunes a vision-language model using a topology-aware reward and GRPO to detect anomalies such as broken or spurious connections in tubular segmentation masks, outperforming standard VLMs.
-
PinPoint: Prompting with Informative Interior Points
PinPoint is a deterministic point selector that replaces naive sampling with stable interior points, raising cIoU by 12-18 points on RefCOCO/+/g and matching supervised/RL methods without any training.
- Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation