DiGSeg repurposes diffusion U-Nets as generalist segmentation learners by conditioning on image-mask latents and multi-scale CLIP text features, achieving strong cross-domain performance.
Image segmentation in foundation model era: A survey.arXiv preprint arXiv:2408.12957
5 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CV 5years
2026 5verdicts
UNVERDICTED 5representative citing papers
RS4D distills ViT knowledge into SSM backbones for remote sensing instance segmentation, delivering 8x fewer parameters and 9x fewer FLOPs than ViT methods while matching or exceeding accuracy on SSDD, WHU, and NWPU datasets.
SegDINO adds Token Pyramid Adaptation and Scale-Aware Decoding to DINOv3 to deliver efficient state-of-the-art medical image segmentation on a new pancreatic CT dataset and public benchmarks.
A visualization protocol using unsupervised semantic segmentation outputs reveals positional biases, scaling behaviors, and boundary artifacts in self-supervised ViTs and distinguishes them from locality bias.
MAgSeg is a decoder-free MLLM approach that uses a new instruction-tuning format to segment complex smallholder agricultural landscapes directly from high-resolution satellite imagery.
citing papers explorer
-
Diffusion Model as a Generalist Segmentation Learner
DiGSeg repurposes diffusion U-Nets as generalist segmentation learners by conditioning on image-mask latents and multi-scale CLIP text features, achieving strong cross-domain performance.
-
Efficient Remote Sensing Instance Segmentation with Linear-Time State Space Distilled Visual Foundation Models
RS4D distills ViT knowledge into SSM backbones for remote sensing instance segmentation, delivering 8x fewer parameters and 9x fewer FLOPs than ViT methods while matching or exceeding accuracy on SSDD, WHU, and NWPU datasets.
-
SegDINO: Introducing Multi-Scale Structure into DINO for Efficient Medical Image Segmentation
SegDINO adds Token Pyramid Adaptation and Scale-Aware Decoding to DINOv3 to deliver efficient state-of-the-art medical image segmentation on a new pancreatic CT dataset and public benchmarks.
-
Unsupervised Semantic Segmentation Facilitates Model Understanding
A visualization protocol using unsupervised semantic segmentation outputs reveals positional biases, scaling behaviors, and boundary artifacts in self-supervised ViTs and distinguishes them from locality bias.
-
MAgSeg: Segmentation of Agricultural Landscapes in High-Resolution Satellite Imagery using Multimodal Large Language Models
MAgSeg is a decoder-free MLLM approach that uses a new instruction-tuning format to segment complex smallholder agricultural landscapes directly from high-resolution satellite imagery.