Pith. sign in

REVIEW 5 cited by

Image Segmentation in Foundation Model Era: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.12957 v3 pith:YC3K3W2I submitted 2024-08-23 cs.CV

classification cs.CV
keywords segmentationimageresearchfoundationsurveyapproacheschallengesclip
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Image segmentation is a long-standing challenge in computer vision, studied continuously over several decades, as evidenced by seminal algorithms such as N-Cut, FCN, and MaskFormer. With the advent of foundation models (FMs), contemporary segmentation methodologies have embarked on a new epoch by either adapting FMs (e.g., CLIP, Stable Diffusion, DINO) for image segmentation or developing dedicated segmentation foundation models (e.g., SAM). These approaches not only deliver superior segmentation performance, but also herald newfound segmentation capabilities previously unseen in deep learning context. However, current research in image segmentation lacks a detailed analysis of distinct characteristics, challenges, and solutions associated with these advancements. This survey seeks to fill this gap by providing a thorough review of cutting-edge research centered around FM-driven image segmentation. We investigate two basic lines of research -- generic image segmentation (i.e., semantic segmentation, instance segmentation, panoptic segmentation), and promptable image segmentation (i.e., interactive segmentation, referring segmentation, few-shot segmentation) -- by delineating their respective task settings, background concepts, and key challenges. Furthermore, we provide insights into the emergence of segmentation knowledge from FMs like CLIP, Stable Diffusion, and DINO. An exhaustive overview of over 300 segmentation approaches is provided to encapsulate the breadth of current research efforts. Subsequently, we engage in a discussion of open issues and potential avenues for future research. We envisage that this fresh, comprehensive, and systematic survey catalyzes the evolution of advanced image segmentation systems. A public website is created to continuously track developments in this fast advancing field: \url{https://github.com/stanley-313/ImageSegFM-Survey}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SCOPE: Speech-guided COllaborative PErception Framework for Surgical Scene Segmentation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A speech-guided framework uses an LLM and open-set vision models to segment and track surgical instruments and anatomy hands-free in live video.

  2. ConText: Driving In-context Learning for Text Removal and Segmentation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    ConText is the first visual in-context learning model for text removal and segmentation, chaining the two tasks and using self-prompting to reach new state-of-the-art scores.

  3. G4Seg: Generation for Inexact Segmentation Refinement with Diffusion Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Using the discrepancy between an image and its mask-conditioned Stable Diffusion reconstruction, G4Seg refines coarse segmentation masks by aligning pixels in CLIP feature space and mixing foreground probabilities.

  4. Segment Anyword: Mask Prompt Inversion for Open-Set Grounded Segmentation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A training-free pipeline uses per-image textual inversion in a frozen diffusion model, then feeds linguistic-guided cross-attention prompts to SAM, achieving state-of-the-art open-set grounded segmentation on several ...

  5. SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3

    cs.CV 2025-08 reject novelty 4.0 of 10

    A frozen DINOv3 backbone plus a simple MLP head reportedly beats specialized segmentation models on six benchmarks, but the evidence lacks statistical rigor.

Pith tools