Pith. sign in

REVIEW 2 cited by

Explore In-Context Segmentation via Latent Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.09616 v2 pith:U6RI7YUJ submitted 2024-03-14 cs.CV

classification cs.CV
keywords segmentationin-contextimagemodelsapproachesbuilddesigndiffusion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In-context segmentation has drawn increasing attention with the advent of vision foundation models. Its goal is to segment objects using given reference images. Most existing approaches adopt metric learning or masked image modeling to build the correlation between visual prompts and input image queries. This work approaches the problem from a fresh perspective - unlocking the capability of the latent diffusion model (LDM) for in-context segmentation and investigating different design choices. Specifically, we examine the problem from three angles: instruction extraction, output alignment, and meta-architectures. We design a two-stage masking strategy to prevent interfering information from leaking into the instructions. In addition, we propose an augmented pseudo-masking target to ensure the model predicts without forgetting the original images. Moreover, we build a new and fair in-context segmentation benchmark that covers both image and video datasets. Experiments validate the effectiveness of our approach, demonstrating comparable or even stronger results than previous specialist or visual foundation models. We hope our work inspires others to rethink the unification of segmentation and generation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unleashing In-context Learning of Autoregressive Models for Few-shot Image Manipulation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    An autoregressive model with group self-attention that separates learning from applying achieves state-of-the-art few-shot image manipulation on unseen instructions.

  2. Vision and Language Reference Prompt into SAM for Few-shot Segmentation

    cs.CV 2025-02 conditional novelty 5.0 of 10

    Using a frozen vision-language model, VLP-SAM injects text-label semantics into SAM's prompt encoder and raises one-shot segmentation mIoU by 6.3 points on PASCAL-5i and 9.5 on COCO-20i.

Pith tools