Pith. sign in

REVIEW 3 cited by

DiffuMask: Synthesizing Images with Pixel-level Annotations for Semantic Segmentation Using Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.11681 v4 pith:2M3G4NT3 submitted 2023-03-21 cs.CV

classification cs.CV
keywords diffumaskdatadiffusionimagessegmentationsemanticsyntheticclass
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Collecting and annotating images with pixel-wise labels is time-consuming and laborious. In contrast, synthetic data can be freely available using a generative model (e.g., DALL-E, Stable Diffusion). In this paper, we show that it is possible to automatically obtain accurate semantic masks of synthetic images generated by the Off-the-shelf Stable Diffusion model, which uses only text-image pairs during training. Our approach, called DiffuMask, exploits the potential of the cross-attention map between text and image, which is natural and seamless to extend the text-driven image synthesis to semantic mask generation. DiffuMask uses text-guided cross-attention information to localize class/word-specific regions, which are combined with practical techniques to create a novel high-resolution and class-discriminative pixel-wise mask. The methods help to reduce data collection and annotation costs obviously. Experiments demonstrate that the existing segmentation methods trained on synthetic data of DiffuMask can achieve a competitive performance over the counterpart of real data (VOC 2012, Cityscapes). For some classes (e.g., bird), DiffuMask presents promising performance, close to the stateof-the-art result of real data (within 3% mIoU gap). Moreover, in the open-vocabulary segmentation (zero-shot) setting, DiffuMask achieves a new SOTA result on Unseen class of VOC 2012. The project website can be found at https://weijiawu.github.io/DiffusionMask/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. S3OD: Towards Generalizable Salient Object Detection with Synthetic Data

    cs.CV 2025-10 conditional novelty 7.0 of 10

    A 139k-image synthetic dataset with diffusion- and DINO-derived masks, trained with a multi-mask decoder, improves cross-dataset salient-object detection and reaches state-of-the-art after fine-tuning.

  2. Direct Segmentation without Logits Optimization for Training-Free Open-Vocabulary Semantic Segmentation

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    The approach uses the analytic solution of distribution discrepancy consistency within categories as semantic maps, eliminating training and model-specific modulation while claiming state-of-the-art results on eight b...

  3. Novel Category Discovery with X-Agent Attention for Open-Vocabulary Semantic Segmentation

    cs.CV 2025-09 conditional novelty 5.0 of 10

    X-Agent adds agent tokens, chosen by optimal-transport affinity between text and visual keys, to CLIP attention, reporting marginal mIoU gains (0.1-0.6%) over prior OVSS methods.

Pith tools