Pith. sign in

REVIEW 1 cited by

MaskDiffusion: Exploiting Pre-trained Diffusion Models for Semantic Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.11194 v1 pith:5A6EMRNH submitted 2024-03-17 cs.CV

classification cs.CV
keywords maskdiffusionsegmentationcomparedsemanticannotationapplicationscategoriesclasses
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Semantic segmentation is essential in computer vision for various applications, yet traditional approaches face significant challenges, including the high cost of annotation and extensive training for supervised learning. Additionally, due to the limited predefined categories in supervised learning, models typically struggle with infrequent classes and are unable to predict novel classes. To address these limitations, we propose MaskDiffusion, an innovative approach that leverages pretrained frozen Stable Diffusion to achieve open-vocabulary semantic segmentation without the need for additional training or annotation, leading to improved performance compared to similar methods. We also demonstrate the superior performance of MaskDiffusion in handling open vocabularies, including fine-grained and proper noun-based categories, thus expanding the scope of segmentation applications. Overall, our MaskDiffusion shows significant qualitative and quantitative improvements in contrast to other comparable unsupervised segmentation methods, i.e. on the Potsdam dataset (+10.5 mIoU compared to GEM) and COCO-Stuff (+14.8 mIoU compared to DiffSeg). All code and data will be released at https://github.com/Valkyrja3607/MaskDiffusion.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HiGarment: Cross-modal Harmony Based Diffusion Model for Flat Sketch to Realistic Garment Image

    cs.CV 2025-05 conditional novelty 6.0 of 10

    HiGarment generates realistic garment images from a flat sketch and text prompt by retrieving fabric references and dynamically weighting sketch versus text cues, outperforming five existing controllable generation me...

Pith tools