Pith. sign in

REVIEW 3 cited by

DFormer: Diffusion-guided Transformer for Universal Image Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.03437 v2 pith:57L6ZMRF submitted 2023-06-06 cs.CV

classification cs.CV
keywords segmentationdformermasksimagediffusion-baseduniversaldenoisingfeatures
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces an approach, named DFormer, for universal image segmentation. The proposed DFormer views universal image segmentation task as a denoising process using a diffusion model. DFormer first adds various levels of Gaussian noise to ground-truth masks, and then learns a model to predict denoising masks from corrupted masks. Specifically, we take deep pixel-level features along with the noisy masks as inputs to generate mask features and attention masks, employing diffusion-based decoder to perform mask prediction gradually. At inference, our DFormer directly predicts the masks and corresponding categories from a set of randomly-generated masks. Extensive experiments reveal the merits of our proposed contributions on different image segmentation tasks: panoptic segmentation, instance segmentation, and semantic segmentation. Our DFormer outperforms the recent diffusion-based panoptic segmentation method Pix2Seq-D with a gain of 3.6% on MS COCO val2017 set. Further, DFormer achieves promising semantic segmentation performance outperforming the recent diffusion-based method by 2.2% on ADE20K val set. Our source code and models will be publicly on https://github.com/cp3wan/DFormer

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. D3RM: A Discrete Denoising Diffusion Refinement Model for Piano Transcription

    cs.SD 2025-01 conditional novelty 6.0 of 10

    D3RM uses discrete diffusion with neighborhood attention and an asymmetric train/inference masking schedule to improve piano transcription F1 over a feed-forward baseline and DiffRoll on MAESTRO.

  2. D-Cube: Exploiting Hyper-Features of Diffusion Model for Robust Medical Classification

    cs.CV 2024-11 conditional novelty 5.0 of 10

    D-Cube combines selected diffusion-model feature maps with ResNet sub-features and custom losses to improve medical image classification.

  3. Unleashing Diffusion and State Space Models for Medical Image Segmentation

    cs.CV 2025-06 conditional novelty 4.0 of 10

    DSM integrates k-means attention, Mamba state-space layers, diffusion-guided boundary refinement, and CLIP text prompts to segment seen organs and unseen tumors in CT images.

Pith tools