REVIEW 3 cited by
DFormer: Diffusion-guided Transformer for Universal Image Segmentation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper introduces an approach, named DFormer, for universal image segmentation. The proposed DFormer views universal image segmentation task as a denoising process using a diffusion model. DFormer first adds various levels of Gaussian noise to ground-truth masks, and then learns a model to predict denoising masks from corrupted masks. Specifically, we take deep pixel-level features along with the noisy masks as inputs to generate mask features and attention masks, employing diffusion-based decoder to perform mask prediction gradually. At inference, our DFormer directly predicts the masks and corresponding categories from a set of randomly-generated masks. Extensive experiments reveal the merits of our proposed contributions on different image segmentation tasks: panoptic segmentation, instance segmentation, and semantic segmentation. Our DFormer outperforms the recent diffusion-based panoptic segmentation method Pix2Seq-D with a gain of 3.6% on MS COCO val2017 set. Further, DFormer achieves promising semantic segmentation performance outperforming the recent diffusion-based method by 2.2% on ADE20K val set. Our source code and models will be publicly on https://github.com/cp3wan/DFormer
Forward citations
Cited by 3 Pith papers
-
D3RM: A Discrete Denoising Diffusion Refinement Model for Piano Transcription
D3RM uses discrete diffusion with neighborhood attention and an asymmetric train/inference masking schedule to improve piano transcription F1 over a feed-forward baseline and DiffRoll on MAESTRO.
-
D-Cube: Exploiting Hyper-Features of Diffusion Model for Robust Medical Classification
D-Cube combines selected diffusion-model feature maps with ResNet sub-features and custom losses to improve medical image classification.
-
Unleashing Diffusion and State Space Models for Medical Image Segmentation
DSM integrates k-means attention, Mamba state-space layers, diffusion-guided boundary refinement, and CLIP text prompts to segment seen organs and unseen tumors in CT images.
Discussion (0). Continue with ORCID to comment.