REVIEW 20 cited by
Label-Efficient Semantic Segmentation with Diffusion Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Denoising diffusion probabilistic models have recently received much research attention since they outperform alternative approaches, such as GANs, and currently provide state-of-the-art generative performance. The superior performance of diffusion models has made them an appealing tool in several applications, including inpainting, super-resolution, and semantic editing. In this paper, we demonstrate that diffusion models can also serve as an instrument for semantic segmentation, especially in the setup when labeled data is scarce. In particular, for several pretrained diffusion models, we investigate the intermediate activations from the networks that perform the Markov step of the reverse diffusion process. We show that these activations effectively capture the semantic information from an input image and appear to be excellent pixel-level representations for the segmentation problem. Based on these observations, we describe a simple segmentation method, which can work even if only a few training images are provided. Our approach significantly outperforms the existing alternatives on several datasets for the same amount of human supervision.
Forward citations
Cited by 20 Pith papers
-
Functionalization via Structure Completion and Motion Rectification
Object functionalization is cast as neural graph completion over a functional graph of parts, contacts, and motions, followed by geometry realization that also rectifies erroneous motions, demonstrated on furniture wi...
-
JEDI: Joint Embedding Diffusion World Model for Online Model-Based Reinforcement Learning
JEDI is the first online end-to-end latent diffusion world model that trains latents from denoising loss rather than reconstruction, achieving competitive Atari100k results with 43% less VRAM and over 3x faster sampli...
-
From Diffusion to Rectified Flow: Rethinking Text-Based Segmentation
RLFSeg repurposes pretrained generative models via Rectified Flow for direct latent-space image-to-mask mapping in text-based segmentation, outperforming diffusion-based methods especially in zero-shot cases.
-
SD-FSMIS: Adapting Stable Diffusion for Few-Shot Medical Image Segmentation
SD-FSMIS adapts Stable Diffusion for few-shot medical image segmentation via support-query interaction and visual-to-textual translation, yielding competitive performance and strong cross-domain generalization.
-
Revisiting Autoregressive Models for Generative Image Classification
Order-marginalized any-order AR models (RandAR) outperform diffusion generative classifiers on ImageNet and OOD sets and match strong SSL models at far lower cost.
-
What Makes Synthetic Data Effective in Image Segmentation
Dense scene composition and instance fidelity in synthetic diffusion images drive better segmentation performance; SENSE framework exploits this to improve models on Cityscapes, COCO, and ADE20K.
-
Diffusion Model as a Generalist Segmentation Learner
DiGSeg repurposes diffusion U-Nets as generalist segmentation learners by conditioning on image-mask latents and multi-scale CLIP text features, achieving strong cross-domain performance.
-
Q-Sched: Pushing the Boundaries of Few-Step Diffusion Models with Quantization-Aware Scheduling
Q-Sched's quantization-aware scheduler with a reference-free JAQ loss lets 2-8 step quantized diffusion models reach lower FID than full-precision baselines.
-
Uni-DocDiff: A Unified Document Restoration Model Based on Diffusion
A dual-stream diffusion model with a handcrafted prior pool and a prior fusion module unifies six document restoration tasks and matches task-specific specialists.
-
Contrastive Conditional-Unconditional Alignment for Long-tailed Diffusion Model
A two-part training regularizer, an unconditional-only contrastive repulsion plus a large-timestep conditional-unconditional alignment, improves tail-class diversity and fidelity in diffusion models, cutting ImageNet-...
-
Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback
InnerControl trains lightweight probes on intermediate UNet features to enforce control alignment throughout the denoising trajectory, improving controllability for edges and depth.
-
Contrastive Learning with Diffusion Features for Weakly Supervised Medical Image Segmentation
Contrastive learning on diffusion-model features with foreground pixels selected by fusing class activation maps and diffusion gradient maps yields state-of-the-art weakly supervised medical image segmentation.
-
CA-Diff: Collaborative Anatomy Diffusion for Brain Tissue Segmentation
CA-Diff jointly denoises a brain label and an atlas-derived distance field with one diffusion U-Net, plus a consistency loss and time-adapted attention, and reports state-of-the-art Dice on MALC, SchizBull, and Hammers.
-
gen2seg: Generative Models Enable Generalizable Instance Segmentation
Finetuning generative models on limited instance segmentation data produces zero-shot generalization to unseen object categories and styles, matching or exceeding supervised baselines like SAM on ambiguous boundaries.
-
DiFaReli++: Diffusion Face Relighting with Consistent Cast Shadows
DiFaReli++ conditions a DDIM on shading references and inferred shadow maps to relight single-view faces with consistent shadows, trained only on 2D images and claiming SOTA on Multi-PIE.
-
Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation
Massive activations in DiTs are timestep-driven detail channels; suppressing them guides finer sampling and AdaLN-modulating them yields more discriminative dense features.
-
Mutual Enhancement Between Global Tokens and Patch Tokens: From Theory to Practice
TaTok is a theoretically grounded adaptive tokenization method that uses global tokens and cumulative conditional entropy filtering to reduce redundancy while improving reconstruction quality over fixed-rate patch tok...
-
CPKD: Clinical Prior Knowledge-Constrained Diffusion Models for Surgical Phase Recognition in Endoscopic Submucosal Dissection
A diffusion-based generative model with training-time masking and clinical logic constraints achieves state-of-the-art surgical phase recognition on ESD videos and a small gain on cholecystectomy videos.
-
Perturbation-Aware Diffusion-Guided Hybrid Segmentation for Robust and Annotation-Efficient Plant Stress Phenotyping
Matched backbone–diffusion refiners plus perturbation-driven augmentation reach 71.83% mIoU on PlantSegV3 and stay stable with 10% labels and limited target adaptation.
-
Latent Space Synergy: Text-Guided Data Augmentation for Direct Diffusion Biomedical Segmentation
Text-guided synthetic polyp images plus a single-step latent diffusion segmentation model reach 96.0 Dice on CVC-ClinicDB, but key data-generation and evaluation details are unverified.
Discussion (0). Sign in to comment.