Pith. sign in

REVIEW 20 cited by

Label-Efficient Semantic Segmentation with Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.03126 v3 pith:EWJCPPBK submitted 2021-12-06 cs.CV cs.LG

classification cs.CVcs.LG
keywords diffusionmodelssegmentationsemanticseveralactivationsperformancealternative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Denoising diffusion probabilistic models have recently received much research attention since they outperform alternative approaches, such as GANs, and currently provide state-of-the-art generative performance. The superior performance of diffusion models has made them an appealing tool in several applications, including inpainting, super-resolution, and semantic editing. In this paper, we demonstrate that diffusion models can also serve as an instrument for semantic segmentation, especially in the setup when labeled data is scarce. In particular, for several pretrained diffusion models, we investigate the intermediate activations from the networks that perform the Markov step of the reverse diffusion process. We show that these activations effectively capture the semantic information from an input image and appear to be excellent pixel-level representations for the segmentation problem. Based on these observations, we describe a simple segmentation method, which can work even if only a few training images are provided. Our approach significantly outperforms the existing alternatives on several datasets for the same amount of human supervision.

Discussion (0). Sign in to comment.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Functionalization via Structure Completion and Motion Rectification

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    Object functionalization is cast as neural graph completion over a functional graph of parts, contacts, and motions, followed by geometry realization that also rectifies erroneous motions, demonstrated on furniture wi...

  2. JEDI: Joint Embedding Diffusion World Model for Online Model-Based Reinforcement Learning

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    JEDI is the first online end-to-end latent diffusion world model that trains latents from denoising loss rather than reconstruction, achieving competitive Atari100k results with 43% less VRAM and over 3x faster sampli...

  3. From Diffusion to Rectified Flow: Rethinking Text-Based Segmentation

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    RLFSeg repurposes pretrained generative models via Rectified Flow for direct latent-space image-to-mask mapping in text-based segmentation, outperforming diffusion-based methods especially in zero-shot cases.

  4. SD-FSMIS: Adapting Stable Diffusion for Few-Shot Medical Image Segmentation

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    SD-FSMIS adapts Stable Diffusion for few-shot medical image segmentation via support-query interaction and visual-to-textual translation, yielding competitive performance and strong cross-domain generalization.

  5. Revisiting Autoregressive Models for Generative Image Classification

    cs.CV 2026-03 accept novelty 6.5 of 10

    Order-marginalized any-order AR models (RandAR) outperform diffusion generative classifiers on ImageNet and OOD sets and match strong SSL models at far lower cost.

  6. What Makes Synthetic Data Effective in Image Segmentation

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    Dense scene composition and instance fidelity in synthetic diffusion images drive better segmentation performance; SENSE framework exploits this to improve models on Cityscapes, COCO, and ADE20K.

  7. Diffusion Model as a Generalist Segmentation Learner

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    DiGSeg repurposes diffusion U-Nets as generalist segmentation learners by conditioning on image-mask latents and multi-scale CLIP text features, achieving strong cross-domain performance.

  8. Q-Sched: Pushing the Boundaries of Few-Step Diffusion Models with Quantization-Aware Scheduling

    cs.CV 2025-09 conditional novelty 6.0 of 10

    Q-Sched's quantization-aware scheduler with a reference-free JAQ loss lets 2-8 step quantized diffusion models reach lower FID than full-precision baselines.

  9. Uni-DocDiff: A Unified Document Restoration Model Based on Diffusion

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A dual-stream diffusion model with a handcrafted prior pool and a prior fusion module unifies six document restoration tasks and matches task-specific specialists.

  10. Contrastive Conditional-Unconditional Alignment for Long-tailed Diffusion Model

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A two-part training regularizer, an unconditional-only contrastive repulsion plus a large-timestep conditional-unconditional alignment, improves tail-class diversity and fidelity in diffusion models, cutting ImageNet-...

  11. Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback

    cs.CV 2025-07 conditional novelty 6.0 of 10

    InnerControl trains lightweight probes on intermediate UNet features to enforce control alignment throughout the denoising trajectory, improving controllability for edges and depth.

  12. Contrastive Learning with Diffusion Features for Weakly Supervised Medical Image Segmentation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Contrastive learning on diffusion-model features with foreground pixels selected by fusing class activation maps and diffusion gradient maps yields state-of-the-art weakly supervised medical image segmentation.

  13. CA-Diff: Collaborative Anatomy Diffusion for Brain Tissue Segmentation

    eess.IV 2025-06 conditional novelty 6.0 of 10

    CA-Diff jointly denoises a brain label and an atlas-derived distance field with one diffusion U-Net, plus a consistency loss and time-adapted attention, and reports state-of-the-art Dice on MALC, SchizBull, and Hammers.

  14. gen2seg: Generative Models Enable Generalizable Instance Segmentation

    cs.CV 2025-05 unverdicted novelty 6.0 of 10

    Finetuning generative models on limited instance segmentation data produces zero-shot generalization to unseen object categories and styles, matching or exceeding supervised baselines like SAM on ambiguous boundaries.

  15. DiFaReli++: Diffusion Face Relighting with Consistent Cast Shadows

    cs.CV 2023-04 unverdicted novelty 6.0 of 10

    DiFaReli++ conditions a DDIM on shading references and inferred shadow maps to relight single-view faces with consistent shadows, trained only on 2D images and claiming SOTA on Multi-PIE.

  16. Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Massive activations in DiTs are timestep-driven detail channels; suppressing them guides finer sampling and AdaLN-modulating them yields more discriminative dense features.

  17. Mutual Enhancement Between Global Tokens and Patch Tokens: From Theory to Practice

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    TaTok is a theoretically grounded adaptive tokenization method that uses global tokens and cumulative conditional entropy filtering to reduce redundancy while improving reconstruction quality over fixed-rate patch tok...

  18. CPKD: Clinical Prior Knowledge-Constrained Diffusion Models for Surgical Phase Recognition in Endoscopic Submucosal Dissection

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A diffusion-based generative model with training-time masking and clinical logic constraints achieves state-of-the-art surgical phase recognition on ESD videos and a small gain on cholecystectomy videos.

  19. Perturbation-Aware Diffusion-Guided Hybrid Segmentation for Robust and Annotation-Efficient Plant Stress Phenotyping

    cs.CV 2026-07 conditional novelty 4.0 of 10

    Matched backbone–diffusion refiners plus perturbation-driven augmentation reach 71.83% mIoU on PlantSegV3 and stay stable with 10% labels and limited target adaptation.

  20. Latent Space Synergy: Text-Guided Data Augmentation for Direct Diffusion Biomedical Segmentation

    eess.IV 2025-07 conditional novelty 3.0 of 10

    Text-guided synthetic polyp images plus a single-step latent diffusion segmentation model reach 96.0 Dice on CVC-ClinicDB, but key data-generation and evaluation details are unverified.

Pith tools