REVIEW 18 cited by
SegDiff: Image Segmentation with Diffusion Probabilistic Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
SegDiff: Image Segmentation with Diffusion Probabilistic Models
read the original abstract
Diffusion Probabilistic Methods are employed for state-of-the-art image generation. In this work, we present a method for extending such models for performing image segmentation. The method learns end-to-end, without relying on a pre-trained backbone. The information in the input image and in the current estimation of the segmentation map is merged by summing the output of two encoders. Additional encoding layers and a decoder are then used to iteratively refine the segmentation map, using a diffusion model. Since the diffusion model is probabilistic, it is applied multiple times, and the results are merged into a final segmentation map. The new method produces state-of-the-art results on the Cityscapes validation set, the Vaihingen building segmentation benchmark, and the MoNuSeg dataset.
Forward citations
Cited by 18 Pith papers
-
Generating HDR Video from SDR Video
A multi-exposure video model predicts bracketed linear SDR sequences from single nonlinear SDR input, which a merging model combines into HDR video preserving shadow and highlight detail.
-
Weakly Supervised Segmentation as Semantic-Based Regularization
Differentiable fuzzy logic constraints fine-tune SAM to generate higher-quality pseudo-labels, enabling a second-stage model to reach state-of-the-art weakly supervised segmentation on Pascal VOC and REFUGE2, sometime...
-
RF-HiT: Rectified Flow Hierarchical Transformer for General Medical Image Segmentation
TransSplat formulates language-driven 3D Gaussian Splatting editing as a multi-view unbalanced semantic transport problem, achieving better cross-view consistency and local editing precision than prior fusion-based me...
-
SD-FSMIS: Adapting Stable Diffusion for Few-Shot Medical Image Segmentation
SD-FSMIS adapts Stable Diffusion for few-shot medical image segmentation via support-query interaction and visual-to-textual translation, yielding competitive performance and strong cross-domain generalization.
-
COP-GEN: Latent Diffusion Transformer for Copernicus Earth Observation Data
COP-GEN models multimodal Copernicus Earth observation data as conditional distributions via a latent diffusion transformer, producing diverse physically consistent outputs and covering 90% of the real observation man...
-
Weakly Supervised Segmentation as Semantic-Based Regularization
A neurosymbolic approach uses fuzzy logic constraints to refine SAM under weak supervision, producing improved pseudo-labels that enable state-of-the-art segmentation on Pascal VOC and REFUGE2.
-
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
UniVidX unifies diverse video generation tasks into one conditional diffusion model using stochastic condition masking, decoupled gated LoRAs, and cross-modal self-attention.
-
Diffusion Model as a Generalist Segmentation Learner
DiGSeg repurposes diffusion U-Nets as generalist segmentation learners by conditioning on image-mask latents and multi-scale CLIP text features, achieving strong cross-domain performance.
-
MedFlowSeg: Flow Matching for Medical Image Segmentation with Frequency-Aware Attention
MedFlowSeg is a conditional flow matching model for medical image segmentation that adds dual-branch spatial attention and frequency-aware attention to achieve more efficient inference than diffusion models while impr...
-
Enhancing Foundation VLM Robustness to Missing Modality: Scalable Diffusion for Bi-directional Feature Restoration
A diffusion model with dynamic modality gating and cross-modal mutual learning restores missing features in VLMs bi-directionally while preserving the original model's generalization.
-
Segment Any-Quality Images with Generative Latent Space Enhancement
GleSAM integrates latent diffusion into SAM and SAM2 to boost segmentation robustness on low-quality images using minimal extra parameters and a new LQSeg dataset.
-
DiFaReli++: Diffusion Face Relighting with Consistent Cast Shadows
DiFaReli++ conditions a DDIM on shading references and inferred shadow maps to relight single-view faces with consistent shadows, trained only on 2D images and claiming SOTA on Multi-PIE.
-
Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation
Massive activations in DiTs are timestep-driven detail channels; suppressing them guides finer sampling and AdaLN-modulating them yields more discriminative dense features.
-
MLFFM-SegDiff: A Multi-Level Feature Fusion Diffusion Model for Skin Lesion Segmentation
MLFFM-SegDiff adds a multi-level feature fusion module and dual-path encoder to a diffusion U-Net, reporting improved Jaccard (0.8546) and Dice (0.9207) scores over baselines on three skin lesion datasets.
-
RF-HiT: Rectified Flow Hierarchical Transformer for General Medical Image Segmentation
RF-HiT uses rectified flow and a multi-scale hierarchical transformer to reach 91.27% Dice on ACDC and 87.40% on BraTS 2021 with only 10.14 GFLOPs, 13.6M parameters, and three inference steps.
-
Towards Any-Quality Image Segmentation via Generative and Adaptive Latent Space Enhancement
GleSAM++ improves SAM robustness on degraded images by using generative enhancement, feature alignment, and adaptive degradation prediction while adding few parameters.
-
Diffusion Models for Solving Inverse Problems via Posterior Sampling with Piecewise Guidance
Piecewise guidance in diffusion posterior sampling cuts inference time 23-25% on inpainting and super-resolution with negligible PSNR/SSIM loss while handling measurement noise.
-
Volumetric Directional Diffusion: Anchoring Uncertainty Quantification in Anatomical Consensus for Ambiguous Medical Image Segmentation
Anchoring a 3D diffusion model to a deterministic consensus prior and sampling only boundary residuals improves uncertainty alignment while keeping anatomical structure intact.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.