REVIEW 36 cited by
SegDiff: Image Segmentation with Diffusion Probabilistic Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Diffusion Probabilistic Methods are employed for state-of-the-art image generation. In this work, we present a method for extending such models for performing image segmentation. The method learns end-to-end, without relying on a pre-trained backbone. The information in the input image and in the current estimation of the segmentation map is merged by summing the output of two encoders. Additional encoding layers and a decoder are then used to iteratively refine the segmentation map, using a diffusion model. Since the diffusion model is probabilistic, it is applied multiple times, and the results are merged into a final segmentation map. The new method produces state-of-the-art results on the Cityscapes validation set, the Vaihingen building segmentation benchmark, and the MoNuSeg dataset.
Forward citations
Cited by 36 Pith papers
-
Multi-Channel Uncertainty-Weighted Score Matching for Conditional Diffusion in Medical UDA
A UDA method uses Bezier style transfer plus an uncertainty-weighted conditional diffusion model to generate labeled target-style images, improving cross-modality medical segmentation.
-
ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features
ConceptAttention shows that linear projections in the output space of DiT attention layers yield sharper concept-localizing saliency maps than cross-attention maps, reaching state-of-the-art zero-shot segmentation.
-
Taking a Big Step: Large Learning Rates in Denoising Score Matching Prevent Memorization
In one-dimensional denoising score matching with two-layer ReLU networks, a large SGD learning rate provably prevents the learned score from getting close to the empirical optimal score, mitigating memorization.
-
DUDE: Diffusion-Based Unsupervised Cross-Domain Image Retrieval
A diffusion-based disentanglement method that separates object content from domain style achieves state-of-the-art unsupervised cross-domain image retrieval on three benchmarks.
-
Q-Sched: Pushing the Boundaries of Few-Step Diffusion Models with Quantization-Aware Scheduling
Q-Sched's quantization-aware scheduler with a reference-free JAQ loss lets 2-8 step quantized diffusion models reach lower FID than full-precision baselines.
-
Generic Event Boundary Detection via Denoising Diffusion
A conditional diffusion model, DiffGEBD, generates diverse but plausible event boundary predictions for videos, with a new symmetric F1 and diversity score protocol for evaluating multi-prediction quality.
-
SDMatte: Grafting Diffusion Models for Interactive Matting
SDMatte adapts Stable Diffusion to interactive matting via visual-prompt cross-attention, opacity/coordinate embeddings, and masked self-attention, reporting SOTA results on multiple benchmarks.
-
Flow Stochastic Segmentation Networks
Flow-SSNs model high-rank pixel covariances for ambiguous medical image segmentation by mapping a learned diagonal-Gaussian prior through a lightweight flow, outperforming prior SOTA with fewer parameters.
-
Diffusion-FS: Multimodal Free-Space Prediction via Diffusion for Autonomous Driving
A self-supervised diffusion model predicts multimodal drivable corridors as contour points in monocular images, and beats two segmentation baselines on CARLA and nuScenes.
-
UniSegDiff: Boosting Unified Lesion Segmentation via a Staged Diffusion Model
UniSegDiff uses staged training and inference with alternating mask/noise prediction targets plus STAPLE fusion of multiple samples to reach state-of-the-art lesion segmentation across six datasets and modalities.
-
FreeDNA: Endowing Domain Adaptation of Diffusion-Based Dense Prediction with Training-Free Domain Noise Alignment
Aligning the L2-norm statistics of noise predictions during diffusion sampling improves domain adaptation for dense prediction, with a source-free version guided by high-confidence regions.
-
Compositional Scene Understanding through Inverse Generative Modeling
Composing per-concept diffusion models and inverting them with denoising loss enables multi-object scene understanding that generalizes beyond the training distribution.
-
DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion Transformers
DiffDecompose recovers foreground and background layers from alpha-composited images using in-context diffusion with position encoding cloning, trained and evaluated on a new six-task synthetic dataset.
-
On Denoising Walking Videos for Gait Recognition
DenoisingGait combines frozen Stable Diffusion features with learned direction-vector matching to create Gait Feature Fields, reporting new state-of-the-art rank-1 accuracy on CCPG and most settings of CASIA-B*, SUSTech1K.
-
DIVE: Inverting Conditional Diffusion Models for Discriminative Tasks
A frozen conditional diffusion model can be inverted via gradient-based discrete optimization, plus a learned layout prior, to perform object detection and faster classification without training a discriminative head.
-
Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models
DeGF reduces hallucinations in vision-language models by generating an image from the model's own response and using the divergence between predictions on original and generated images to switch between complementary ...
-
PhysMotion: Physics-Grounded Dynamics From a Single Image
PhysMotion generates physically plausible videos from a single image by simulating 3D object motion with a material point method, then enhancing the rendering with a diffusion model.
-
Enhancing Diffusion Posterior Sampling for Inverse Problems by Integrating Crafted Measurements
DPS-CM improves diffusion posterior sampling for inverse problems by generating a denoised reverse-measurement trajectory and using it in the likelihood gradient, yielding better restoration in experiments.
-
Scaling Properties of Diffusion Models for Perceptual Tasks
Diffusion models for depth, optical flow, and amodal segmentation improve along power laws as training and test-time compute scale, and the fitted recipes match prior specialist models with less data.
-
Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation
Massive activations in DiTs are timestep-driven detail channels; suppressing them guides finer sampling and AdaLN-modulating them yields more discriminative dense features.
-
IDCNet: Guided Video Diffusion for Metric-Consistent RGBD Scene Generation with Precise Camera Control
The claimed IDC-Net framework is absent; the body text is an unrelated instance-segmentation paper.
-
From Variability To Accuracy: Conditional Bernoulli Diffusion Models with Consensus-Driven Correction for Thin Structure Segmentation
A consensus-driven correction on top of a conditional Bernoulli diffusion model improves recall in thin orbital bone segmentation, though not all metrics beat baselines.
-
LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion
LangScene-X generates RGB, normal, and semantic videos from sparse views to reconstruct 3D language-embedded Gaussian fields that support open-ended text queries.
-
FedSaaS: Class-Consistency Federated Semantic Segmentation via Global Prototype Supervision and Local Adversarial Harmonization
FedSaaS aligns class representations across federated clients via class exemplars, global prototype supervision, and local adversarial harmonization, improving segmentation accuracy under domain shift.
-
LDPoly: Latent Diffusion for Polygonal Road Outline Extraction in Large-Scale Topographic Mapping
LDPoly jointly generates road masks and vertex heatmaps with a dual-latent diffusion model, then polygonizes them into compact road outlines that beat prior methods on Dutch topographic benchmark Map2ImLas.
-
Exploring the latent space of diffusion models directly through singular value decomposition
The authors report that singular value decomposition of diffusion latent codes reveals stable, order-mobile attribute directions and propose Attribute Vector Integration, a per-pair MLP-based editor that transfers tex...
-
CrossDiff: Diffusion Probabilistic Model With Cross-conditional Encoder-Decoder for Crack Segmentation
CrossDiff applies a diffusion probabilistic model with a cross-conditional encoder-decoder to crack segmentation, reporting state-of-the-art Dice and IoU scores on five datasets.
-
CityLoc: 6DoF Pose Distributional Localization for Text Descriptions in Large-Scale Scenes with Gaussian Representation
A text-conditioned diffusion model with 3D Gaussian splatting refinement estimates 6DoF camera pose distributions in city-scale scenes, beating a Monte Carlo dropout baseline on five datasets.
-
Video-Guided Foley Sound Generation with Multimodal Controls
A video-guided diffusion model generates synchronized foley sound from text, audio, and video controls, using joint training on noisy internet videos and professional sound-effect libraries to reach 48kHz output.
-
Volumetric Directional Diffusion: Anchoring Uncertainty Quantification in Anatomical Consensus for Ambiguous Medical Image Segmentation
Anchoring a 3D diffusion model to a deterministic consensus prior and sampling only boundary residuals improves uncertainty alignment while keeping anatomical structure intact.
-
SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3
A frozen DINOv3 backbone plus a simple MLP head reportedly beats specialized segmentation models on six benchmarks, but the evidence lacks statistical rigor.
-
Probabilistic Spatial Interpolation of Sparse Data using Diffusion Models
KrigSCD, a kriging-smoothed diffusion inpainting method, reconstructs 2D temperature fields from sparse masks and beats IDW, kriging, and plain diffusion on LPIPS at all tested coverage levels.
-
Conditional diffusion model with spatial attention and latent embedding for medical image segmentation
A conditional diffusion model with a per-timestep discriminator, spatial attention, and latent embedding reports state-of-the-art accuracy on three medical segmentation datasets using only 2 to 4 diffusion steps.
-
Adversarially Domain-adaptive Latent Diffusion for Unsupervised Semantic Segmentation
An adversarially trained latent diffusion model with long encoder-decoder skip connections reports state-of-the-art mIoU of 74.4 and 67.2 on two unsupervised domain adaptation benchmarks.
-
Computationally Efficient Diffusion Models in Medical Imaging: A Comprehensive Review
A survey of DDPM, LDM, and WDM diffusion models for medical imaging, organized around training and inference efficiency.
-
Diffusion Models for Hyperspectral Image Analysis: A Comprehensive Review
A literature review that organizes diffusion-model work for hyperspectral imaging into eight task categories and compiles comparative performance tables from prior papers.
Discussion (0). Continue with ORCID to comment.