Pith. sign in

REVIEW 5 cited by

StableNormal: Reducing Diffusion Variance for Stable and Sharp Normal

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.16864 v1 pith:7SZTF6Q6 submitted 2024-06-24 cs.CV cs.AIcs.GR

classification cs.CVcs.AIcs.GR
keywords normalstablenormaldiffusionestimationprocessbeendeterministicensembling
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This work addresses the challenge of high-quality surface normal estimation from monocular colored inputs (i.e., images and videos), a field which has recently been revolutionized by repurposing diffusion priors. However, previous attempts still struggle with stochastic inference, conflicting with the deterministic nature of the Image2Normal task, and costly ensembling step, which slows down the estimation process. Our method, StableNormal, mitigates the stochasticity of the diffusion process by reducing inference variance, thus producing "Stable-and-Sharp" normal estimates without any additional ensembling process. StableNormal works robustly under challenging imaging conditions, such as extreme lighting, blurring, and low quality. It is also robust against transparent and reflective surfaces, as well as cluttered scenes with numerous objects. Specifically, StableNormal employs a coarse-to-fine strategy, which starts with a one-step normal estimator (YOSO) to derive an initial normal guess, that is relatively coarse but reliable, then followed by a semantic-guided refinement process (SG-DRN) that refines the normals to recover geometric details. The effectiveness of StableNormal is demonstrated through competitive performance in standard datasets such as DIODE-indoor, iBims, ScannetV2 and NYUv2, and also in various downstream tasks, such as surface reconstruction and normal enhancement. These results evidence that StableNormal retains both the "stability" and "sharpness" for accurate normal estimation. StableNormal represents a baby attempt to repurpose diffusion priors for deterministic estimation. To democratize this, code and models have been publicly available in hf.co/Stable-X

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. WildShadowRemover: In-the-Wild Video Shadow Removal via Detail-Preserving Video Diffusion Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    WildShadowRemover fine-tunes a pretrained video diffusion model with LoRA plus detail-injection and depth conditioning to produce temporally consistent shadow-free videos, trained on a new synthetic dataset.

  2. Detail-Preserving Latent Diffusion for Stable Shadow Removal

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A two-stage fine-tuning of Stable Diffusion, followed by a shadow-aware detail injection module, yields mask-free shadow removal with preserved textures and cross-dataset generalization.

  3. FiffDepth: Feed-forward Transformation of Diffusion-Based Generators for Detailed Depth Estimation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    FiffDepth transforms a pre-trained diffusion image generator into a feed-forward monocular depth estimator that combines generative detail with DINOv2-based robustness.

  4. SynthDrive: Scalable Real2Sim2Real Sensor Simulation Pipeline for High-Fidelity Asset Generation and Driving Data Synthesis

    cs.CV 2025-09 conditional novelty 5.0 of 10

    SynthDrive automatically mines images of rare objects, reconstructs them as 3D assets from a single view, and synthesizes driving footage that modestly improves detection of those objects.

  5. LaVin-DiT: Large Vision Diffusion Transformer

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A single diffusion transformer with a spatial-temporal VAE and in-context example pairs unifies over 20 vision tasks, with strong scores on several image benchmarks and uneven evidence on video tasks.

Pith tools