Pith. sign in

REVIEW 23 cited by

Exploiting Diffusion Prior for Real-World Image Super-Resolution

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.07015 v4 pith:SKB2Q2ID submitted 2023-05-11 cs.CV

classification cs.CV
keywords diffusionmodelspre-trainedpriorfidelityreal-worldsuper-resolutionachieve
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present a novel approach to leverage prior knowledge encapsulated in pre-trained text-to-image diffusion models for blind super-resolution (SR). Specifically, by employing our time-aware encoder, we can achieve promising restoration results without altering the pre-trained synthesis model, thereby preserving the generative prior and minimizing training cost. To remedy the loss of fidelity caused by the inherent stochasticity of diffusion models, we employ a controllable feature wrapping module that allows users to balance quality and fidelity by simply adjusting a scalar value during the inference process. Moreover, we develop a progressive aggregation sampling strategy to overcome the fixed-size constraints of pre-trained diffusion models, enabling adaptation to resolutions of any size. A comprehensive evaluation of our method using both synthetic and real-world benchmarks demonstrates its superiority over current state-of-the-art approaches. Code and models are available at https://github.com/IceClear/StableSR.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 23 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MVGBench: Comprehensive Benchmark for Multi-view Generation Models

    cs.GR 2025-06 conditional novelty 7.0 of 10

    MVGBench evaluates multi-view generators through self-consistency of 3D reconstructions and uses this protocol to rank 12 models and build a better one.

  2. When Latents Forget Pixels: Restoring Fidelity in Diffusion Transformer Super-Resolution

    cs.CV 2026-08 conditional novelty 6.0 of 10

    By injecting pre-VAE pixel features into both the latent denoising trajectory and the frozen VAE decoder, PGSR improves fidelity of latent diffusion transformer super-resolution while maintaining perceptual quality.

  3. Detail Continuation over a Trustworthy Coarse Scale for Autoregressive Super-Resolution

    cs.CV 2026-08 conditional novelty 6.0 of 10

    In VAR-based super-resolution, K2N predicts the first three coarse scales in parallel from the low-resolution input and generates only the remaining fine scales autoregressively, reducing hallucination while staying c...

  4. Fine-structure Preserved Real-world Image Super-resolution via Transfer VAE Training

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A transfer training scheme converts Stable Diffusion's 8x VAE into a 4x VAE that stays compatible with the pretrained UNet, improving fine-structure preservation in real-world super-resolution at lower FLOPs.

  5. Robust ID-Specific Face Restoration via Alignment Learning

    cs.CV 2025-07 conditional novelty 6.0 of 10

    RIDFR injects a reference person's identity into diffusion-based face restoration and uses Alignment Learning across multiple same-identity references to suppress pose, expression, and makeup interference.

  6. PICD: Versatile Perceptual Image Compression with Diffusion Rendering

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A screen-and-natural-image codec that losslessly encodes OCR text and uses a diffusion renderer with three levels of conditioning achieves state-of-the-art perceptual quality and high text accuracy at 0.005 to 0.05 bpp.

  7. InstaRevive: One-Step Image Enhancement via Dynamic Score Matching

    cs.CV 2025-04 conditional novelty 6.0 of 10

    A one-step image enhancer that combines dynamically controlled score-based distillation with caption prompts, matching multi-step diffusion quality on face and super-resolution benchmarks.

  8. Zero-shot Depth Completion via Test-time Alignment with Affine-invariant Depth Prior

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Zero-shot depth completion by aligning an affine-invariant depth diffusion prior to sparse metric measurements through test-time optimization achieves domain-generalizable dense depth without training on depth complet...

  9. INDIGO+: A Unified INN-Guided Probabilistic Diffusion Algorithm for Blind and Non-Blind Image Restoration

    cs.CV 2025-01 conditional novelty 6.0 of 10

    An invertible neural network trained to mimic image degradations is used to steer a pretrained diffusion model in every sampling step, giving a blind and non-blind image restoration algorithm.

  10. Navigating Image Restoration with VAR's Distribution Alignment Prior

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A unified image restoration framework, VarFormer, repurposes the scale-wise latent features of the pretrained generative model VAR as a distribution-alignment prior and reports state-of-the-art results across six degr...

  11. F-Bench: Rethinking Human Preference Evaluation Metrics for Benchmarking Face Generation, Customization, and Restoration

    cs.CV 2024-12 conditional novelty 6.0 of 10

    FaceQ, a new 12K-image benchmark with multi-dimensional human preference scores, reveals that existing quality metrics poorly match human judgment on AI-generated faces, and F-Eval, an instruction-tuned LMM, outperforms them.

  12. Pixel-level and Semantic-level Adjustable Super-resolution: A Dual-LoRA Approach

    cs.CV 2024-12 conditional novelty 6.0 of 10

    PiSA-SR decouples pixel-level and semantic-level super-resolution into two LoRA spaces on a frozen Stable Diffusion model, enabling one-step restoration and user-tunable fidelity-perception control.

  13. MicroZoom: Structure-Preserving Detail Synthesis at Extreme Scale

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A cascaded, segmentation-conditioned, per-instance diffusion method synthesizes globally coherent gigapixel microscopic detail from a phone photo and sparse microscope references at up to 350×.

  14. DECAF: De-Clustering for Adaptive Representational Unlearning

    cs.LG 2026-07 conditional novelty 5.0 of 10

    DECAF is a forget-only unlearning method that adds input noise, suppresses the forget-class probability, and diversifies outputs, achieving 0.10% forget accuracy and 79.4% retain accuracy on CIFAR-10/ResNet-18 while d...

  15. FS-Diff: Semantic guidance and clarity-aware simultaneous multimodal image fusion and super-resolution

    cs.CV 2025-09 conditional novelty 5.0 of 10

    FS-Diff is a diffusion model that jointly fuses and super-resolves low-resolution multimodal image pairs using clarity-aware CLIP semantics and a bidirectional Mamba feature extractor.

  16. Efficient Burst Super-Resolution with One-step Diffusion

    cs.CV 2025-07 reject novelty 5.0 of 10

    E-BSRD applies EDM sampling and consistency-model distillation to burst super-resolution, achieving one-step diffusion at 0.44 s/image while roughly matching or slightly underperforming the multi-step BSRD baseline on...

  17. Visual Autoregressive Modeling for Image Super-Resolution

    cs.CV 2025-01 conditional novelty 5.0 of 10

    VARSR shows that next-scale visual autoregressive prediction, augmented with diffusion-based quantization residual refinement, can produce competitive perceptual-quality super-resolution at roughly ten times lower inf...

  18. DiffStereo: High-Frequency Aware Diffusion Model for Stereo Image Restoration

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A diffusion model that generates latent high-frequency maps from low-quality stereo images and injects them into a transformer restoration network yields modest gains on stereo super-resolution, deblurring, and low-li...

  19. Diffusion Prior Interpolation for Flexibility Real-World Face Super-Resolution

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A diffusion-based face super-resolution method using fixed and random masks plus a trained corrector network reports state-of-the-art perceptual quality and face recognition consistency on common benchmarks.

  20. Controllable Distortion-Perception Tradeoff Through Latent Diffusion for Neural Image Compression

    eess.IV 2024-12 conditional novelty 5.0 of 10

    A pretrained neural image codec can be augmented with a decoder-side latent diffusion module and a τ slider that interpolates between low-distortion and high-perception reconstructions without changing the bitstream.

  21. Efficient Difficulty-Aware Dynamic Routing for Diffusion-Based Real-World Image Super-Resolution

    cs.CV 2026-07 reject novelty 4.0 of 10

    DDR-SR routes each real-world low-resolution image to one of two diffusion experts based on a high-frequency-loss difficulty score, using a low-compression VAE for hard images and a high-compression VAE for easy image...

  22. Incorporating Uncertainty-Guided and Top-k Codebook Matching for Real-World Blind Image Super-Resolution

    cs.CV 2025-06 conditional novelty 4.0 of 10

    UGTSR improves codebook-based blind super-resolution by combining uncertainty-guided loss weighting, top-3 codebook matching, and an align-attention module for fusing low- and high-quality features.

  23. Optimizing Few-Step Sampler for Diffusion Probabilistic Model

    cs.CV 2024-12 reject novelty 3.0 of 10

    A diffusion model can be made faster by alternately optimizing the noise schedule and finetuning the denoiser, yielding lower FID at small step counts on three datasets.

Pith tools