Pith. sign in

REVIEW 4 cited by

XPSR: Cross-modal Priors for Diffusion-based Image Super-Resolution

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.05049 v2 pith:BGEHONEU submitted 2024-03-08 cs.CV

classification cs.CV
keywords xpsrcross-modalimagespriorssuper-resolutiontextitattentiondegradation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion-based methods, endowed with a formidable generative prior, have received increasing attention in Image Super-Resolution (ISR) recently. However, as low-resolution (LR) images often undergo severe degradation, it is challenging for ISR models to perceive the semantic and degradation information, resulting in restoration images with incorrect content or unrealistic artifacts. To address these issues, we propose a \textit{Cross-modal Priors for Super-Resolution (XPSR)} framework. Within XPSR, to acquire precise and comprehensive semantic conditions for the diffusion model, cutting-edge Multimodal Large Language Models (MLLMs) are utilized. To facilitate better fusion of cross-modal priors, a \textit{Semantic-Fusion Attention} is raised. To distill semantic-preserved information instead of undesired degradations, a \textit{Degradation-Free Constraint} is attached between LR and its high-resolution (HR) counterpart. Quantitative and qualitative results show that XPSR is capable of generating high-fidelity and high-realism images across synthetic and real-world datasets. Codes are released at \url{https://github.com/qyp2000/XPSR}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Pixel-level and Semantic-level Adjustable Super-resolution: A Dual-LoRA Approach

    cs.CV 2024-12 conditional novelty 6.0 of 10

    PiSA-SR decouples pixel-level and semantic-level super-resolution into two LoRA spaces on a frozen Stable Diffusion model, enabling one-step restoration and user-tunable fidelity-perception control.

  2. Adversarial Diffusion Compression for Real-World Image Super-Resolution

    eess.IV 2024-11 conditional novelty 6.0 of 10

    AdcSR distills OSEDiff into a pruned diffusion-GAN that cuts inference time 3.7x and parameters 74% while achieving comparable super-resolution quality.

  3. Zoomed In, Diffused Out: Towards Local Degradation-Aware Multi-Diffusion for Extreme Image Super-Resolution

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A training-free recipe uses MultiDiffusion with per-tile degradation-aware text prompts to make frozen text-to-image diffusion models super-resolve images up to 8K.

  4. Efficient Difficulty-Aware Dynamic Routing for Diffusion-Based Real-World Image Super-Resolution

    cs.CV 2026-07 reject novelty 4.0 of 10

    DDR-SR routes each real-world low-resolution image to one of two diffusion experts based on a high-frequency-loss difficulty score, using a low-compression VAE for hard images and a high-compression VAE for easy image...

Pith tools