REVIEW 4 cited by
XPSR: Cross-modal Priors for Diffusion-based Image Super-Resolution
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Diffusion-based methods, endowed with a formidable generative prior, have received increasing attention in Image Super-Resolution (ISR) recently. However, as low-resolution (LR) images often undergo severe degradation, it is challenging for ISR models to perceive the semantic and degradation information, resulting in restoration images with incorrect content or unrealistic artifacts. To address these issues, we propose a \textit{Cross-modal Priors for Super-Resolution (XPSR)} framework. Within XPSR, to acquire precise and comprehensive semantic conditions for the diffusion model, cutting-edge Multimodal Large Language Models (MLLMs) are utilized. To facilitate better fusion of cross-modal priors, a \textit{Semantic-Fusion Attention} is raised. To distill semantic-preserved information instead of undesired degradations, a \textit{Degradation-Free Constraint} is attached between LR and its high-resolution (HR) counterpart. Quantitative and qualitative results show that XPSR is capable of generating high-fidelity and high-realism images across synthetic and real-world datasets. Codes are released at \url{https://github.com/qyp2000/XPSR}.
Forward citations
Cited by 4 Pith papers
-
Pixel-level and Semantic-level Adjustable Super-resolution: A Dual-LoRA Approach
PiSA-SR decouples pixel-level and semantic-level super-resolution into two LoRA spaces on a frozen Stable Diffusion model, enabling one-step restoration and user-tunable fidelity-perception control.
-
Adversarial Diffusion Compression for Real-World Image Super-Resolution
AdcSR distills OSEDiff into a pruned diffusion-GAN that cuts inference time 3.7x and parameters 74% while achieving comparable super-resolution quality.
-
Zoomed In, Diffused Out: Towards Local Degradation-Aware Multi-Diffusion for Extreme Image Super-Resolution
A training-free recipe uses MultiDiffusion with per-tile degradation-aware text prompts to make frozen text-to-image diffusion models super-resolve images up to 8K.
-
Efficient Difficulty-Aware Dynamic Routing for Diffusion-Based Real-World Image Super-Resolution
DDR-SR routes each real-world low-resolution image to one of two diffusion experts based on a high-frequency-loss difficulty score, using a low-compression VAE for hard images and a high-compression VAE for easy image...
Discussion (0). Continue with ORCID to comment.