Pith. sign in

REVIEW 6 cited by

SwinFIR: Revisiting the SwinIR with Fast Fourier Convolution and Improved Training for Image Super-Resolution

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.11247 v3 pith:3L3PJJ5H submitted 2022-08-24 cs.CV

classification cs.CV
keywords performanceswinirimagemethodsswinfirachievedconvolutionensemble
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transformer-based methods have achieved impressive image restoration performance due to their capacities to model long-range dependency compared to CNN-based methods. However, advances like SwinIR adopts the window-based and local attention strategy to balance the performance and computational overhead, which restricts employing large receptive fields to capture global information and establish long dependencies in the early layers. To further improve the efficiency of capturing global information, in this work, we propose SwinFIR to extend SwinIR by replacing Fast Fourier Convolution (FFC) components, which have the image-wide receptive field. We also revisit other advanced techniques, i.e, data augmentation, pre-training, and feature ensemble to improve the effect of image reconstruction. And our feature ensemble method enables the performance of the model to be considerably enhanced without increasing the training and testing time. We applied our algorithm on multiple popular large-scale benchmarks and achieved state-of-the-art performance comparing to the existing methods. For example, our SwinFIR achieves the PSNR of 32.83 dB on Manga109 dataset, which is 0.8 dB higher than the state-of-the-art SwinIR method.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. 4KAgent: Agentic Any Image to 4K Super-Resolution

    cs.CV 2025-07 reject novelty 6.0 of 10

    An agentic pipeline that plans and executes image restoration from a toolbox of pretrained models to upscale arbitrary images to 4K, reporting state-of-the-art results on many benchmarks.

  2. WaveHiT-SR: Hierarchical Wavelet Network for Efficient Image Super-Resolution

    cs.CV 2025-08 conditional novelty 5.0 of 10

    WaveHiT-SR embeds discrete wavelet transforms into hierarchical transformer blocks, producing efficient super-resolution models with modest PSNR gains over SwinIR and SRFormer baselines.

  3. NTIRE 2025 Challenge on RAW Image Restoration and Super-Resolution

    eess.IV 2025-06 conditional novelty 5.0 of 10

    The NTIRE 2025 challenge report benchmarks raw-image restoration and 2x super-resolution, with Samsung AI's methods winning both tracks on a synthetic test set.

  4. DiMoSR: Feature Modulation via Multi-Branch Dilated Convolutions for Efficient Image Super-Resolution

    cs.CV 2025-05 conditional novelty 5.0 of 10

    DiMoSR, a lightweight super-resolution network using multi-branch dilated convolutions and feature modulation, reports state-of-the-art PSNR/SSIM on four benchmark datasets at x4 upscaling.

  5. Frequency-Domain Fusion Transformer for Image Inpainting

    cs.CV 2025-06 reject novelty 4.0 of 10

    Dabformer combines wavelet and Gabor filtered attention with an FFT gating network, but the reported gains over prior methods are inconsistent across datasets.

  6. Revealing the Ancient Beauty: Digital Reconstruction of Temple Tiles using Computer Vision

    cs.CV 2025-07 reject novelty 2.0 of 10

    A pipeline combining YOLOv8, a GAN variant, and stage-wise super-resolution is proposed for temple tile restoration, but no quantitative validation is provided for the end-to-end system.

Pith tools