Pith. sign in

REVIEW 4 major objections 5 minor 45 references

Frequency-Domain Fusion Transformer for Image Inpainting

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read A Transformer inpainting architecture that fuses wavelet and Gabor frequency features preserves high-frequency texture better than self-attention alone.

desk verdict Incremental but real architecture; the paper's own Table II contradicts its 'consistently superior' claim, so it needs major revision plus the missing Gabformer baseline. read the letter →

arxiv 2506.18437 v1 pith:7MWFQMS5 submitted 2025-06-23 cs.CV

classification cs.CV
keywords imageinpaintingfrequency-domainfusionwavelettransformGaborfilterfastFourierTransformerderaininghigh-frequencydetails
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Image inpainting with Transformers tends to smooth over fine texture because self-attention behaves like a low-pass filter, and it is computationally heavy. This paper tries to fix both problems in one architecture, Dabformer, by giving the attention mechanism a frequency-domain fusion front end and by replacing the feedforward network with a learnable FFT-based gating filter. The proposed modules decompose features with a wavelet transform, apply directionally matched Gabor filters to the high-frequency subbands, and then carry out cross-channel attention, so the model can see both global structure and directional detail. On deraining benchmarks and on images corrupted with random noise blocks, the reported PSNR and SSIM numbers match or beat strong baselines including a diffusion model, with the largest gains on dense, directional rain and heavy occlusion. The paper's central claim is that frequency-domain fusion is what preserves high-frequency information that ordinary Transformer inpainting loses.

What carries the argument

The load-bearing mechanism is the Frequency-Domain Fusion Attention (FDFA) paired with the Frequency-Domain Adaptive Gating Network (FDAGN). FDFA applies a discrete wavelet transform to the query, keeps the low-frequency LL subband for a depthwise convolution, and runs Gabor filters on the HL, LH, and HH subbands with a learnable wavelength so that filter scale adapts to local texture; the resulting query drives cross-channel attention of complexity $O(C\times C)$ rather than $O(M\times M)$. FDAGN replaces the standard feedforward network: it transforms feature blocks with the fast Fourier transform, applies a learnable complex filter initialized near identity, transforms back, and gates the result with GELU, which lets the network suppress redundant frequencies while preserving structure. The combination is what carries the paper's claim that high-frequency detail survives inpainting.

What would settle it

Run Dabformer on Places2 and CelebA with free-form irregular masks, holding everything else fixed, and compare PSNR and SSIM against Restormer, StrDiffusion, and FADformer. If the gains seen under random noise blocks shrink, vanish, or reverse on irregular masks, then the reported superiority is tied to the square-block corruption model rather than to general inpainting.

Watch

Extended reading notes

Core claim

On its own terms, the paper claims that the low-pass behavior of standard Transformer attention is the main obstacle to high-quality image inpainting, and that this obstacle can be removed by fusing two classical frequency tools into the Transformer. A wavelet decomposition splits the query into one low-frequency and three high-frequency subbands; Gabor filters aligned with each high-frequency subband's orientation extract directional textures; and a learnable FFT-domain filter takes over the feedforward network's role, adaptively suppressing noisy frequency components while retaining useful ones. Guided by a combination of L1, perceptual, edge, and SSIM losses, the four-level encoder-decoder then produces restorations that, in the paper's experiments, preserve more high-frequency detail than the compared methods. The authors present Dabformer as an extension of their earlier Gabor-guided deraining model, generalized from rain removal to damaged-image restoration.

Load-bearing premise

The inpainting experiments corrupt each image with random noise blocks of varying size and position, and the central claim assumes this synthetic corruption is a fair stand-in for real missing regions and for the irregular mask patterns used in standard inpainting benchmarks.

Editorial extensions

If this is right

  • On dense, directional rain (Rain200H, DID-Data), the method reports the best or near-best PSNR and SSIM among the compared deraining models.
  • Under 40–70% occlusion on Places2 and CelebA, Dabformer matches or exceeds diffusion-based inpainting baselines on PSNR and SSIM while using far fewer parameters (29.73M versus 114.05M for StrDiffusion).
  • Ablations show that combining wavelet and Gabor query features raises PSNR from 23.80 to 24.71 dB on CelebA at 40–50% occlusion, and adding the FFT gating network raises it further to 26.62 dB.
  • Adding perceptual, edge, and SSIM losses on top of L1 raises Rain200H PSNR from 31.97 to 32.34 dB, so each loss term contributes to detail preservation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: The same FDFA and FDAGN modules are likely transferable to other pixel-level restoration tasks such as super-resolution and deblurring, because the problem they solve—retaining high-frequency texture while suppressing noise—is not specific to rain or square-block damage.
  • Editorial inference: Because the damage model uses random square blocks rather than free-form masks, the method's inpainting claim would be tested more sharply on irregular-mask benchmarks; the texture-preservation advantage may or may not survive that change.
  • Editorial inference: The learnable wavelength in the Gabor filters suggests a natural extension to learnable orientation as well, which could handle textures whose directions do not align with the three fixed subband orientations.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes Dabformer, a Transformer-based image inpainting network that integrates wavelet transform and Gabor filtering into a frequency-domain fusion attention mechanism (FDFA) and replaces the feedforward network with a learnable frequency-domain adaptive gating network (FDAGN). The model is evaluated on four deraining datasets and on a 'damaged image restoration' task built by adding random noise blocks of varying size and position to Places2 and CelebA. The authors claim consistent state-of-the-art performance and improved high-frequency preservation for image inpainting.

Significance. The architectural idea of adaptively combining wavelet multiscale decomposition with directional Gabor filtering is timely, and the ablation study suggests that each proposed component contributes to the reported scores. However, the paper's central empirical claim is not established: Table II contradicts the claim of consistent superiority in inpainting, the evaluation protocol diverges from standard free-form-mask inpainting benchmarks, and the claimed high-frequency benefit is never directly measured. The deraining results are also not consistently superior across datasets. The paper would need a substantially revised experimental study to support its claims.

major comments (4)
  1. [Section IV-F2, Table II] The abstract and Section IV-F2 claim that the proposed method achieves 'consistently superior performance' and the highest PSNR/SSIM under moderate and heavy occlusion. Table II contradicts this: under 60-70% occlusion on Places2, StrDiffusion achieves PSNR/SSIM 20.43/0.858 versus Dabformer's 20.04/0.792; on CelebA 60-70%, StrDiffusion achieves 21.75/0.874 versus 21.44/0.842; on Places2 40-50%, Dabformer's PSNR lead over StrDiffusion is 0.03 dB (22.42 vs 22.39) while its SSIM is lower (0.780 vs 0.882). The central inpainting claim is therefore unsupported by the paper's own numbers.
  2. [Section IV-B] The damaged image restoration protocol corrupts each image with 'noise blocks of varying size and position' and does not provide a mask to the model. This is a blind restoration task, not the standard free-form-mask inpainting benchmark used in the field. Consequently, even if the results were strong, they would not establish the abstract's claim about image inpainting. The paper should evaluate on standard inpainting masks (for example, irregular masks) and compare with inpainting-specific methods.
  3. [Section IV-E, Table I] The text claims that on sparse-rain datasets 'our method still maintains a leading overall performance,' but Table I shows FADformer outperforms Dabformer on Rain200L (41.69/0.990 vs 41.66/0.990) and DDN-Data (34.42/0.960 vs 34.09/0.957). The claim of consistent superiority in deraining is inaccurate.
  4. [Section III and Section IV] The paper claims that the method preserves high-frequency information, but no frequency-domain metric or analysis is provided. The only evidence is PSNR and SSIM, which do not directly measure high-frequency fidelity. Please include spectral comparisons or a frequency-domain error metric to support the mechanism that is central to the paper's title and abstract.
minor comments (5)
  1. [Section III-B, Eq. (3)] Only Q is defined in Eq. (3); K and V are used in Eq. (4) but never defined. Please state explicitly how K and V are computed from the input feature maps.
  2. [Section III-D, Eq. (8)] The definition L_M = 1 - SSIM(O) is incomplete because SSIM takes two images; it should be written as 1 - SSIM(G_t, O).
  3. [Section III-A, Eq. (1)] The text below Eq. (1) refers to 'GTFFN modules', but the proposed module is called FDAGN. Please align the terminology.
  4. [Table II] The parameter count reported for Dabformer (29.73M) is identical to that of ICT (29.73M). Please verify that this is not a typo and provide the correct value.
  5. [Section IV-B] The noise-block corruption procedure is underspecified: the sizes, positions, number of blocks, and whether the blocks are contiguous are not stated. Please provide full details for reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central claim is an empirical architecture comparison, and the one self-citation is not load-bearing.

full rationale

The paper's derivation chain is a standard architecture paper: Eq. (1) defines the encoder block as a residual composition of FDFA and FDAGN; Eqs. (2)-(5) define Gabor filtering, frequency-domain fusion, and cross-channel attention; Eqs. (6)-(9) define a multi-term loss. None of these definitions contain the target quantity (inpainting PSNR/SSIM or high-frequency preservation) as an input, and none of the reported results are fitted parameters renamed as predictions. The learnable Gabor wavelength and the loss weights in Section III-D are trained or empirically set in the ordinary sense, and Table IV tests the loss components incrementally rather than treating a fitted weight as a performance claim. The only self-citation is Section III-C, where the post-IFFT operations of FDAGN are said to be 'consistent with our method in Gabformer [26]'; this identifies a reused building block but does not carry the inpainting claim, and the FDAGN's contribution is separately ablated in Table III. The paper is self-contained against external benchmarks (Restormer, StrDiffusion, RePaint, etc.), so the central comparison is not forced by a self-citation chain. The narrative in Section IV-F overstates Table II in places (e.g., Dabformer trails StrDiffusion in most heavy-occlusion cells and its best PSNR lead is 0.03 dB with a 0.102 SSIM deficit), but that is an evidence-vs-claim mismatch, not circularity.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The central method rests on standard signal-processing tools (wavelet, Gabor, FFT) plus several hand-set or learned hyperparameters. The evaluative protocol itself is an ad hoc assumption about what counts as inpainting.

free parameters (4)
  • Learnable Gabor wavelength = not reported (initialized to 'reasonable prior')
    Section III-B1: the wavelength of each Gabor filter is treated as a learnable parameter and optimized during training; the value is not disclosed.
  • Gabor filter fixed parameters = sigma=2*pi, psi=0, gamma=0.5
    Section III-C1: standard deviation, phase offset, and spatial aspect ratio are set by hand; only wavelength and orientation receive sensitivity analysis.
  • Loss weights = lambda1=10, lambdaP=0.6, lambdaE=0.4, lambdaM=0.5
    Section III-D: weights are 'empirically set' with no sensitivity study.
  • Channel expansion ratio = 2.66
    Section III-C2: fixed expansion ratio in training, no ablation reported.
assumptions (4)
  • standard math Wavelet transform decomposes feature maps into LL, HL, LH, HH subbands with standard DWT
    Section III-B1 uses 2D discrete wavelet transform as a fixed preprocessing step.
  • standard math Gabor filters respond selectively to oriented textures with the given parameters
    Section III-B1 applies Gabor filtering to high-frequency subbands.
  • domain assumption Self-attention is low-pass and loses high-frequency information
    Abstract and Introduction assert this without citation or derivation; it motivates the entire design.
  • ad hoc to paper Random noise block corruption models real-world damage for inpainting
    Section IV-B uses random noise blocks on Places2 and CelebA; this is not a standard inpainting mask protocol.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Frequency-Domain Fusion Transformer for Image Inpainting." pith.science (2026). https://pith.science/paper/7MWFQMS5

@misc{pith2026250618437,
  author       = {Pith},
  title        = {Pith review of: Frequency-Domain Fusion Transformer for Image Inpainting},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7MWFQMS5}},
  note         = {Machine review of arXiv:2506.18437}
}
read the original abstract

Image inpainting plays a vital role in restoring missing image regions and supporting high-level vision tasks, but traditional methods struggle with complex textures and large occlusions. Although Transformer-based approaches have demonstrated strong global modeling capabilities, they often fail to preserve high-frequency details due to the low-pass nature of self-attention and suffer from high computational costs. To address these challenges, this paper proposes a Transformer-based image inpainting method incorporating frequency-domain fusion. Specifically, an attention mechanism combining wavelet transform and Gabor filtering is introduced to enhance multi-scale structural modeling and detail preservation. Additionally, a learnable frequency-domain filter based on the fast Fourier transform is designed to replace the feedforward network, enabling adaptive noise suppression and detail retention. The model adopts a four-level encoder-decoder structure and is guided by a novel loss strategy to balance global semantics and fine details. Experimental results demonstrate that the proposed method effectively improves the quality of image inpainting by preserving more high-frequency information.

Figures

Figures reproduced from arXiv: 2506.18437 by the authors.

Figure 1
Figure 1. Comparison of Gabor and Wavelet Transform Effects on Multi [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Detailed framework of Dabformer with the main constituent modules of (a) Overall framework(Dabformer), (b) Frequency Domain Fusion Attention [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Qualitative comparison of deraining methods on the Rain200H dataset. From left to right: (a) Input rainy image, (b) Ground Truth (GT) local region, [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Qualitative comparison of deraining methods on the DID-Data dataset. From left to right: (a) Input rainy image, (b) Ground Truth (GT) local region, [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Qualitative comparison of deraining methods on the DDN-Data dataset. From left to right: (a) Input rainy image, (b) Ground Truth (GT) local region, [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Qualitative comparison of image inpainting methods on the Places2 dataset with different mask ratios. From left to right: (a) Input, (b) Ground Truth, [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: Qualitative comparison of image inpainting methods on the CelebA dataset with different mask ratios. From left to right: (a) Input, (b) Ground Truth, [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 9
Figure 9. Figure 9: Qualitative Results of Ablation Experiments on Different Loss [PITH_FULL_IMAGE:figures/full_fig_p010_9.png]
Figure 8
Figure 8. Figure 8: Ablation study with qualitative results on the CelebA dataset under [PITH_FULL_IMAGE:figures/full_fig_p010_8.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

45 extracted references · 28 canonical work pages

  1. [1]

    Recurrent feature reasoning for image inpainting,

    J. Li, N. Wang, L. Zhang, B. Du, and D. Tao, “Recurrent feature reasoning for image inpainting,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 7760– 7768

  2. [2]

    Misf: Multi- level interactive siamese filtering for high-fidelity image inpainting,

    X. Li, Q. Guo, D. Lin, P. Li, W. Feng, and S. Wang, “Misf: Multi- level interactive siamese filtering for high-fidelity image inpainting,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 1869–1878

  3. [3]

    Image inpainting with learnable bidirectional attention maps,

    C. Xie, S. Liu, C. Li, M.-M. Cheng, W. Zuo, X. Liu, S. Wen, and E. Ding, “Image inpainting with learnable bidirectional attention maps,” inProceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 8858–8867

  4. [4]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,”Advances in neural information processing systems, vol. 30, 2017

  5. [5]

    Transfill: Reference-guided image inpainting by merging multiple color and spatial transformations,

    Y . Zhou, C. Barnes, E. Shechtman, and S. Amirghodsi, “Transfill: Reference-guided image inpainting by merging multiple color and spatial transformations,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 2266–2276

  6. [6]

    High-fidelity pluralistic image completion with transformers,

    Z. Wan, J. Zhang, D. Chen, and J. Liao, “High-fidelity pluralistic image completion with transformers,” inProceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 4692–4701

  7. [7]

    Swinir: Image restoration using swin transformer,

    J. Liang, J. Cao, G. Sun, K. Zhang, L. Van Gool, and R. Timofte, “Swinir: Image restoration using swin transformer,” inProceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 1833–1844

  8. [8]

    Swinfir: Revisiting the swinir with fast fourier convolution and improved training for image super-resolution,

    D. Zhang, F. Huang, S. Liu, X. Wang, and Z. Jin, “Swinfir: Revisiting the swinir with fast fourier convolution and improved training for image super-resolution,”arXiv preprint arXiv:2208.11247, 2022

Show all 45 references
  1. [9]

    Activating more pixels in image super-resolution transformer,

    X. Chen, X. Wang, J. Zhou, Y . Qiao, and C. Dong, “Activating more pixels in image super-resolution transformer,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 22 367–22 377

  2. [10]

    Hint: High-quality inpainting transformer with mask-aware encoding and enhanced atten- tion,

    S. Chen, A. Atapour-Abarghouei, and H. P. Shum, “Hint: High-quality inpainting transformer with mask-aware encoding and enhanced atten- tion,”IEEE Transactions on Multimedia, 2024

  3. [11]

    Transref: Multi-scale reference embedding transformer for reference- guided image inpainting,

    T. Liu, L. Liao, D. Chen, J. Xiao, Z. Wang, C.-W. Lin, and S. Satoh, “Transref: Multi-scale reference embedding transformer for reference- guided image inpainting,”Neurocomputing, vol. 632, p. 129749, 2025

  4. [12]

    Restormer: Efficient transformer for high-resolution image restoration,

    S. W. Zamir, A. Arora, S. Khan, M. Hayat, F. S. Khan, and M.-H. Yang, “Restormer: Efficient transformer for high-resolution image restoration,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 5728–5739

  5. [13]

    A practical guide to wavelet analysis,

    C. Torrence and G. P. Compo, “A practical guide to wavelet analysis,” Bulletin of the American Meteorological society, vol. 79, no. 1, pp. 61– 78, 1998

  6. [14]

    A review of convolutional neural networks and gabor filters in object recognition,

    M. Rai and P. Rivas, “A review of convolutional neural networks and gabor filters in object recognition,” in2020 International Conference on Computational Science and Computational Intelligence (CSCI). IEEE, 2020, pp. 1560–1567

  7. [15]

    Ft-tdr: Frequency-guided transformer and top-down refinement network for blind face inpainting,

    J. Wang, S. Chen, Z. Wu, and Y .-G. Jiang, “Ft-tdr: Frequency-guided transformer and top-down refinement network for blind face inpainting,” IEEE Transactions on Multimedia, vol. 25, pp. 2382–2392, 2022

  8. [16]

    Bridging global context interactions for high-fidelity image completion,

    C. Zheng, T.-J. Cham, J. Cai, and D. Phung, “Bridging global context interactions for high-fidelity image completion,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 11 512–11 522. 12

  9. [17]

    Incremental transformer structure en- hanced image inpainting with masking positional encoding,

    Q. Dong, C. Cao, and Y . Fu, “Incremental transformer structure en- hanced image inpainting with masking positional encoding,” inPro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 11 358–11 368

  10. [18]

    Reduce information loss in transformers for pluralistic image inpainting,

    Q. Liu, Z. Tan, D. Chen, Q. Chu, X. Dai, Y . Chen, M. Liu, L. Yuan, and N. Yu, “Reduce information loss in transformers for pluralistic image inpainting,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 11 347–11 357

  11. [19]

    Mat: Mask- aware transformer for large hole image inpainting,

    W. Li, Z. Lin, K. Zhou, L. Qi, Y . Wang, and J. Jia, “Mat: Mask- aware transformer for large hole image inpainting,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 10 758–10 768

  12. [20]

    Image restoration refine- ment with uformer gan,

    X. Ouyang, Y . Chen, K. Zhu, and G. Agam, “Image restoration refine- ment with uformer gan,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 5919–5928

  13. [21]

    Blind image inpainting via omni-dimensional gated attention and wavelet queries,

    S. S. Phutke, A. Kulkarni, S. K. Vipparthi, and S. Murala, “Blind image inpainting via omni-dimensional gated attention and wavelet queries,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 1251–1260

  14. [22]

    Multi-dimensional visual prompt enhanced image restoration via mamba-transformer ag- gregation,

    A. Jiang, H. Chen, Z. Chen, J. Ye, and M. Wang, “Multi-dimensional visual prompt enhanced image restoration via mamba-transformer ag- gregation,”arXiv preprint arXiv:2412.15845, 2024

  15. [23]

    Real-time single image and video super- resolution using an efficient sub-pixel convolutional neural network,

    W. Shi, J. Caballero, F. Husz ´ar, J. Totz, A. P. Aitken, R. Bishop, D. Rueckert, and Z. Wang, “Real-time single image and video super- resolution using an efficient sub-pixel convolutional neural network,” in Proceedings of the IEEE conference on computer vision and pattern r...

  16. [24]

    U-net: Convolutional networks for biomedical image segmentation,

    O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” inMedical image computing and computer-assisted intervention–MICCAI 2015: 18th international con- ference, Munich, Germany, October 5-9, 2015, proceedings, part III 18. ...

  17. [25]

    Gabor filter-based edge detection,

    R. Mehrotra, K. R. Namuduri, and N. Ranganathan, “Gabor filter-based edge detection,”Pattern recognition, vol. 25, no. 12, pp. 1479–1494, 1992

  18. [26]

    Gabor-guided transformer for single image deraining,

    S. He, G. Lin, and J. Feng, “Gabor-guided transformer for single image deraining,” in2024 10th International Conference on Systems and Informatics (ICSAI), 2024, pp. 1–6

  19. [27]

    Very deep convolutional networks for large-scale image recognition,

    K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,”arXiv preprint arXiv:1409.1556, 2014

  20. [28]

    Removing rain from a single image via discriminative sparse coding,

    Y . Luo, Y . Xu, and H. Ji, “Removing rain from a single image via discriminative sparse coding,” inProceedings of the IEEE international conference on computer vision, 2015, pp. 3397–3405

  21. [29]

    Rain streak removal using layer priors,

    Y . Li, R. T. Tan, X. Guo, J. Lu, and M. S. Brown, “Rain streak removal using layer priors,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 2736–2744

  22. [30]

    Multi-scale progressive fusion network for single image deraining,

    K. Jiang, Z. Wang, P. Yi, C. Chen, B. Huang, Y . Luo, J. Ma, and J. Jiang, “Multi-scale progressive fusion network for single image deraining,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 8346–8355

  23. [31]

    Progressive image deraining networks: A better and simpler baseline,

    D. Ren, W. Zuo, Q. Hu, P. Zhu, and D. Meng, “Progressive image deraining networks: A better and simpler baseline,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 3937–3946

  24. [32]

    Multi-stage progressive image restoration,

    S. W. Zamir, A. Arora, S. Khan, M. Hayat, F. S. Khan, M.-H. Yang, and L. Shao, “Multi-stage progressive image restoration,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 14 821–14 831

  25. [33]

    Image de-raining trans- former,

    J. Xiao, X. Fu, A. Liu, F. Wu, and Z.-J. Zha, “Image de-raining trans- former,”IEEE transactions on pattern analysis and machine intelligence, vol. 45, no. 11, pp. 12 978–12 995, 2022

  26. [34]

    Hybrid cnn-transformer feature fusion for single image deraining,

    X. Chen, J. Pan, J. Lu, Z. Fan, and H. Li, “Hybrid cnn-transformer feature fusion for single image deraining,” inProceedings of the AAAI conference on artificial intelligence, vol. 37, no. 1, 2023, pp. 378–386

  27. [35]

    Learning a sparse transformer network for effective image deraining,

    X. Chen, H. Li, M. Li, and J. Pan, “Learning a sparse transformer network for effective image deraining,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 5896– 5905

  28. [36]

    Efficient frequency- domain image deraining with contrastive regularization,

    N. Gao, X. Jiang, X. Zhang, and Y . Deng, “Efficient frequency- domain image deraining with contrastive regularization,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 240–257

  29. [37]

    Deep joint rain detection and removal from a single image,

    W. Yang, R. T. Tan, J. Feng, J. Liu, Z. Guo, and S. Yan, “Deep joint rain detection and removal from a single image,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 1357–1366

  30. [38]

    Removing rain from single images via a deep detail network,

    X. Fu, J. Huang, D. Zeng, Y . Huang, X. Ding, and J. Paisley, “Removing rain from single images via a deep detail network,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 3855–3863

  31. [39]

    Density-aware single image de-raining using a multi-stream dense network,

    H. Zhang and V . M. Patel, “Density-aware single image de-raining using a multi-stream dense network,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 695–704

  32. [40]

    Places: A 10 million image database for scene recognition,

    B. Zhou, A. Lapedriza, A. Khosla, A. Oliva, and A. Torralba, “Places: A 10 million image database for scene recognition,”IEEE transactions on pattern analysis and machine intelligence, vol. 40, no. 6, pp. 1452–1464, 2017

  33. [41]

    Deep learning face attributes in the wild,

    Z. Liu, P. Luo, X. Wang, and X. Tang, “Deep learning face attributes in the wild,” inProceedings of the IEEE international conference on computer vision, 2015, pp. 3730–3738

  34. [42]

    Cmt: Convolutional neural networks meet vision transformers,

    J. Guo, K. Han, H. Wu, Y . Tang, X. Chen, Y . Wang, and C. Xu, “Cmt: Convolutional neural networks meet vision transformers,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 12 175–12 185

  35. [43]

    Repaint: Inpainting using denoising diffusion probabilistic models,

    A. Lugmayr, M. Danelljan, A. Romero, F. Yu, R. Timofte, and L. Van Gool, “Repaint: Inpainting using denoising diffusion probabilistic models,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 11 461–11 471

  36. [44]

    Image restoration with mean-reverting stochastic differential equations,

    Z. Luo, F. K. Gustafsson, Z. Zhao, J. Sj ¨olund, and T. B. Sch ¨on, “Image restoration with mean-reverting stochastic differential equations,”arXiv preprint arXiv:2301.11699, 2023

  37. [45]

    Structure matters: Tackling the semantic discrepancy in diffusion models for image inpaint- ing,

    H. Liu, Y . Wang, B. Qian, M. Wang, and Y . Rui, “Structure matters: Tackling the semantic discrepancy in diffusion models for image inpaint- ing,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 8038–8047

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.