Pith. sign in

REVIEW 5 cited by

Xformer: Hybrid X-Shaped Transformer for Image Denoising

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.06440 v2 pith:PBBCNOMF submitted 2023-03-11 cs.CV

classification cs.CV
keywords transformerxformerdenoisingglobalimageperformstokensacross
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this paper, we present a hybrid X-shaped vision Transformer, named Xformer, which performs notably on image denoising tasks. We explore strengthening the global representation of tokens from different scopes. In detail, we adopt two types of Transformer blocks. The spatial-wise Transformer block performs fine-grained local patches interactions across tokens defined by spatial dimension. The channel-wise Transformer block performs direct global context interactions across tokens defined by channel dimension. Based on the concurrent network structure, we design two branches to conduct these two interaction fashions. Within each branch, we employ an encoder-decoder architecture to capture multi-scale features. Besides, we propose the Bidirectional Connection Unit (BCU) to couple the learned representations from these two branches while providing enhanced information fusion. The joint designs make our Xformer powerful to conduct global information modeling in both spatial and channel dimensions. Extensive experiments show that Xformer, under the comparable model complexity, achieves state-of-the-art performance on the synthetic and real-world image denoising tasks. We also provide code and models at https://github.com/gladzhang/Xformer.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MambaIRv2: Attentive State Space Restoration

    eess.IV 2024-11 conditional novelty 6.0 of 10

    MambaIRv2 modifies Mamba's state-space output with semantic prompts and reorders tokens by semantic group, achieving single-scan non-causal restoration that beats several transformer baselines.

  2. Reversing Flow for Image Restoration

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ResFlow models HQ-to-LQ degradation as a deterministic augmented flow and inverts it via velocity matching, reporting state-of-the-art restoration in fewer than four sampling steps.

  3. Complementary Advantages: Exploiting Cross-Field Frequency Correlation for NIR-Assisted Image Denoising

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A neural network that fuses NIR and RGB images in the frequency domain to denoise RGB images, achieving state-of-the-art PSNR on DVD and IVRG benchmarks.

  4. Hierarchical Information Flow for Generalized Efficient Image Restoration

    cs.CV 2024-11 conditional novelty 5.0 of 10

    Hi-IR is a hierarchical three-level attention network that achieves top results on several image restoration benchmarks while using fewer parameters than prior transformer methods.

  5. The Tenth NTIRE 2025 Image Denoising Challenge Report

    cs.CV 2025-04 conditional novelty 4.0 of 10

    A benchmark report of the NTIRE 2025 denoising challenge, where the top solution achieved 31.20 dB PSNR, outperforming the 2023 winner by 1.24 dB.

Pith tools