Pith. sign in

REVIEW 1 cited by

IPT-V2: Efficient Image Processing Transformer using Hierarchical Attentions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.00633 v1 pith:N4DB4P7T submitted 2024-03-31 cs.CV

classification cs.CV
keywords imagegloballocalprocessingself-attentiontransformerarchitectureattentions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances have demonstrated the powerful capability of transformer architecture in image restoration. However, our analysis indicates that existing transformerbased methods can not establish both exact global and local dependencies simultaneously, which are much critical to restore the details and missing content of degraded images. To this end, we present an efficient image processing transformer architecture with hierarchical attentions, called IPTV2, adopting a focal context self-attention (FCSA) and a global grid self-attention (GGSA) to obtain adequate token interactions in local and global receptive fields. Specifically, FCSA applies the shifted window mechanism into the channel self-attention, helps capture the local context and mutual interaction across channels. And GGSA constructs long-range dependencies in the cross-window grid, aggregates global information in spatial dimension. Moreover, we introduce structural re-parameterization technique to feed-forward network to further improve the model capability. Extensive experiments demonstrate that our proposed IPT-V2 achieves state-of-the-art results on various image processing tasks, covering denoising, deblurring, deraining and obtains much better trade-off for performance and computational complexity than previous methods. Besides, we extend our method to image generation as latent diffusion backbone, and significantly outperforms DiTs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Degradation-Modeled Multipath Diffusion for Tunable Metalens Photography

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A multipath diffusion model, guided by simulated lens blur and image-quality scores, restores sharp images from a custom 1 mm3 metalens camera, beating published baselines on the authors' test set.

Pith tools