Pith. sign in

REVIEW 3 cited by

Unsupervised Misaligned Infrared and Visible Image Fusion via Cross-Modality Image Generation and Registration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.11876 v1 pith:6QGHDTJ4 submitted 2022-05-24 cs.CV

Unsupervised Misaligned Infrared and Visible Image Fusion via Cross-Modality Image Generation and Registration

classification cs.CV
keywords imageinfraredfusioncross-modalitymisalignedvisibleimagesnetwork
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Recent learning-based image fusion methods have marked numerous progress in pre-registered multi-modality data, but suffered serious ghosts dealing with misaligned multi-modality data, due to the spatial deformation and the difficulty narrowing cross-modality discrepancy. To overcome the obstacles, in this paper, we present a robust cross-modality generation-registration paradigm for unsupervised misaligned infrared and visible image fusion (IVIF). Specifically, we propose a Cross-modality Perceptual Style Transfer Network (CPSTN) to generate a pseudo infrared image taking a visible image as input. Benefiting from the favorable geometry preservation ability of the CPSTN, the generated pseudo infrared image embraces a sharp structure, which is more conducive to transforming cross-modality image alignment into mono-modality registration coupled with the structure-sensitive of the infrared image. In this case, we introduce a Multi-level Refinement Registration Network (MRRN) to predict the displacement vector field between distorted and pseudo infrared images and reconstruct registered infrared image under the mono-modality setting. Moreover, to better fuse the registered infrared images and visible images, we present a feature Interaction Fusion Module (IFM) to adaptively select more meaningful features for fusion in the Dual-path Interaction Fusion Network (DIFN). Extensive experimental results suggest that the proposed method performs superior capability on misaligned cross-modality image fusion.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. BeyondFusion: Self-Aligned Latent Diffusion for Calibration-Free Infrared Super-Resolution and Infrared-Visible Fusion

    cs.CV 2026-07 conditional novelty 5.0

    One latent diffusion model, with token-level cross-modal attention, performs calibration-free visible-guided infrared super-resolution and infrared-visible fusion as two outputs of the same process.

  2. Uncertainty-aware Spatial-Frequency Registration and Fusion for Infrared and Visible Images

    cs.CV 2026-05 unverdicted novelty 5.0

    SFRF combines uncertainty-aware multi-scale registration with frequency-domain thermal consistency and dual-branch fusion to handle unregistered infrared-visible image pairs.

  3. FS-Diff: Semantic guidance and clarity-aware simultaneous multimodal image fusion and super-resolution

    cs.CV 2025-09 conditional novelty 5.0

    FS-Diff is a diffusion model that jointly fuses and super-resolves low-resolution multimodal image pairs using clarity-aware CLIP semantics and a bidirectional Mamba feature extractor.