Pith. sign in

REVIEW 2 cited by

PAN-Crafter: Learning Modality-Consistent Alignment for PAN-Sharpening

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.23367 v2 pith:PXF6J5IT submitted 2025-05-29 cs.CV cs.AI

PAN-Crafter: Learning Modality-Consistent Alignment for PAN-Sharpening

classification cs.CV cs.AI
keywords alignmentimagesmisalignmentpan-crafteracrosscross-modalitydatasetshigh-resolution
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

PAN-sharpening aims to fuse high-resolution panchromatic (PAN) images with low-resolution multi-spectral (MS) images to generate high-resolution multi-spectral (HRMS) outputs. However, cross-modality misalignment -- caused by sensor placement, acquisition timing, and resolution disparity -- induces a fundamental challenge. Conventional deep learning methods assume perfect pixel-wise alignment and rely on per-pixel reconstruction losses, leading to spectral distortion, double edges, and blurring when misalignment is present. To address this, we propose PAN-Crafter, a modality-consistent alignment framework that explicitly mitigates the misalignment gap between PAN and MS modalities. At its core, Modality-Adaptive Reconstruction (MARs) enables a single network to jointly reconstruct HRMS and PAN images, leveraging PAN's high-frequency details as auxiliary self-supervision. Additionally, we introduce Cross-Modality Alignment-Aware Attention (CM3A), a novel mechanism that bidirectionally aligns MS texture to PAN structure and vice versa, enabling adaptive feature refinement across modalities. Extensive experiments on multiple benchmark datasets demonstrate that our PAN-Crafter outperforms the most recent state-of-the-art method in all metrics, even with 50.11$\times$ faster inference time and 0.63$\times$ the memory size. Furthermore, it demonstrates strong generalization performance on unseen satellite datasets, showing its robustness across different conditions.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MetaTele: Compact Refractive Metasurface Computational Telephoto Camera

    eess.IV 2026-04 conditional novelty 6.5

    Hybrid refractive-metasurface optics plus one-step diffusion fusion of structure and color measurements yields RGB telephoto imaging at telephoto ratio 0.44 with 13 mm TTL.

  2. MetaTele: Compact Refractive Metasurface Computational Telephoto Camera

    eess.IV 2026-04 unverdicted novelty 5.0

    A metasurface-refractive optics plus diffusion-model co-design achieves 0.44 telephoto ratio and 13 mm total track length for RGB imaging in a compact prototype.