REVIEW 5 cited by
HyperSIGMA: Hyperspectral Intelligence Comprehension Foundation Model
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Accurate hyperspectral image (HSI) interpretation is critical for providing valuable insights into various earth observation-related applications such as urban planning, precision agriculture, and environmental monitoring. However, existing HSI processing methods are predominantly task-specific and scene-dependent, which severely limits their ability to transfer knowledge across tasks and scenes, thereby reducing the practicality in real-world applications. To address these challenges, we present HyperSIGMA, a vision transformer-based foundation model that unifies HSI interpretation across tasks and scenes, scalable to over one billion parameters. To overcome the spectral and spatial redundancy inherent in HSIs, we introduce a novel sparse sampling attention (SSA) mechanism, which effectively promotes the learning of diverse contextual features and serves as the basic block of HyperSIGMA. HyperSIGMA integrates spatial and spectral features using a specially designed spectral enhancement module. In addition, we construct a large-scale hyperspectral dataset, HyperGlobal-450K, for pre-training, which contains about 450K hyperspectral images, significantly surpassing existing datasets in scale. Extensive experiments on various high-level and low-level HSI tasks demonstrate HyperSIGMA's versatility and superior representational capability compared to current state-of-the-art methods. Moreover, HyperSIGMA shows significant advantages in scalability, robustness, cross-modal transferring capability, real-world applicability, and computational efficiency. The code and models will be released at https://github.com/WHU-Sigma/HyperSIGMA.
Forward citations
Cited by 5 Pith papers
-
SpectralX: Parameter-efficient Domain Generalization for Spectral Remote Sensing Foundation Models
SpectralX adapts RGB-pretrained remote sensing models to spectral data with a two-stage parameter-efficient training scheme and reports state-of-the-art cross-domain segmentation on three benchmarks.
-
M-SpecGene: Generalized Foundation Model for RGBT Multispectral Vision
M-SpecGene is a Siamese masked-autoencoder foundation model for RGB-thermal vision, trained on the RGBT550K dataset with a GMM-CMSS progressive masking strategy, and evaluated on four downstream tasks.
-
MAPEX: Modality-Aware Pruning of Experts for Remote Sensing Foundation Models
MAPEX shows that a modality-conditioned mixture-of-experts vision transformer, pre-trained on six remote sensing modalities and then pruned to keep only the experts for a target modality, can outperform or match large...
-
MergeSAM: Unsupervised change detection of remote sensing images based on the Segment Anything Model
A SAM-based unsupervised change detection method that matches and splits segmentation masks across two dates, improving F1 over AnyChange on GZ_CD_data.
-
Parameter-Efficient Fine-Tuning of Multispectral Foundation Models for Hyperspectral Image Classification
KronA+ fine-tunes SpectralGPT for hyperspectral image classification using only 0.056% trainable parameters and reaches accuracy close to full fine-tuning on five public datasets.
Discussion (0). Continue with ORCID to comment.