REVIEW 5 major objections 5 minor 42 references
DiffFormer: a Differential Spatial-Spectral Transformer for Hyperspectral Image Classification
T0 review · 5 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read DiffFormer, a spatial-spectral transformer whose attention module scores differences between neighboring tokens, reports state-of-the-art accuracy on four hyperspectral benchmarks.
desk verdict A genuinely new differential-attention trick buried under an evaluation that is not trustworthy as written. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Differential Multi-Head Self-Attention (DMHSA) module. Standard multi-head attention computes scores $S = QK^\top/\sqrt{d_{\text{head}}}$; DMHSA forms $S_{\text{diff}} = S[:, 1:] - S[:, :-1]$, the difference between each token's attention-score row and the row before it, and uses these differential scores to weight the value vectors. This one subtraction is meant to replace absolute similarities with local transitions between neighboring spectral-spatial patches, producing attention that is sparse, noise-robust, and sensitive to subtle spectral variation. The supporting machinery—3D convolution patch embeddings, sinusoidal positional encoding, a learnable class token, transformer layers with SWiGLU activation—feeds and stabilizes that differential attention.
What would settle it
Re-run the comparison on all four datasets with spatially disjoint training and test regions, and check whether DiffFormer still beats WaveFormer, AGCN, and the other baselines by comparable OA and kappa margins.
Extended reading notes
Core claim
The central claim is that DiffFormer, through its Differential Multi-Head Self-Attention (DMHSA) module, achieves state-of-the-art hyperspectral image classification. On the HanChuan, University of Houston, Salinas, and Pavia University datasets it reports the highest overall accuracy, average accuracy, and kappa coefficient among eight compared methods, with OA values of 99.3137%, 99.6229%, 99.8152%, and 99.4623% respectively. In the attention ablation, DMHSA outperforms both standard multi-head self-attention and multi-head cross-attention on all four datasets by 0.51 to 1.39 percentage points in OA, which the paper takes as evidence that the differential operation, not the transformer backbone alone, drives the improvement.
Load-bearing premise
The reported accuracies assume the random train/validation/test split yields independent samples, but the paper never states whether patches are extracted before the split or whether overlapping patches are allowed; if the same or neighboring pixels appear in both training and test sets, the accuracy tables would be inflated.
Editorial extensions
If this is right
- Existing spatial-spectral transformers could adopt DMHSA by adding one subtraction per attention head, potentially lifting hyperspectral classification accuracy at negligible extra cost.
- The reported margins are not tied to the best patch size: the comparison tables use 12x12 patches, while the patch-size study finds that larger patches (18x18 or 20x20) often do even better on kappa and OA.
- Training-set scaling shows diminishing returns beyond roughly 35% training data, and shallow models with one to three transformer layers are often sufficient, so the method is practical in data-limited settings.
- DMHSA's advantage over MHSA and MHCA holds across all four datasets, suggesting the differential operation generalizes across sensors, spatial resolutions, and class distributions.
- By the paper's complexity analysis, the differential operation does not change the quadratic attention complexity $O(N_{\text{patch}}^2 d_{\text{head}})$, so the accuracy gain is not bought with asymptotic extra computation.
Reading between the lines
- Because DMHSA is defined on any sequence of tokens, the same differential-attention trick could be dropped into transformer models for other remote-sensing tasks—change detection, multi-temporal analysis, boundary extraction—where local transitions carry signal.
- The paper fixes the differential operation as a subtraction of adjacent rows; a learnable variant, for example a small convolution along the token axis, could adaptively mix absolute and differential scores, and comparing it with DMHSA would test whether the fixed subtraction is the optimal form.
- The reported evaluation uses random pixel-level splits; a stricter stress test would partition each image into spatially disjoint training and testing regions and check whether the OA and kappa advantages persist, because hyperspectral ground truth is strongly spatially autocorrelated.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes DiffFormer, a transformer architecture for hyperspectral image classification whose main novelty is a Differential Multi-Head Self-Attention (DMHSA) mechanism. The architecture combines 3D-convolution-based spectral-spatial patch tokenization, sinusoidal positional encoding, transformer layers with SWiGLU activation, and a class-token classification head. Experiments are reported on four benchmark datasets (HanChuan, University of Houston, Salinas, and Pavia University), with comparisons against seven recent methods (AGCN, Former, PyFormer, WaveFormer, HViT, MHSSMamba, and WaveMamba). The paper reports state-of-the-art aggregate metrics (e.g., OA 99.62% on UH, 99.82% on SA), runtimes, and ablation studies over patch size, training-sample percentage, transformer depth, attention heads, and attention mechanism.
Significance. If the reported results are reproducible, DiffFormer would be a competitive and efficient HSIC architecture, and the differential-attention idea is worth investigating: the ablation in Figure 6 suggests that DMHSA consistently outperforms standard MHSA and MHCA, which is an interesting empirical finding. The paper also provides a complexity analysis and compares against a diverse set of recent baselines, including Mamba-based models. However, the central mechanism is not fully specified, the evaluation protocol is under-described, and several reported numbers are internally inconsistent. The absence of code, error bars, or a precise description of the train/test patch construction means the claimed superiority cannot currently be verified. The paper does not ship machine-checked proofs or a reproducible artifact, so the assessment rests entirely on the textual description and tables.
major comments (5)
- [Section II, Eqs. (5)-(7)] The differential attention computation is not fully specified and appears dimensionally inconsistent. With Q and K of shape (N, d_head), the score matrix S = QK^T / sqrt(d_head) has shape (N, N). Equation (6) then produces Sdiff = S[:, 1:] - S[:, :-1], which has shape (N, N-1). Equation (7) then writes Z = A V without defining A. If A is intended to be softmax(Sdiff), the token dimension of the attention weights no longer matches the token dimension of V (shape (N, d_head)); if A is something else, it must be defined explicitly. As written, the core mechanism cannot be implemented, and the claimed benefits of DMHSA cannot be evaluated.
- [Section III and Section IX] The train/validation/test split is not described at the pixel or patch level. The text states only that the dataset is partitioned into 25% training, 25% validation, and 50% testing, and that patches of size 8 (Section III) or 12x12 (Section IX) are used. It never states whether patches are extracted before or after the pixel-level split, nor whether overlapping patches are permitted. If overlapping patches are extracted from the full image and then randomly assigned to the train and test sets, the same or neighboring pixels can appear in both sets, which would directly inflate all accuracy figures in Tables III-VI. This concern is load-bearing for every comparative result in the paper, and the authors must specify the exact split order and patch-overlap policy, and ideally release the exact train/test masks.
- [Table III] The reported AA of 99.4136% is inconsistent with the per-class accuracies in the same table. Averaging the 16 per-class values for DiffFormer (Strawberry through Water) gives approximately 98.87%, not 99.41%, a discrepancy of about 0.54 percentage points. Similarly, Table V reports a kappa coefficient of 1477.32 for MHMamba, which is impossible because kappa is bounded above by 1. These inconsistencies must be corrected and all aggregate metrics recomputed from the per-class values.
- [Sections IV-VII and IX] Hyperparameters appear to be selected using test-set metrics, which biases the reported results. Table II reports OA, AA, and kappa for seven patch sizes on all four datasets, and the prose in Section IV identifies the best patch sizes (e.g., 18x18 for HC, 20x20 for UH), yet Section IX states that a 12x12 patch is used uniformly in the final comparisons. According to Table II, 12x12 is not the best patch size for any dataset. If the final configuration was chosen after inspecting test-set performance, the claimed SOTA numbers are optimistically biased. Model selection should be performed on the validation split, with test metrics reported only for the final model.
- [All of Section IX] All reported results appear to come from a single run, without error bars or statistical significance tests. The claimed improvements over the closest competitors are often small (e.g., OA 99.6229% vs. 99.3878% on UH in Table IV), so the absence of variance estimates makes it impossible to tell whether the differences are meaningful. The code is not provided for review (the abstract states it will be released only after revision), so the exact split, hyperparameter search, and implementation details cannot be audited.
minor comments (5)
- [Section II, Eq. (8)] The formula SWiGLU(x,g) = x * sigmoid(g) + x does not match the usual definition of SwiGLU, which is typically based on a Swish-gated linear unit; please clarify the intended activation and cite the original source precisely.
- [Section II, after Eq. (2)] The word 'presentec' should be 'presented'.
- [Section IV, PU paragraph] The text says the 18x18 patch gives the best PU results with kappa=97.26, OA=97.66, AA=95.65, but those values correspond to the HC dataset in Table II; for PU, Table II lists kappa=97.98, OA=98.48, AA=97.65 for 18x18, and the 14x14 patch gives even higher values. The discussion should be corrected to match the table.
- [Figure 5] The three-dimensional plot showing four datasets with different markers is difficult to read; separate two-dimensional plots or a table would convey the dependence on attention heads more clearly.
- [Section X] The final paragraph contains a run-on sentence that lacks punctuation; please revise it for readability.
Circularity Check
No significant circularity: DiffFormer's accuracy claims rest on benchmark comparisons, not on a derivation that reduces to its inputs.
full rationale
This is an empirical architecture paper, not a derivation, and I find no step in which a predicted quantity is defined in terms of its own target or in which a fitted parameter is renamed as a prediction. The core mechanism, DMHSA, is specified directly in Section II (Eqs. 5-7): S = QK^T / sqrt(d_head), followed by S_diff = S[:, 1:] - S[:, :-1]; this is a fixed architectural operation on the input Q, K, V, not a quantity fitted from classification accuracy. Similarly, the SWiGLU expression (Eq. 8) and the class-token classification head (Eq. 9) are standard components with no hidden dependence on the reported OA, AA, or kappa values. The stated superiority claim in Section IX (Tables III-VI) is supported by benchmark runs against published baselines; several baselines, such as Former [30], WaveFormer [32], WaveMamba [35], PyFormer [40], and MHSSMamba [42], are prior works by overlapping authors, but the comparison is empirical rather than a theorem whose validity rests on those citations. A self-citation becomes circular only when the load-bearing argument reduces to the unverified citation itself, which is not the case here. The manuscript's main weakness is an unreported experimental-protocol detail: Section III states only that the dataset is split 25%/25%/50% and that patches of spatial size 8 (later 12x12) are used, without saying whether patches are extracted before or after the split or whether overlapping patches are allowed. If overlapping patches straddle the train/test partition, the reported accuracies could be inflated by spatial leakage. That is a data-contamination and soundness concern, not a circularity concern, because a clean implementation would still yield a self-contained empirical result. No self-definitional equation, no fitted-input-called-prediction, no imported uniqueness theorem, and no ansatz smuggled in via citation were found. Accordingly, the circularity score is low, reflecting only the presence of minor self-citations that are not load-bearing.
Assumptions & free parameters
free parameters (9)
- Number of PCA components =
15
- Input patch size =
12x12 for comparative tables
- Transformer layers =
4
- Attention heads =
8
- Dropout rate =
0.1
- L2 penalty =
0.01
- Learning rate and decay =
0.001, 1e-6
- Batch size =
56
- Number of epochs =
50
assumptions (3)
- domain assumption Patch-based random splitting yields independent train and test samples.
- domain assumption Self-attention over 3D-conv patches captures spatial-spectral dependencies relevant to HSIC.
- standard math A normalization such as softmax turns raw attention scores into weights.
Cite this review
Pith. "Pith review of DiffFormer: a Differential Spatial-Spectral Transformer for Hyperspectral Image Classification." pith.science (2026). https://pith.science/paper/NR4PD6C4
@misc{pith2026241217350,
author = {Pith},
title = {Pith review of: DiffFormer: a Differential Spatial-Spectral Transformer for Hyperspectral Image Classification},
year = {2026},
howpublished = {\url{https://pith.science/paper/NR4PD6C4}},
note = {Machine review of arXiv:2412.17350}
}
read the original abstract
Hyperspectral image classification (HSIC) has gained significant attention because of its potential in analyzing high-dimensional data with rich spectral and spatial information. In this work, we propose the Differential Spatial-Spectral Transformer (DiffFormer), a novel framework designed to address the inherent challenges of HSIC, such as spectral redundancy and spatial discontinuity. The DiffFormer leverages a Differential Multi-Head Self-Attention (DMHSA) mechanism, which enhances local feature discrimination by introducing differential attention to accentuate subtle variations across neighboring spectral-spatial patches. The architecture integrates Spectral-Spatial Tokenization through three-dimensional (3D) convolution-based patch embeddings, positional encoding, and a stack of transformer layers equipped with the SWiGLU activation function for efficient feature extraction (SwiGLU is a variant of the Gated Linear Unit (GLU) activation function). A token-based classification head further ensures robust representation learning, enabling precise labeling of hyperspectral pixels. Extensive experiments on benchmark hyperspectral datasets demonstrate the superiority of DiffFormer in terms of classification accuracy, computational efficiency, and generalizability, compared to existing state-of-the-art (SOTA) methods. In addition, this work provides a detailed analysis of computational complexity, showcasing the scalability of the model for large-scale remote sensing applications. The source code will be made available at \url{https://github.com/mahmad000/DiffFormer} after the first round of revision.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
R. N. Patro, S. Subudhi, P. K. Biswal, and F. Dell’acqua, “A review of unsupervised band selection techniques: Land cover classification for hyperspectral earth observation data,” IEEE Geoscience and Remote Sensing Magazine, vol. 9, no. 3, pp. 72–111, 2021
work page 2021
-
[2]
Hyperspectral image clustering: Current achievements and future lines,
H. Zhai, H. Zhang, P. Li, and L. Zhang, “Hyperspectral image clustering: Current achievements and future lines,” IEEE Geoscience and Remote Sensing Magazine, vol. 9, no. 4, pp. 35–67, 2021
work page 2021
-
[3]
Hyperspectral imaging in environmental monitoring and analysis,
R. Rajabi, A. Zehtabian, K. D. Singh, A. Tabatabaeenejad, P. Ghamisi, and S. Homayouni, “Hyperspectral imaging in environmental monitoring and analysis,” Frontiers in Environmental Science, vol. 11, p. 1353447, 2024
work page 2024
-
[4]
Modern trends in hyperspectral image analysis: A review,
M. J. Khan, H. S. Khan, A. Yousaf, K. Khurshid, and A. Abbas, “Modern trends in hyperspectral image analysis: A review,” Ieee Access, vol. 6, pp. 14 118–14 129, 2018
work page 2018
-
[5]
S. Peyghambari and Y . Zhang, “Hyperspectral remote sensing in litho- logical mapping, mineral exploration, and environmental geology: an updated review,” Journal of Applied Remote Sensing , vol. 15, no. 3, pp. 031 501–031 501, 2021
work page 2021
-
[6]
M. Imani and H. Ghassemian, “An overview on spectral and spatial information fusion for hyperspectral image classification: Current trends and challenges,” Information fusion, vol. 59, pp. 59–83, 2020
work page 2020
-
[7]
Hyperspectral remote sensing data analysis and future challenges,
J. M. Bioucas-Dias, A. Plaza, G. Camps-Valls, P. Scheunders, N. Nasrabadi, and J. Chanussot, “Hyperspectral remote sensing data analysis and future challenges,” IEEE Geoscience and remote sensing magazine, vol. 1, no. 2, pp. 6–36, 2013
work page 2013
-
[8]
Future perspectives and challenges in hyperspectral remote sensing,
P. C. Pandey, H. Balzter, P. K. Srivastava, G. P. Petropoulos, and B. Bhattacharya, “Future perspectives and challenges in hyperspectral remote sensing,” Hyperspectral Remote Sensing , pp. 429–439, 2020
work page 2020
Show all 42 references
-
[9]
Multiple Sub-Pixel Target Detection for Hyperspectral Imaging Systems,
P. Addabbo, N. Fiscante, G. Giunta, D. Orlando, G. Ricci, and S. L. Ullo, “Multiple Sub-Pixel Target Detection for Hyperspectral Imaging Systems,” IEEE Transactions on Signal Processing , vol. 71, pp. 1599– 1611, 2023
2023
-
[10]
Hyperspectral image clas- sification—Traditional to deep models: A survey for future prospects,
M. Ahmad, S. Shabbir, S. K. Roy, D. Hong, X. Wu, J. Yao, A. M. Khan, M. Mazzara, S. Distefano, and J. Chanussot, “Hyperspectral image clas- sification—Traditional to deep models: A survey for future prospects,” IEEE Journal of Selected Topics in Applied Earth Observations and ...
2021
-
[11]
Hyperspectral Image Classification Using Groupwise Separable Convolutional Vision Transformer Net- work,
Z. Zhao, X. Xu, S. Li, and A. Plaza, “Hyperspectral Image Classification Using Groupwise Separable Convolutional Vision Transformer Net- work,” IEEE Transactions on Geoscience and Remote Sensing , vol. 62, pp. 1–17, 2024
2024
-
[12]
MS2I2Former: Multiscale Spa- tial–Spectral Information Interactive Transformer for Hyperspectral Im- age Classification,
S. Cheng, R. Chan, and A. Du, “MS2I2Former: Multiscale Spa- tial–Spectral Information Interactive Transformer for Hyperspectral Im- age Classification,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–19, 2024
2024
-
[13]
Bridging CNN and Transformer With Cross-Attention Fusion Network for Hyperspectral Image Classification,
F. Xu, S. Mei, G. Zhang, N. Wang, and Q. Du, “Bridging CNN and Transformer With Cross-Attention Fusion Network for Hyperspectral Image Classification,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–14, 2024
2024
-
[14]
A center-masked transformer for hyperspectral image classification,
S. Jia, Y . Wang, S. Jiang, and R. He, “A center-masked transformer for hyperspectral image classification,” IEEE Transactions on Geoscience and Remote Sensing , vol. 62, pp. 1–16, 2024
2024
-
[15]
Hierarchical attention transformer for hyper- spectral image classification,
T. Arshad and J. Zhang, “Hierarchical attention transformer for hyper- spectral image classification,” IEEE Geoscience and Remote Sensing Letters, vol. 21, pp. 1–5, 2024
2024
-
[16]
GraphGST: Graph Generative Structure-Aware Transformer for Hyper- spectral Image Classification,
M. Jiang, Y . Su, L. Gao, A. Plaza, X.-L. Zhao, X. Sun, and G. Liu, “GraphGST: Graph Generative Structure-Aware Transformer for Hyper- spectral Image Classification,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–16, 2024
2024
-
[17]
Foundation Model-Based Spec- tral–Spatial Transformer for Hyperspectral Image Classification,
L. Huang, Y . Chen, and X. He, “Foundation Model-Based Spec- tral–Spatial Transformer for Hyperspectral Image Classification,” IEEE Transactions on Geoscience and Remote Sensing , vol. 62, pp. 1–25, 2024
2024
-
[18]
Dual attention transformer network for hyperspectral image classification,
Z. Shu, Y . Wang, and Z. Yu, “Dual attention transformer network for hyperspectral image classification,” Engineering Applications of Artificial Intelligence , vol. 127, p. 107351, 2024. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S095219762301535X
2024
-
[19]
Spectral–Spatial Trans- former Network for Hyperspectral Image Classification: A Factorized Architecture Search Framework,
Z. Zhong, Y . Li, L. Ma, J. Li, and W.-S. Zheng, “Spectral–Spatial Trans- former Network for Hyperspectral Image Classification: A Factorized Architecture Search Framework,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–15, 2022
2022
-
[20]
QTN: Quaternion Transformer Network for Hyperspectral Image Classification,
X. Yang, W. Cao, Y . Lu, and Y . Zhou, “QTN: Quaternion Transformer Network for Hyperspectral Image Classification,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 33, no. 12, pp. 7370– 7384, 2023
2023
-
[21]
A Lightweight Transformer Network for Hyperspectral Image Classifica- tion,
X. Zhang, Y . Su, L. Gao, L. Bruzzone, X. Gu, and Q. Tian, “A Lightweight Transformer Network for Hyperspectral Image Classifica- tion,” IEEE Transactions on Geoscience and Remote Sensing , vol. 61, pp. 1–17, 2023
2023
-
[22]
Hyperspectral Image Transformer Classification Networks,
X. Yang, W. Cao, Y . Lu, and Y . Zhou, “Hyperspectral Image Transformer Classification Networks,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–15, 2022
2022
-
[23]
MSTNet: A Multilevel Spectral–Spatial Transformer Network for Hyperspectral Image Classification,
H. Yu, Z. Xu, K. Zheng, D. Hong, H. Yang, and M. Song, “MSTNet: A Multilevel Spectral–Spatial Transformer Network for Hyperspectral Image Classification,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–13, 2022
2022
-
[24]
MATNet: A Combining Multi-Attention and Transformer Network for Hyperspectral Image Classification,
B. Zhang, Y . Chen, Y . Rong, S. Xiong, and X. Lu, “MATNet: A Combining Multi-Attention and Transformer Network for Hyperspectral Image Classification,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1–15, 2023
2023
-
[25]
Differential Transformer,
T. Ye, L. Dong, Y . Xia, Y . Sun, Y . Zhu, G. Huang, and F. Wei, “Differential Transformer,” 2024. [Online]. Available: https: //arxiv.org/abs/2410.05258
2024 arXiv
-
[26]
Adaptive Graph Modeling With Self-Training for Heterogeneous Cross-Scene Hyperspectral Image Classification,
M. Ye, J. Chen, F. Xiong, and Y . Qian, “Adaptive Graph Modeling With Self-Training for Heterogeneous Cross-Scene Hyperspectral Image Classification,” IEEE Transactions on Geoscience and Remote Sensing , vol. 62, pp. 1–15, 2024
2024
-
[27]
Fusing Transformers in a Tuning Fork Structure for Hyperspectral Image Classification Across Disjoint Samples,
M. Ahmad, M. Usama, M. Mazzara, S. Distefano, H. A. Altuwaijri, and S. L. Ullo, “Fusing Transformers in a Tuning Fork Structure for Hyperspectral Image Classification Across Disjoint Samples,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vo...
2024
-
[28]
MIMO-SST: Multi-Input Multi-Output Spatial-Spectral Transformer for Hyperspectral and Multi- spectral Image Fusion,
J. Fang, J. Yang, A. Khader, and L. Xiao, “MIMO-SST: Multi-Input Multi-Output Spatial-Spectral Transformer for Hyperspectral and Multi- spectral Image Fusion,” IEEE Transactions on Geoscience and Remote Sensing, pp. 1–1, 2024
2024
-
[29]
Sparse self- attention transformer for image inpainting,
W. Huang, Y . Deng, S. Hui, Y . Wu, S. Zhou, and J. Wang, “Sparse self- attention transformer for image inpainting,” Pattern Recognition, vol. 145, p. 109897, 2024
2024
-
[30]
Spatial–Spectral Transformer With Conditional Position Encoding for Hyperspectral Image Classification,
M. Ahmad, M. Usama, A. M. Khan, S. Distefano, H. A. Altuwaijri, and M. Mazzara, “Spatial–Spectral Transformer With Conditional Position Encoding for Hyperspectral Image Classification,” IEEE Geoscience and Remote Sensing Letters , vol. 21, pp. 1–5, 2024
2024
-
[31]
Image fusion for the novelty rotating synthetic aperture system based on vision transformer,
Y . Sun, X. Zhi, S. Jiang, G. Fan, X. Yan, and W. Zhang, “Image fusion for the novelty rotating synthetic aperture system based on vision transformer,” Information Fusion, vol. 104, p. 102163, 2024
2024
-
[32]
WaveFormer: Spectral–Spatial Wavelet Transformer for Hyperspectral Image Classifi- cation,
M. Ahmad, U. Ghous, M. Usama, and M. Mazzara, “WaveFormer: Spectral–Spatial Wavelet Transformer for Hyperspectral Image Classifi- cation,” IEEE Geoscience and Remote Sensing Letters , vol. 21, pp. 1–5, 2024
2024
-
[33]
Deep Transformer Based Video Inpainting Using Fast Fourier Tokenization,
T. Kim, J. Kim, H. Oh, and J. Kang, “Deep Transformer Based Video Inpainting Using Fast Fourier Tokenization,” IEEE Access, vol. 12, pp. 21 723–21 736, 2024
2024
-
[34]
A Dual-Feature-Based Adaptive Shared Transformer Network for Image Captioning,
Y . Shi, J. Xia, M. Zhou, and Z. Cao, “A Dual-Feature-Based Adaptive Shared Transformer Network for Image Captioning,” IEEE Transactions on Instrumentation and Measurement , vol. 73, pp. 1–13, 2024
2024
-
[35]
WaveMamba: Spatial-Spectral Wavelet Mamba for Hyperspectral Image Classifica- tion,
M. Ahmad, M. Usama, M. Mazzara, and S. Distefano, “WaveMamba: Spatial-Spectral Wavelet Mamba for Hyperspectral Image Classifica- tion,” IEEE Geoscience and Remote Sensing Letters , vol. 22, pp. 1–5, 2025
2025
-
[36]
Swish: a self-gated activation function,
P. Ramachandran, B. Zoph, and Q. Le, “Swish: a self-gated activation function,” 10 2017. JOURNAL OF LATEX CLASS FILES 13
2017
-
[37]
Mini-UA V-borne hyperspectral remote sensing: From observation and processing to applications,
Y . Zhong, X. Wang, Y . Xu, S. Wang, T. Jia, X. Hu, J. Zhao, L. Wei, and L. Zhang, “Mini-UA V-borne hyperspectral remote sensing: From observation and processing to applications,” IEEE Geoscience and Remote Sensing Magazine , vol. 6, no. 4, pp. 46–62, 2018
2018
-
[38]
Hyperspectral and lidar data fusion: Outcome of the 2013 grss data fusion contest,
C. Debes, A. Merentitis, R. Heremans, J. Hahn, N. Frangiadakis, T. van Kasteren, W. Liao, R. Bellens, A. Pi ˇzurica, S. Gautama et al. , “Hyperspectral and lidar data fusion: Outcome of the 2013 grss data fusion contest,” IEEE Journal of Selected Topics in Applied Earth Observ...
2013
-
[39]
Attention Graph Convolutional Network for Disjoint Hyperspectral Image Classification,
A. Jamali, S. K. Roy, D. Hong, P. M. Atkinson, and P. Ghamisi, “Attention Graph Convolutional Network for Disjoint Hyperspectral Image Classification,” IEEE Geoscience and Remote Sensing Letters , vol. 21, pp. 1–5, 2024
2024
-
[40]
Pyramid Hierarchical Spatial-Spectral Transformer for Hyperspectral Image Classification,
M. Ahmad, M. H. F. Butt, M. Mazzara, S. Distefano, A. M. Khan, and H. A. Altuwaijri, “Pyramid Hierarchical Spatial-Spectral Transformer for Hyperspectral Image Classification,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing , vol. 17, pp. 17 6...
2024
-
[41]
A hybrid convolution transformer for hyperspectral image classification,
J. Z. Tahir Arshad and I. Ullah, “A hybrid convolution transformer for hyperspectral image classification,” European Journal of Remote Sensing, vol. 0, no. 0, p. 2330979, 2024
2024
-
[42]
Multi-head Spatial-Spectral Mamba for Hyperspectral Image Classification,
M. Ahmad, M. H. F. Butt, M. Usama, H. A. Altuwaijri, M. Mazzara, and S. Distefano, “Multi-head Spatial-Spectral Mamba for Hyperspectral Image Classification,” 2024. [Online]. Available: https://arxiv.org/abs/ 2408.01224
2024 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.