{"id":"6230dd02-ec85-463e-9506-8ada72675d24","arxiv_id":"2504.15612","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"HS-Mamba fuses non-overlapping local Mamba patches with whole-image lightweight attention and reports state-of-the-art accuracy on four hyperspectral classification datasets.","lead":"A new neural network, HS-Mamba, combines patch-based Mamba blocks with a whole-image attention branch for classifying hyperspectral remote sensing images. The authors report top accuracy on four benchmarks, but do not release code and include an internal inconsistency in their efficiency claims.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Data-count contradiction: Section V-A states HanChuan has 123,797 and HongHu ~120,000 labeled samples, yet Table I lists test sets of 256,890 and 385,813; the HC/HH SOTA claims rest on an unverified split.","rationale":"The reader's verdict is reasonable, and the unidirectional-scanning premise in Section IV-B1 is indeed unsupported: the paper asserts 'the unidirectional scanning strategy becomes the best choice' with no controlled ablation, and Table VI only ablates domain branches, not scan direction. However, that gap concerns architectural justification, not the truth of the reported SOTA numbers. A model does not cease to achieve 94.65% OA on IP merely because its scan direction is not proven optimal. The more load-bearing concern is empirical reproducibility: the manuscript gives irreconcilable dataset counts for two of the four benchmarks. Section V-A3 says HanChuan has 123,797 labeled samples, while Table I's test column alone sums to 256,890 (total 257,530 with train/val); Section V-A4 says HongHu has ~120,000 labeled samples, while Table I sums to 386,693. These are not stylistic issues; they determine whether the HC and HH rows of Tables IV-V can exist. A concrete check against the official WHU-Hi ground truths settles the matter without code. Until that reconciliation is public, the central 'SOTA on four benchmarks' claim is unverdictable, although the IP and PU results appear plausible and the 10-run protocol with mean plus/minus std is a plus.","tokens_in":22004,"tokens_out":15474,"duration_ms":148266,"concrete_test":"Download the official WHU-Hi-HanChuan and WHU-Hi-HongHu ground-truth maps from the WHU-Hi dataset page and count labeled pixels per class. Compare with Table I's per-class train/val/test columns. If HanChuan's official total is 123,797 (and HongHu's is ~120,000), then Table I's test totals of 256,890 and 385,813 are impossible and the HC/HH SOTA claims are invalid. If the official totals instead match 257,530 and 386,693, the prose counts are typos and the table may be correct; then rerun the HS-Mamba HC/HH experiments with the exact Table I splits to confirm the reported OA.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The central four-dataset SOTA claim depends on the integrity of the experimental splits, and the manuscript contradicts itself on two of the four datasets. Section V-A3 (WHU-Hi-HanChuan) says 'The dataset includes 123797 annotated samples'; Table I's HanChuan columns sum to 480 train + 160 val + 256,890 test = 257,530 total, about 2.08x the stated count. Section V-A4 says HongHu has 'approximately 120,000 labeled samples'; Table I sums to 660 + 220 + 385,813 = 386,693. Unless the official ground truths contain 257,530 and 386,693 labeled pixels (in which case the prose is wrong), the reported test sets exceed the entire labeled dataset, so the HC and HH rows of Tables IV-V cannot be reproduced from the described data. Because the strongest claim explicitly extends to all four benchmarks, an unresolved contradiction in the two largest datasets is more load-bearing than the unidirectional-scanning design premise: it puts the empirical basis of half the headline results in doubt. Resolving this does not require model code or assumptions about scanning direction, only the public ground-truth maps.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes HS-Mamba, a hyperspectral image classification architecture that combines a dual-channel spatial-spectral encoder (DCSS-Encoder) processing non-overlapping patches with multi-group Mamba blocks and a lightweight global inline attention (LGI-Att) branch processing the full image. The two branches are fused with a learned gated fusion and a three-stage hierarchical up/down-sampling scheme. The authors evaluate on Indian Pines, Pavia University, WHU-Hi-HanChuan, and WHU-Hi-HongHu, reporting OA/AA/Kappa over 10 runs against eight prior methods, and claim state-of-the-art results with margins over MambaHSI of 0.96-3.86% OA. Ablations examine dual-domain design, positional encoding, the LGI-Att branch, fusion strategy, patch size, and number of Mamba groups.","tokens_in":22315,"tokens_out":10818,"duration_ms":89419,"significance":"The local/global fusion idea is a sensible middle ground between pixel-patch and whole-image strategies, and the multi-group Mamba design with grouped adaptive weighting is interesting. Strengths include the 10-run mean±std protocol, ablation coverage of the main components, and comparison against eight baselines on four standard datasets. If the SOTA margins survive a correctly specified split and statistical testing, HS-Mamba would be a competitive contribution to Mamba-based HSI classification. The main claims as stated are not fully supported by the manuscript's own data: the dataset-size contradictions and the contradicted efficiency claim prevent the paper from being accepted in current form.","major_comments":[{"comment":"Section V-A3 states that WHU-Hi-HanChuan includes 123,797 annotated samples, but Table I sums to 480 train + 160 val + 256,890 test = 257,530, disagreeing by more than a factor of two. Section V-A4 states that WHU-Hi-HongHu has approximately 120,000 labeled samples, but Table I sums to 660 + 220 + 385,813 = 386,693. Since Tables IV and V report the headline SOTA results on these two datasets, the authors must state which numbers are correct and confirm that the train/val/test splits are exactly derived from the public ground truth; as written, the test sets on HC and HH exceed the label budgets described in the prose.","section":"V-A3, V-A4, Table I"},{"comment":"The claimed computational superiority in Section V-F is contradicted by Table X. The text says HS-Mamba achieves '52% faster inference and 26% lower FLOPs than MambaHSI on medium-high resolution datasets,' but the reported inference times are 0.11 vs 0.03 s (IP), 0.50 vs 0.35 s (PU), 0.85 vs 0.65 s (HC), and only 0.65 vs 0.77 s on HH. Thus HS-Mamba is slower than MambaHSI on three of four datasets, and the efficiency conclusion in Section VI ('maintaining high computational efficiency') is not supported by the paper's own table.","section":"V-F, Table X"},{"comment":"No significance testing is reported for the SOTA margins. Several claimed improvements are smaller than the reported standard deviations: on PU, HS-Mamba OA is 96.43±1.35 versus MambaHSI 95.47±0.84 (Table III), and on IP the OA margin is 1.8 points with stds of 0.91 and 0.81 (Table II). Without paired significance tests or confidence intervals, the claim of 'best performance' is not statistically established for the smaller margins, and the qualitative statement in Section V-C should be qualified accordingly.","section":"V-C, Tables II-V"},{"comment":"The hyperparameter analysis in Section V-E selects patch size and the number of Mamba groups per dataset (e.g., groups = 16 for IP, PU, HC and 8 for HH), but the paper does not state whether these curves were computed on the validation or test set, and Section V-B3 lists neither the default patch size nor the final M/N values used to produce Tables II-V. If the test set was used for model selection, the reported results are not held-out and the SOTA margins are inflated; the authors must state the selection procedure and the final hyperparameter values for full reproducibility.","section":"V-E, V-B3"},{"comment":"Section IV-B1 asserts that 'the unidirectional scanning strategy becomes the best choice' for HSI sequences and justifies it by citing redundancy and computational inefficiency of multi-directional scanning, but the ablation study in Section V-D contains no comparison of unidirectional against bidirectional or multi-directional scanning within the DCSS-Encoder. Since every local feature in the encoder is read in a single direction, this unvalidated premise is load-bearing: either add a controlled scanning-direction ablation or state the claim as an assumption rather than a finding.","section":"IV-B1, V-D"},{"comment":"The split protocol stated in Section V-B3 ('30 pixel samples are allocated for training, 10 for validation, and the remaining for testing') is not followed for the small classes in Table I: Indian Pines class 7 has 15 train + 5 val + 8 test = 28 samples, and class 9 has 10 + 5 + 5 = 20 samples. The manuscript must specify the actual rule for classes with fewer than 40 labeled pixels; otherwise the experimental setup is not reproducible.","section":"V-B3, Table I"}],"minor_comments":[{"comment":"The mathematical notation is incomplete: L is introduced as the patch count, then reused as sequence length; the split operations in Eq. (8) use U and V without definition, and the sequence shapes in Eq. (7) do not match the later grouping dimensions. Please provide precise tensor shapes for the scanning and splitting steps.","section":"IV-B1, Eqs. (7)-(8)"},{"comment":"The axis labels in Figure 9 are corrupted by glyph-escape text ('/uni0000001a/uni0000001c/...'), making the patch-size and group-count curves unreadable; the figure must be regenerated with normal text labels.","section":"Fig. 9"},{"comment":"The same configuration (DCSS-Encoder without LGI-Attention) is reported with HC OA 95.93 in Table VI and 95.73 in Table VIII; please reconcile these numbers or state the differences in configuration.","section":"Tables VI and VIII"},{"comment":"The claim that LGI-Attention reduces standard deviations by approximately 0.3% is not supported by Table VIII for HC, where the std increases from 0.74 to 0.79; the summary sentence should be revised.","section":"V-D3"},{"comment":"The sentence 'pixel-patch methods achieve 36% faster training times' is not derived from Table X and is too vague; give the calculation or remove the statistic.","section":"V-F"},{"comment":"There are several typos: 'an full-field' in the abstract, 'MorpyFormer' in Figure 5 caption, and the incomplete sentence 'The division details for training, validation and test sets are provided in.' in Section V-A.","section":"Abstract and figure captions"},{"comment":"The header 'Abbreviated Category Colors, Names and Sample Numbers' mentions colors, but the table lists names and numbers only; clarify or remove 'Colors'.","section":"Table I"}],"recommendation":"major_revision","confidential_remarks":"The manuscript would benefit from a strict round of reproducibility checking: the dataset-count errors and the contradicted efficiency claim will be picked up by any reviewer. No code is included, so the accuracy tables cannot be independently checked against the corrected splits once the authors specify them."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"X,\n\nThe paper is an incremental but reasonable engineering contribution: it combines non-overlapping multi-group Mamba patches with a whole-image lightweight attention branch and gated fusion. The experimental protocol is better than most (10 runs, mean ± std, ablations, hyperparameter analysis). But I wouldn't let the SOTA claims stand as-is. There is a data-count contradiction that needs to be resolved before the HC results can be trusted.\n\nSection V-A3 says HanChuan has 123,797 annotated samples. Table I lists 480 train, 160 val, and 256,890 test pixels for that dataset, total 257,530. That's more than twice the stated number of labeled samples. Unless the official ground truth has changed, the test set alone exceeds the entire labeled pool. The HongHu section says approximately 120,000 labeled samples, but Table I sums to 386,693, which is likely the official count; the prose is simply wrong there. One of these is a simple typo, but the HanChuan discrepancy cannot be a typo—it means the HC numbers in Tables IV and V are not reproducible from the described data. That is load-bearing: the headline SOTA claim covers all four datasets.\n\nThe efficiency claim in Section V-F is also contradicted by its own Table X: HS-Mamba is slower than MambaHSI at inference on three of four datasets (0.11 vs 0.03 on IP, 0.50 vs 0.35 on PU, 0.85 vs 0.65 on HC). The paper claims 52% faster inference. The FLOPs are lower, but the inference time data don't support the stated claim.\n\nSmaller issues: the default patch size and number of Mamba groups are never stated in the implementation details, only shown in hyperparameter figures. No significance tests accompany the 10-run means. And the unidirectional scanning choice is asserted without a controlled comparison, though I think that's a secondary concern.\n\nFor a reader in remote sensing, this is a useful comparison point—if the data issue gets fixed. The architecture itself is fine and the accuracy gains over MambaHSI are plausible. But the contradiction on HanChuan means I would not cite the HC/HH numbers until the authors provide the actual split or correct the dataset description.\n\nRecommendation: send it to peer review, but with a clear request to resolve the data-count contradiction and the efficiency claim. It deserves referee time; it just needs minor-to-moderate revision.","headline":"Solid incremental architecture with a serious data-count inconsistency on the HanChuan benchmark that must be fixed before the SOTA numbers are credible.","tokens_in":22804,"tokens_out":6167,"would_cite":false,"duration_ms":49415,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"HS-Mamba fuses local patches with whole-image attention and reports state-of-the-art accuracy on four hyperspectral benchmarks.","keywords":["hyperspectral image classification","Mamba","selective state space model","full-field interaction","dual-channel spatial-spectral encoder","global attention","remote sensing","pixel-level classification"],"falsifier":"Run HS-Mamba's DCSS-Encoder with the same hyperparameters but with bidirectional or four-way (cross-scan) scanning on Indian Pines, using the same 30-train/10-validation split; if the multi-directional version's mean OA exceeds 94.65% or closes the gap to MambaHSI by more than the reported margin, the claim that unidirectional scanning is the best choice would be refuted.","tokens_in":21825,"feed_emoji":"🛰️","tokens_out":4653,"duration_ms":35316,"temperature":0.7,"pith_summary":"The paper proposes HS-Mamba, a hyperspectral image classification framework that deliberately combines two previously separate strategies: it feeds non-overlapping local patches through multiple parallel Mamba state-space encoders, while simultaneously passing the entire image through a lightweight attention branch. The claim is that this full-field interaction design captures both fine local structure and global scene context, and that doing so yields state-of-the-art accuracy on four standard benchmarks: Indian Pines, Pavia University, WHU-Hi-HanChuan, and WHU-Hi-HongHu. A sympathetic reader would care because the method reports consistent gains over the previous best (MambaHSI) while also cutting computation, suggesting a practical recipe for high-resolution pixel-level classification.","feed_headline":"Dual-branch Mamba sets new records on four hyperspectral datasets","feed_subtitle":"HS-Mamba combines non-overlapping patches with whole-image attention, outpacing MambaHSI by up to 3.9% OA.","key_machinery":"The central object is the HS-Mamba block, composed of a DCSS-Encoder and a LGI-Att branch. The DCSS-Encoder flattens non-overlapping patches into spatial-priority and spectral-priority 1D sequences, splits them into groups, processes each group with a parallel S6 (selective state-space) block, and adaptively combines the groups with learned weights and an adaptive concatenation, adding cosine positional encoding to retain location. The LGI-Att branch applies a compressed attention on the spectral dimension and a dilated-convolution extended attention on the spatial dimension of the whole image. A gated fusion layer then weights the two branches.","core_discovery":"This paper establishes that a Mamba network can achieve the best published classification accuracy on four hyperspectral image benchmarks by fusing dual-domain local features and global attention. The framework, HS-Mamba, uses a dual-channel spatial-spectral encoder (DCSS-Encoder) to model non-overlapping patches with multi-group Mamba blocks and cosine positional encoding, and a lightweight global inline attention (LGI-Att) branch to capture whole-image spectral and spatial context. Gated fusion combines the two streams. Against eight state-of-the-art baselines, HS-Mamba reports OA/AA/Kappa of 94.65/96.86/93.87 on Indian Pines, 96.43/97.14/95.29 on Pavia University, 96.64/96.13/95.72 on HanChuan, and 96.10/96.00/95.08 on HongHu, surpassing the second-best MambaHSI on all four.","pith_inferences":["The unidirectional scanning decision in Section IV-B1 is not backed by a controlled comparison against bidirectional or four-way scanning, so the reported gains may depend partly on that choice; a test swapping the scanning direction could separate the contribution.","The largest margins appear on the wide-resolution HanChuan dataset (+3.86% OA), hinting that the fusion is most valuable where local detail and global layout both matter; this could be tested on other high-resolution remote sensing images.","Since the method's efficiency claims rely on a fixed train/validation split and specific GPU, reproducing the FLOPs and inference comparisons on other hardware and splits would test the practicality."],"forward_implications":["On the four tested benchmarks, fusing local patch modelling with whole-image global attention beats both pure pixel-patch and pure whole-image strategies, as HS-Mamba outscores MambaHSI, a whole-image approach, and all pixel-patch baselines.","The reported efficiency numbers, about 52% faster inference and 26% lower FLOPs than MambaHSI on medium-high resolution datasets, imply the full-field strategy is usable on large HSI scenes without GPU-memory overflow.","The ablation results indicate that gated fusion, rather than sum, adaptive sum, or concatenation, is the most robust way to combine the two branches across datasets.","Positional encoding contributes 2% to 7% OA across datasets, so explicit location information is load-bearing for non-overlapping patch scanning.","The design suggests a general template: local fine-grained sequence modeling plus lightweight global attention, which could extend to other dense prediction tasks.","The multi-group Mamba composition with adaptive concatenation preserves long-range dependency modelling at linear complexity, a direct consequence of the S6 design."],"supporting_citations":[{"why":"Supplies the S6 selective state-space block that powers every group in the multi-group Mamba.","marker":"[21]"},{"why":"The whole-image-based Mamba baseline that is the second-best method and the main comparison target.","marker":"[24]"},{"why":"Cited for the claim that multi-directional scanning introduces redundancy and inefficiency, motivating the unidirectional scanning choice.","marker":"[25]"},{"why":"Establishes how Mamba is adapted to image sequences through bidirectional scanning.","marker":"[22]"},{"why":"Provides the cross-scan vision Mamba context that informs the scanning design discussion.","marker":"[23]"},{"why":"A pixel-patch Mamba baseline that HS-Mamba must outperform in the comparisons.","marker":"[42]"},{"why":"The 3D-CNN baseline used in the comparative tables.","marker":"[27]"}],"fun_headline_variants":["HS-Mamba fuses local and global features to beat SOTA on four HSI sets","Mamba-based HS-Mamba tops four hyperspectral benchmarks with dual-stream fusion","Full-field Mamba wins on four HSI datasets with local-global attention fusion","HS-Mamba: local patches plus whole-image attention sets new hyperspectral records","Dual-stream Mamba surpasses all baselines on four hyperspectral datasets"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The design assumes that reading each local feature sequence in a single fixed direction is sufficient; no experiment in the paper compares this against bidirectional or multi-directional scanning, so if direction order discards important spectral-spatial correlations, the reported accuracy would drop.","fun_headline_variants_meta":{"raw":{"variants":["HS-Mamba fuses local and global features to beat SOTA on four HSI sets","Mamba-based HS-Mamba tops four hyperspectral benchmarks with dual-stream fusion","Full-field Mamba wins on four HSI datasets with local-global attention fusion","HS-Mamba: local patches plus whole-image attention sets new hyperspectral records","Dual-stream Mamba surpasses all baselines on four hyperspectral datasets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000288,"raw_usage":{"total_tokens":1739,"prompt_tokens":1042,"completion_tokens":697,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":658,"completion_tokens_details":{"reasoning_tokens":592}},"tokens_in":658,"tokens_out":697,"duration_ms":6041,"temperature":1.0,"reasoning_tokens":592,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:21:55.785620+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run HS-Mamba's DCSS-Encoder with the same hyperparameters but with bidirectional or four-way (cross-scan) scanning on Indian Pines, using the same 30-train/10-validation split; if the multi-directional version's mean OA exceeds 94.65% or closes the gap to MambaHSI by more than the reported margin, the claim that unidirectional scanning is the best choice would be refuted.","supporting_citations":[{"cited_title":"MambaHSI: Spatial–Spectral Mamba for Hyperspectral Image Classification","cited_arxiv_id":null,"evidence_quote":"The whole-image-based Mamba baseline that is the second-best method and the main comparison target."},{"cited_title":"DualMamba: A Lightweight Spec- tral–Spatial Mamba-Convolution Network for Hyper- spectral Image Classification","cited_arxiv_id":null,"evidence_quote":"Cited for the claim that multi-directional scanning introduces redundancy and inefficiency, motivating the unidirectional scanning choice."},{"cited_title":"3DSS-Mamba: 3D-Spectral-Spatial Mamba for Hyperspectral Image Classification","cited_arxiv_id":null,"evidence_quote":"A pixel-patch Mamba baseline that HS-Mamba must outperform in the comparisons."},{"cited_title":"HybridSN: Exploring 3- D–2-D CNN Feature Hierarchy for Hyperspectral Im- age Classification","cited_arxiv_id":null,"evidence_quote":"The 3D-CNN baseline used in the comparative tables."}],"review_version":1}