REVIEW 4 major objections 5 minor 10 references
Classification non supervis{\'e}es d'acquisitions hyperspectrales cod{\'e}es : quelles v{\'e}rit{\'e}s terrain ?
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Expert ground-truth labels for hyperspectral scenes group spectrally dissimilar pixels, making them a biased yardstick for unsupervised classification.
desk verdict Useful caution about Pavia-U ground truth, but the headline claim that reference labels contain classification errors is loaded: the evidence is two classes and a SAM-based comparison the method itself optimizes. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine of the analysis is a statistical model of intra-class spectral variability in which each pixel in a class is the same reference spectrum multiplied by a scale factor (a model of illumination variation), plus noise. The unsupervised classifier from the authors' earlier work tests, from coded DD-CASSI measurements, whether groups of pixels are consistent with sharing one underlying spectrum up to scale; a threshold T controls how tight the homogeneity must be, and lower T yields smaller, more coherent regions. For each detected region the median of the full pixel spectra is used as its representative spectrum, and the spectral angle mapper (SAM) between each pixel and that median quantifies residual intra-class variability. The coded data themselves are generated by a ray-tracing simulator of the DD-CASSI optical system, so the whole pipeline runs on realistic simulated acquisitions with ten times fewer measurements than the full cube.
What would settle it
Take the Pavia University Meadows class, compute each pixel's SAM to the median, and isolate the high-SAM sub-regions the method flags; if field or higher-resolution data show those sub-regions are distinct materials (different vegetation or surfaces), the ground truth is wrong and the paper's conclusion holds, whereas if they are all the same material under natural variation, the method is over-segmenting lawful variability.
Extended reading notes
Core claim
The paper's central discovery is that reference ground truth is a flawed yardstick for unsupervised classification: on Pavia University, the reference class Meadows contains several spectrally distinct materials, while Bitumen contains a few outlier spectra. When the authors' unsupervised classifier is applied directly to coded acquisitions compressed by a factor of ten, it detects regions whose SAM and RMSE distributions around the median spectrum are more concentrated near zero than the reference class distributions, meaning the detected regions are more spectrally homogeneous. The authors read this as evidence that the reference labels contain heterogeneity and labeling errors, and that their method better captures 'truly homogeneous' regions. The consequence is that numerical scores computed against such ground truths do not faithfully measure spectral coherence, so the evaluation of unsupervised methods needs to be redesigned.
Load-bearing premise
The load-bearing premise is that a correct class must be spectrally homogeneous, so that high spectral angle within a reference class counts as a labeling error; a material can, however, lawfully contain spectral variation from moisture, weathering, or mixing, and the spectral angle alone cannot tell lawful variability from mislabeling.
Editorial extensions
If this is right
- If ground truths overstate spectral homogeneity, then accuracy, precision, and confusion-matrix comparisons against them are systematically biased in favor of classifiers that reproduce the annotation's grouping.
- For scenes like Pavia University, a single reference class such as Meadows should be subdivided into several spectrally homogeneous sub-regions before it is used as an evaluation target.
- Unsupervised classification from coded acquisitions is viable at tenfold compression: the method identifies coherent classes and reference spectra without full cube reconstruction.
- The Bitumen analysis shows that reference labels can be partly salvaged by detecting and excluding a small number of aberrant spectra, rather than re-labeling the whole class.
- The class definition itself—what counts as one material versus a mixture or a spatially varying material—becomes a modeling choice that must be made explicit in any evaluation.
Reading between the lines
- An immediate testable extension is to build an alternative ground truth for Pavia University by subdividing reference classes at SAM thresholds, then re-running standard supervised and unsupervised benchmarks; if rankings change, the original labels were indeed biasing comparisons.
- The same SAM-coherence audit could be applied to datasets whose reference spectra were measured in situ (such as the CAMCATT campaign) to separate lawful material variability, like moisture or weathering, from annotation error—if in-situ spectra also scatter widely within a label, the 'error' is partly natural variability.
- The scale-factor variability model is a deliberate simplification; extending it to direction-dependent variability (e.g., mixtures or non-illumination effects) would let the method distinguish between 'same material, different lighting' and 'same label, different materials,' which the current SAM metric cannot do.
- If ground truths are biased, then paper-to-paper comparisons of unsupervised hyperspectral classifiers may be comparing how well each method mimics annotation artifacts; a community benchmark with multiple thresholded labels or intrinsic homogeneity scores would be a more honest yardstick.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents an unsupervised classification method for DD-CASSI coded-aperture hyperspectral data, based on a simple model of intra-class spectral variability with a homogeneity threshold T, and evaluates it on the Pavia University scene. Its central claim is that common ground truths such as Pavia University group spectrally dissimilar pixels into the same class and therefore do not faithfully reflect true spectral homogeneity, biasing the evaluation of unsupervised classification. The evidence consists of SAM and RMSE maps and histograms for two reference classes (Meadows and Bitumen), showing that the regions detected by the method have narrower intra-class SAM/RMSE distributions than the corresponding reference classes. The paper concludes that reference annotations contain classification errors and that unsupervised classification evaluation should be rethought.
Significance. The paper addresses a genuinely important issue: reference labels in hyperspectral benchmarks are not error-free, and their limitations are often ignored in unsupervised classification evaluation. The use of a realistic coded-acquisition simulator (SIMCA), the explicit comparison of median spectra, and the provision of spatial SAM maps are strengths. However, the central evidence is currently circular: the method is designed to produce spectrally homogeneous regions using a spectral-angle threshold, and the evaluation uses the same spectral-angle (SAM) criterion to show that detected regions are more coherent than the reference. The conclusion that the reference contains 'spectres aberrants' conflates high intra-class SAM with labeling error, without an independent ground-truth definition. If the authors can provide such independent validation, the paper would make a useful contribution; in its present form, it demonstrates that the method produces spectrally more coherent regions, not that the Pavia University ground truth is wrong.
major comments (4)
- [Section 4, Figs. 3-4] The evaluation uses the same metric (SAM) that the segmentation criterion T is designed to minimize. The observed decrease in intra-region SAM for detected regions is therefore partly by construction. To support the claim that the reference ground truth contains classification errors, the paper needs an independent reference—for example, in-situ spectral measurements (available for the CAMCATT dataset), known material samples, or a comparison of subregion spectra against spectral libraries—or at minimum a control experiment showing that the SAM reduction is larger than what would be obtained by any fine-grained segmentation of a physically homogeneous material.
- [Sections 3-4, Fig. 4] The quantitative evidence is limited to two classes (Meadows and Bitumen) and is purely descriptive. No statistical tests, confidence intervals, or accuracy measures are reported for the histograms in Fig. 4, and the claim that the reference is less homogeneous rests on visual inspection. The authors should report, for each class and threshold, the number of pixels, the median and quantiles of SAM and RMSE, and a statistical comparison (e.g., a bootstrap test or a two-sample Kolmogorov-Smirnov test) between reference-class and detected-region distributions.
- [Section 4, threshold T] The value T ≈ 0.09 is selected after observing that the SAM distributions become narrowest near that threshold, and the widening for smaller T is attributed to estimation difficulties on small regions. This post hoc selection, combined with the uncontrolled effect of region size, makes it difficult to attribute the observed coherence to the method's model rather than to the chosen threshold and to varying pixel counts. The comparison should be repeated with region-size-matched samples or with a principled criterion for selecting T, such as a stability or model-selection criterion.
- [Section 1 vs. Section 4] The paper acknowledges that intra-class spectral variability can lawfully arise from illumination differences, within-material spatial variation, and mixing, which SAM cannot distinguish from mislabeling. Yet the conclusion in Section 4 interprets high SAM in the Meadows reference class as evidence of 'spectres dissemblables, voire des spectres aberrants'. This inference is not justified unless the authors can show that the high-SAM pixels are inconsistent with physical within-class variability. A concrete test would be to examine the spatial distribution of high-SAM pixels and test whether they fall on material boundaries or correspond to known sub-materials visible in the RGB image.
minor comments (5)
- [Section 2] The text appears truncated after the introductory description of DD-CASSI; the passage jumps to 'son homogénéité relative' without presenting the model equations, the definition of the threshold T, or the notation for the number of coded acquisitions A and spectral bands W. Please complete this section.
- [Figures 3-4] The figure captions should indicate explicitly which panels correspond to the reference classification and which to each threshold value, and the axes limits and color scales should be kept consistent across panels to make the visual comparison fair.
- [Section 1] The statement that 'des observations similaires ont été obtenues sur d'autres jeux de données comme Indian Pines, CAMCATT' is not substantiated anywhere in the paper; either include the supporting evidence or present it as a conjecture to be tested.
- [Conclusions] The conclusion generalizes from a single scene and two classes to 'les vérités terrain dans les jeux de données hyperspectrales' at large; this extrapolation should be qualified, since the analysis is only illustrative.
- [References] Reference [8] appears to have an atypical page range (8174–8185 for a JOSA A article); please verify the volume and pages, and also check the spelling of the author names in reference [6].
Circularity Check
The 'more coherent regions' result is circular: the algorithm is constructed to find spectrally similar regions and is then evaluated with the same spectral-similarity metric (SAM); the inference that reference labels contain errors rests on this by-construction coherence gain.
-
self definitional
[Section 1, paragraph 3; Abstract]
"Dans ce cadre, nous avons proposé un algorithme de classification spectrale non supervisée [3], c'est-à-dire identifiant automatiquement les régions correspondant à des spectres similaires inconnus a priori, à partir de quelques acquisitions codées, en s'affranchissant d'une étape de reconstruction."
The method's stated objective is to find regions of similar spectra, and the paper's abstract equates a good class with spectral coherence ('détecter des régions spectralement plus cohérentes'). Reference classes with high intra-class spectral variability are then called 'voire erreurs de classification' because they do not match this similarity-based notion of class. The conclusion that reference labels contain errors is thus built into the definition of a class adopted by the method; it does not test ground-truth correctness independently of that definition.
-
fitted input called prediction
[Section 4, 'Pour quantifier la variabilité intra-classe...' and Fig. 4 analysis]
"Pour quantifier la variabilité intra-classe, dans chaque région labellisée, nous calculons le SAM entre les spectres des pixels et le spectre médian. ... Nous constatons que les distributions de SAM et de RMSE des régions classées par notre algorithme sont plus concentrées autour de 0 et présentent des valeurs moyennes plus faibles que celles des régions de la classification de référence, traduisant une meilleure cohérence spectrale intra-classe."
The threshold T directly controls how similar spectra must be to be grouped: decreasing T creates smaller, more homogeneous regions. The evaluation then measures SAM with respect to the median spectrum — the same spectral-similarity criterion used to form the regions. The observed decrease in SAM for decreasing T is therefore a direct consequence of the construction, not independent evidence that the reference label is wrong. The claim 'notre méthode permet de mieux détecter les régions véritablement homogènes' is a restatement of the region-partitioning criterion, and the paper's own Section 1 acknowledges that legitimate sources of intra-class variability (illumination, material differences, mixing) are indistinguishable by this metric.
full rationale
The paper does have an externally anchored component: the Pavia University labels are human annotations, and the observation that the Meadows reference class contains visually distinct sub-areas with different median spectra is an empirical finding that does not depend on the algorithm. However, the quantitative demonstration that the proposed method 'better detects' truly homogeneous regions is circular. The algorithm is designed to group spectrally similar pixels, and the evaluation of its success uses SAM between pixels and the median spectrum of the region — the same spectral-similarity notion that drives the grouping. Lowering the threshold T mechanically reduces intra-region SAM, so the reported improvement over the reference is partly a restatement of the segmentation criterion. The further inference that high-SAM reference classes contain 'erreurs de classification' requires the unvalidated assumption that spectral homogeneity, as measured by SAM, is equivalent to correct class membership. The paper itself acknowledges that variability can lawfully arise from illumination, within-material spatial differences, and mixing (Section 1), so the conclusion is not forced by the data. Overall, this is partial circularity in the central evaluation claim, not a fully self-referential derivation.
Assumptions & free parameters
free parameters (3)
- homogeneity threshold T =
0.2 and 0.05 used; no principled selection criterion given
- SIMCA instrument model parameters
- number of coded acquisitions A =
10
assumptions (4)
- domain assumption Intra-class spectral variability can be reduced to global intensity scaling for the purpose of class definition.
- domain assumption Pavia University ground truth is representative of common hyperspectral benchmarks.
- domain assumption SIMCA simulator faithfully reproduces DD-CASSI coded acquisitions.
- domain assumption Spectral homogeneity (low SAM) is an appropriate criterion for assessing the quality of an unsupervised classification.
Cite this review
Pith. "Pith review of Classification non supervis{\'e}es d'acquisitions hyperspectrales cod{\'e}es : quelles v{\'e}rit{\'e}s terrain ?." pith.science (2026). https://pith.science/paper/W5J4I2XJ
@misc{pith2026250803753,
author = {Pith},
title = {Pith review of: Classification non supervis\'ees d'acquisitions hyperspectrales cod\'ees : quelles v\'erit\'es terrain ?},
year = {2026},
howpublished = {\url{https://pith.science/paper/W5J4I2XJ}},
note = {Machine review of arXiv:2508.03753}
}
read the original abstract
We propose an unsupervised classification method using a limited number of coded acquisitions from a DD-CASSI hyperspectral imager. Based on a simple model of intra-class spectral variability, this approach allow to identify classes and estimate reference spectra, despite data compression by a factor of ten. Here, we highlight the limitations of the ground truths commonly used to evaluate this type of method: lack of a clear definition of the notion of class, high intra-class variability, and even classification errors. Using the Pavia University scene, we show that with simple assumptions, it is possible to detect regions that are spectrally more coherent, highlighting the need to rethink the evaluation of classification methods, particularly in unsupervised scenarios.
Reference graph
Works this paper leans on
-
[1]
NV5 Geospatial Software, ENVI Classic Tutorial: Spectral Angle Map- per (SAM) and Spectral Information Divergence (SID) Classification
- [2]
-
[3]
T.-T. DINH, H. CARFANTAN , A. MONMAYRANT et S. LACROIX : Sta- tistical Tests for Hyperspectral Coded Data Unsupervised Classification. In IEEE Workshop on Hyperspectral Imaging and Signal Processing : Evolution in Remote Sensing (WHISPERS) , Helsinki, Finland, 2024
work page 2024
-
[4]
L. DRUMETZ , M. VEGANZONES , S. HENROT , R. PHLYPO , J. CHANUS - SOT et C. JUTTEN : Blind Hyperspectral Unmixing Using an Extended Linear Mixing Model to Address Spectral Variability.IEEE Transactions on Image Processing, 25(8):3890–3905, août 2016
work page 2016
-
[5]
M. FAUVEL , Y .TARABALKA , J. A. BENEDIKTSSON , J. CHANUSSOT et J. C. TILTON : Advances in Spectral-Spatial Classification of Hyper- spectral Images. Proc. IEEE, 101(3):652–675, mars 2013
work page 2013
- [6]
-
[7]
M. E. GEHM, R. JOHN et D. J. BRADY : Single-shot compressive spectral imaging with a dual-disperser architecture. Optics Express, 15(21):1013–14027, nov. 2007
work page 2007
-
[8]
E. HEMSLEY , I. ARDI, T. ROUVIER , S. LACROIX , H. CARFANTAN et A. MONMAYRANT : Fast reconstruction of hyperspectral images from coded acquisitions using a separability assumption. Journal of the Optical Society of America A , 30(5):8174–8185, fév. 2022
work page 2022
Show all 10 references
-
[9]
ROUPIOZ , X
L. ROUPIOZ , X. BRIOTTET , K. ADELINE , A. AL BITAR, D. BARBON - DUBOSC , R. BARDA -CHATAIN, P. BARILLOT , S. BRIDIER , E. CAR- ROLL et C. CASSA : Multi-source datasets acquired over Toulouse (France) in 2021 for urban microclimate studies during the CAM- CATT/AI4GEO field cam...
2021
-
[10]
ROUXEL , A
A. ROUXEL , A. MONMAYRANT , S. LACROIX , H. CAMON et S. LOPEZ : Accurate ray-tracing optical model for coded aperture spectral snapshot imagers. Applied Optics, 63(7):1828–1838, 2024. 4
2024
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.