{"id":"41182f0d-51f6-4c8f-99d7-9927ebb3f7b2","arxiv_id":"2508.03753","paper_version":1,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Intra-class spectra in the Pavia University ground truth vary more than labels suggest, and a threshold-based unsupervised method isolates spectrally more homogeneous regions, questioning ground-truth-based evaluation.","lead":"This paper warns that human-made ground truth labels for hyperspectral images, like the Pavia University scene, are not as reliable as commonly assumed: some labeled regions contain very different spectra. Using an unsupervised classifier designed for coded snapshot images, the authors find more spectrally coherent sub-regions and argue that evaluation of unsupervised classification needs to be rethought.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim that reference labels contain classification errors rests on an unvalidated identification of low SAM with ground-truth correctness; the Section 4 comparison is circular because the method optimizes the same spectral-coherence criterion it uses to judge the reference.","rationale":"The reader's weakest assumption is exactly the point on which the paper's conclusion depends. Section 4 establishes the superiority of the detected regions by lower SAM and RMSE relative to the reference classes, but the unsupervised algorithm is itself defined as a segmentation into spectrally homogeneous regions controlled by a threshold T. Using the same criterion as the arbiter of ground-truth quality is circular. The paper's Section 1 lists illumination, spatial material differences, and mixtures as legitimate sources of intra-class variability, so a high-SAM reference class is not by itself evidence of labeling error. The evidence is compatible with the alternative reading that a material class may lawfully contain structured spectral variability, and that low SAM is one specific coherence notion, not the truth of the class taxonomy. This concern justifies a conditional verdict: the caution about ground-truth evaluation is valuable, but the strong claim that reference labels contain 'erreurs de classification' is not established. A field-based validation on a dataset with in-situ spectra, such as CAMCATT, would break the circularity. I therefore keep the reader's CONDITIONAL verdict unchanged.","tokens_in":6174,"tokens_out":3719,"duration_ms":54590,"concrete_test":"Use the CAMCATT dataset (reference [9]), which provides in-situ spectrometer measurements, to compute the proposed method's detected regions at T=0.2 for a reference class, then compare those regions against the independently field-measured material labels. If the detected sub-regions coincide with distinct field-measured materials, the homogeneity-as-truth assumption is supported; if they instead subdivide continuous illumination, moisture, or weathering gradients, the central claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central conclusion — that Pavia University ground truth contains pixels with dissemblance spectra, 'voire des spectres aberrants', and that the proposed method detects 'les régions véritablement homogènes' — depends on treating low intra-class SAM as the normative definition of a correct class. This assumption is doing all the work. The method's threshold T and median-spectrum representation are themselves designed to produce spectrally homogeneous regions, so the Figure 3 and 4 comparison is circular: lower SAM in the detected regions is largely a restatement of the segmentation criterion, not independent evidence that the reference labeling is wrong. The paper's own Section 1 acknowledges that intra-class variability can lawfully arise from illumination differences, within-material spatial variation, and mixing, which SAM cannot distinguish from mislabeling. Consequently, a high-SAM reference class such as Meadows may reflect a real material with structured variability rather than an annotation error. This is not an internal inconsistency, but it is a correctness risk: the empirical evidence is compatible with the alternative interpretation that spectral homogeneity, as measured by SAM, is simply not equivalent to ground-truth class membership.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an unsupervised classification method for DD-CASSI coded-aperture hyperspectral data, based on a simple model of intra-class spectral variability with a homogeneity threshold T, and evaluates it on the Pavia University scene. Its central claim is that common ground truths such as Pavia University group spectrally dissimilar pixels into the same class and therefore do not faithfully reflect true spectral homogeneity, biasing the evaluation of unsupervised classification. The evidence consists of SAM and RMSE maps and histograms for two reference classes (Meadows and Bitumen), showing that the regions detected by the method have narrower intra-class SAM/RMSE distributions than the corresponding reference classes. The paper concludes that reference annotations contain classification errors and that unsupervised classification evaluation should be rethought.","tokens_in":6367,"tokens_out":3666,"duration_ms":44598,"significance":"The paper addresses a genuinely important issue: reference labels in hyperspectral benchmarks are not error-free, and their limitations are often ignored in unsupervised classification evaluation. The use of a realistic coded-acquisition simulator (SIMCA), the explicit comparison of median spectra, and the provision of spatial SAM maps are strengths. However, the central evidence is currently circular: the method is designed to produce spectrally homogeneous regions using a spectral-angle threshold, and the evaluation uses the same spectral-angle (SAM) criterion to show that detected regions are more coherent than the reference. The conclusion that the reference contains 'spectres aberrants' conflates high intra-class SAM with labeling error, without an independent ground-truth definition. If the authors can provide such independent validation, the paper would make a useful contribution; in its present form, it demonstrates that the method produces spectrally more coherent regions, not that the Pavia University ground truth is wrong.","major_comments":[{"comment":"The evaluation uses the same metric (SAM) that the segmentation criterion T is designed to minimize. The observed decrease in intra-region SAM for detected regions is therefore partly by construction. To support the claim that the reference ground truth contains classification errors, the paper needs an independent reference—for example, in-situ spectral measurements (available for the CAMCATT dataset), known material samples, or a comparison of subregion spectra against spectral libraries—or at minimum a control experiment showing that the SAM reduction is larger than what would be obtained by any fine-grained segmentation of a physically homogeneous material.","section":"Section 4, Figs. 3-4"},{"comment":"The quantitative evidence is limited to two classes (Meadows and Bitumen) and is purely descriptive. No statistical tests, confidence intervals, or accuracy measures are reported for the histograms in Fig. 4, and the claim that the reference is less homogeneous rests on visual inspection. The authors should report, for each class and threshold, the number of pixels, the median and quantiles of SAM and RMSE, and a statistical comparison (e.g., a bootstrap test or a two-sample Kolmogorov-Smirnov test) between reference-class and detected-region distributions.","section":"Sections 3-4, Fig. 4"},{"comment":"The value T ≈ 0.09 is selected after observing that the SAM distributions become narrowest near that threshold, and the widening for smaller T is attributed to estimation difficulties on small regions. This post hoc selection, combined with the uncontrolled effect of region size, makes it difficult to attribute the observed coherence to the method's model rather than to the chosen threshold and to varying pixel counts. The comparison should be repeated with region-size-matched samples or with a principled criterion for selecting T, such as a stability or model-selection criterion.","section":"Section 4, threshold T"},{"comment":"The paper acknowledges that intra-class spectral variability can lawfully arise from illumination differences, within-material spatial variation, and mixing, which SAM cannot distinguish from mislabeling. Yet the conclusion in Section 4 interprets high SAM in the Meadows reference class as evidence of 'spectres dissemblables, voire des spectres aberrants'. This inference is not justified unless the authors can show that the high-SAM pixels are inconsistent with physical within-class variability. A concrete test would be to examine the spatial distribution of high-SAM pixels and test whether they fall on material boundaries or correspond to known sub-materials visible in the RGB image.","section":"Section 1 vs. Section 4"}],"minor_comments":[{"comment":"The text appears truncated after the introductory description of DD-CASSI; the passage jumps to 'son homogénéité relative' without presenting the model equations, the definition of the threshold T, or the notation for the number of coded acquisitions A and spectral bands W. Please complete this section.","section":"Section 2"},{"comment":"The figure captions should indicate explicitly which panels correspond to the reference classification and which to each threshold value, and the axes limits and color scales should be kept consistent across panels to make the visual comparison fair.","section":"Figures 3-4"},{"comment":"The statement that 'des observations similaires ont été obtenues sur d'autres jeux de données comme Indian Pines, CAMCATT' is not substantiated anywhere in the paper; either include the supporting evidence or present it as a conjecture to be tested.","section":"Section 1"},{"comment":"The conclusion generalizes from a single scene and two classes to 'les vérités terrain dans les jeux de données hyperspectrales' at large; this extrapolation should be qualified, since the analysis is only illustrative.","section":"Conclusions"},{"comment":"Reference [8] appears to have an atypical page range (8174–8185 for a JOSA A article); please verify the volume and pages, and also check the spelling of the author names in reference [6].","section":"References"}],"recommendation":"major_revision","confidential_remarks":"This is a short conference-style paper whose central claim—that the Pavia University ground truth contains classification errors—is currently supported only by a circular spectral-coherence comparison. I recommend major revision because the issue is fixable within the paper's scope: the authors can add an independent validation (e.g., comparison with in-situ spectra or with sub-regions validated by RGB/field knowledge) and quantitative statistical tests. I would not recommend rejection, as the topic is relevant and the authors' approach of studying ground truth limitations through coded acquisitions is potentially valuable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new thing here is the quantitative look at the Pavia University ground truth: SAM maps and histograms for the Meadows and Bitumen classes show that the reference labels contain pixels with quite different spectra, and that threshold-based sub-regions have lower intra-class SAM. That is a real, checkable observation, and it deserves to be taken seriously by anyone who evaluates unsupervised hyperspectral classification against Pavia-U.\n\nThe paper is honest in an important way: Section 1 explicitly lists legitimate sources of intra-class variability—illumination, within-material spatial variation, mixing—before the analysis starts. It also makes a fair choice in Section 4 by comparing median spectra of detected regions against the median spectrum of the reference class, rather than using the regularized reconstructed spectra that only exist for the proposed method.\n\nWhere the paper overreaches is in the leap from \"the reference class has high SAM\" to \"the ground truth contains classification errors.\" The evaluation is circular to a real degree: the method's threshold T is precisely a spectral-coherence threshold, so lower SAM inside the detected regions is partly a restatement of the segmentation criterion. The independent evidence—human labels vs. algorithm—is exactly what is at issue, so it cannot settle the question. The paper's own Section 1 undermines its later conclusion: high SAM in Meadows could be a lawful material class with weathering, moisture, or mixtures, and SAM cannot tell that apart from mislabeling.\n\nThe empirical support is thin even on its own terms: two classes, no statistical tests, no error bars, no quantitative accuracy measure, no baselines, and no code or data artifacts. The Section 5 claim that such ground truths \"can introduce a significant bias\" is plausible as a caution, but it is not proven by the evidence shown. The French text and truncation also prevented me from checking the derivation in Section 2, though the method itself is prior work.\n\nIn short: this is a useful position-style warning about reference labels, with a nice illustrative analysis, but the strong claim about classification errors is not supported at the level the paper states. It should not be taken as a demonstrated indictment of Pavia-U; it should be taken as a reason to look more carefully at what agreement with reference labels actually means.\n\nI would send this to peer review with the expectation of heavy revision: reviewers should ask for more classes, proper statistical comparisons, and a discussion of what spectral coherence can and cannot establish. As a workshop contribution it already works; as a journal claim it needs more.","headline":"Useful caution about Pavia-U ground truth, but the headline claim that reference labels contain classification errors is loaded: the evidence is two classes and a SAM-based comparison the method itself optimizes.","tokens_in":6915,"tokens_out":1720,"would_cite":true,"duration_ms":23994,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Expert ground-truth labels for hyperspectral scenes group spectrally dissimilar pixels, making them a biased yardstick for unsupervised classification.","keywords":["hyperspectral imaging","DD-CASSI","coded aperture snapshot imaging","unsupervised classification","ground truth evaluation","spectral angle mapper","intra-class spectral variability","Pavia University"],"falsifier":"Take the Pavia University Meadows class, compute each pixel's SAM to the median, and isolate the high-SAM sub-regions the method flags; if field or higher-resolution data show those sub-regions are distinct materials (different vegetation or surfaces), the ground truth is wrong and the paper's conclusion holds, whereas if they are all the same material under natural variation, the method is over-segmenting lawful variability.","tokens_in":5991,"feed_emoji":"🛰️","tokens_out":7019,"duration_ms":76123,"temperature":0.7,"pith_summary":"The paper contends that the ground truths commonly used to score hyperspectral image classification—expert-drawn label maps such as Pavia University's—are not spectrally homogeneous: a single reference class often groups pixels with strikingly different spectra, including outliers. It argues that evaluating an unsupervised classifier by agreement with such labels therefore rewards the wrong notion of class, because the labels overstate spectral coherence. Using a DD-CASSI coded-aperture imager model and only ten coded acquisitions (a tenfold data reduction), the authors' unsupervised method splits Pavia University's Meadows class into 7 to 20 spectrally coherent sub-regions and removes aberrant pixels from Bitumen, with lower intra-class spectral angle (SAM) than the reference labels. On this basis the paper calls for rethinking what a 'class' is and how unsupervised classification should be evaluated.","feed_headline":"Reference labels mix unlike spectra, skewing classification scores","feed_subtitle":"A coded-aperture classifier finds lower intra-class variability than expert labels on Pavia University.","key_machinery":"The engine of the analysis is a statistical model of intra-class spectral variability in which each pixel in a class is the same reference spectrum multiplied by a scale factor (a model of illumination variation), plus noise. The unsupervised classifier from the authors' earlier work tests, from coded DD-CASSI measurements, whether groups of pixels are consistent with sharing one underlying spectrum up to scale; a threshold T controls how tight the homogeneity must be, and lower T yields smaller, more coherent regions. For each detected region the median of the full pixel spectra is used as its representative spectrum, and the spectral angle mapper (SAM) between each pixel and that median quantifies residual intra-class variability. The coded data themselves are generated by a ray-tracing simulator of the DD-CASSI optical system, so the whole pipeline runs on realistic simulated acquisitions with ten times fewer measurements than the full cube.","core_discovery":"The paper's central discovery is that reference ground truth is a flawed yardstick for unsupervised classification: on Pavia University, the reference class Meadows contains several spectrally distinct materials, while Bitumen contains a few outlier spectra. When the authors' unsupervised classifier is applied directly to coded acquisitions compressed by a factor of ten, it detects regions whose SAM and RMSE distributions around the median spectrum are more concentrated near zero than the reference class distributions, meaning the detected regions are more spectrally homogeneous. The authors read this as evidence that the reference labels contain heterogeneity and labeling errors, and that their method better captures 'truly homogeneous' regions. The consequence is that numerical scores computed against such ground truths do not faithfully measure spectral coherence, so the evaluation of unsupervised methods needs to be redesigned.","pith_inferences":["An immediate testable extension is to build an alternative ground truth for Pavia University by subdividing reference classes at SAM thresholds, then re-running standard supervised and unsupervised benchmarks; if rankings change, the original labels were indeed biasing comparisons.","The same SAM-coherence audit could be applied to datasets whose reference spectra were measured in situ (such as the CAMCATT campaign) to separate lawful material variability, like moisture or weathering, from annotation error—if in-situ spectra also scatter widely within a label, the 'error' is partly natural variability.","The scale-factor variability model is a deliberate simplification; extending it to direction-dependent variability (e.g., mixtures or non-illumination effects) would let the method distinguish between 'same material, different lighting' and 'same label, different materials,' which the current SAM metric cannot do.","If ground truths are biased, then paper-to-paper comparisons of unsupervised hyperspectral classifiers may be comparing how well each method mimics annotation artifacts; a community benchmark with multiple thresholded labels or intrinsic homogeneity scores would be a more honest yardstick."],"forward_implications":["If ground truths overstate spectral homogeneity, then accuracy, precision, and confusion-matrix comparisons against them are systematically biased in favor of classifiers that reproduce the annotation's grouping.","For scenes like Pavia University, a single reference class such as Meadows should be subdivided into several spectrally homogeneous sub-regions before it is used as an evaluation target.","Unsupervised classification from coded acquisitions is viable at tenfold compression: the method identifies coherent classes and reference spectra without full cube reconstruction.","The Bitumen analysis shows that reference labels can be partly salvaged by detecting and excluding a small number of aberrant spectra, rather than re-labeling the whole class.","The class definition itself—what counts as one material versus a mixture or a spatially varying material—becomes a modeling choice that must be made explicit in any evaluation."],"supporting_citations":[{"why":"Supplies the Pavia University scene and its expert reference labels that the paper audits and subdivides.","marker":"[6]"},{"why":"Defines the unsupervised classification method whose detected regions and median spectra are compared with the reference labels.","marker":"[3]"},{"why":"Introduces the DD-CASSI dual-disperser coded aperture architecture that the coded acquisition model is based on.","marker":"[7]"},{"why":"Provides the ray-tracing optical model used to generate realistic coded acquisitions from the reference cubes.","marker":"[10]"},{"why":"Provides the separability-assumption reconstruction used for regularized spectra, though the classification comparison relies on median spectra instead.","marker":"[8]"},{"why":"Defines the spectral angle mapper (SAM) used to quantify intra-class spectral variability in all comparisons.","marker":"[1]"},{"why":"Cited as a recent, rigorously built dataset with in-situ reference spectra, used to argue ground-truth issues are not specific to Pavia University.","marker":"[9]"},{"why":"Documents spectral variability in hyperspectral data (extended linear mixing), supporting the paper's claim that intra-class variability is a genuine and difficult phenomenon.","marker":"[4]"}],"fun_headline_variants":["Ground truth labels hide spectral mixing, skewing scores","Unsupervised classifier finds purer classes than expert labels","Reference labels flawed: classifier finds more coherent spectra","Pavia ground truth mixes spectra; evaluation scores misleading","Coded hyperspectral data expose flaws in classification labels"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a correct class must be spectrally homogeneous, so that high spectral angle within a reference class counts as a labeling error; a material can, however, lawfully contain spectral variation from moisture, weathering, or mixing, and the spectral angle alone cannot tell lawful variability from mislabeling.","fun_headline_variants_meta":{"raw":{"variants":["Ground truth labels hide spectral mixing, skewing scores","Unsupervised classifier finds purer classes than expert labels","Reference labels flawed: classifier finds more coherent spectra","Pavia ground truth mixes spectra; evaluation scores misleading","Coded hyperspectral data expose flaws in classification labels"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000168,"raw_usage":{"total_tokens":1200,"prompt_tokens":823,"completion_tokens":377,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":439,"completion_tokens_details":{"reasoning_tokens":301}},"tokens_in":439,"tokens_out":377,"duration_ms":4669,"temperature":1.0,"reasoning_tokens":301,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T05:09:43.167630+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the Pavia University Meadows class, compute each pixel's SAM to the median, and isolate the high-SAM sub-regions the method flags; if field or higher-resolution data show those sub-regions are distinct materials (different vegetation or surfaces), the ground truth is wrong and the paper's conclusion holds, whereas if they are all the same material under natural variation, the method is over-segmenting lawful variability.","supporting_citations":[{"cited_title":"GAMBA , M","cited_arxiv_id":null,"evidence_quote":"Supplies the Pavia University scene and its expert reference labels that the paper audits and subdivides."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the unsupervised classification method whose detected regions and median spectra are compared with the reference labels."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the DD-CASSI dual-disperser coded aperture architecture that the coded acquisition model is based on."},{"cited_title":"ROUXEL , A","cited_arxiv_id":null,"evidence_quote":"Provides the ray-tracing optical model used to generate realistic coded acquisitions from the reference cubes."},{"cited_title":"HEMSLEY , I","cited_arxiv_id":null,"evidence_quote":"Provides the separability-assumption reconstruction used for regularized spectra, though the classification comparison relies on median spectra instead."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the spectral angle mapper (SAM) used to quantify intra-class spectral variability in all comparisons."},{"cited_title":"ROUPIOZ , X","cited_arxiv_id":null,"evidence_quote":"Cited as a recent, rigorously built dataset with in-situ reference spectra, used to argue ground-truth issues are not specific to Pavia University."},{"cited_title":"DRUMETZ , M","cited_arxiv_id":null,"evidence_quote":"Documents spectral variability in hyperspectral data (extended linear mixing), supporting the paper's claim that intra-class variability is a genuine and difficult phenomenon."}],"review_version":1}