{"id":"58ce4289-2c0f-47f7-a9ac-1ccfa847470d","arxiv_id":"2501.01484","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"CLUES is a non-parametric clustering workflow that combines Sequencer distance matrices with hierarchical clustering and silhouette scoring to classify debris disk and mineral spectra.","lead":"This paper presents CLUES, a machine-learning workflow that combines the Sequencer algorithm with hierarchical clustering to sort and classify infrared spectra of dusty planetary disks and minerals. It is a methods paper showing the tool can recover known mineral groupings, setting up a larger catalog analysis promised for a follow-up paper.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The real-disk validation inherits the P1 continuum ambiguity: the global polynomial fit uses noisy 30 µm anchors, so the acknowledged 20/30 µm normalization problem can shift even the 10 µm emissivities that drive the HD 113766 clustering result.","rationale":"The paper's central claim is conditional on the preprocessing chain, and the only experiment that exercises that chain on real debris disk data is HD 113766. The authors are transparent about the 20/30 µm limitation, but they do not quantify how it affects the clustering outputs. The lab-based validations (forsterite Fo ordering, meteorite classes) are internal to the clustering stack and do not propagate the continuum ambiguity, so they cannot mitigate this. I therefore align with the reader's CONDITIONAL verdict rather than ACCEPT: the method is promising and honestly benchmarked, but the headline debris-disk application requires a sensitivity test showing the HD 113766 grouping (and, by extension, the Paper II pipeline) is stable under reasonable alternative continuum models. If the proposed test shows the grouping persists, the concern is retired; if it flips, the central validation is a preprocessing artifact and the method needs a more robust continuum treatment before catalogue-wide conclusions. This is not a reason to reject the paper; it is the concrete missing experiment that separates a strong methods paper from a conditional one.","tokens_in":24911,"tokens_out":11572,"duration_ms":119613,"concrete_test":"Re-run the Section 5.1.2 HD 113766 experiment with alternative continuum choices: (a) a locally fit blackbody or spline using only anchors at 5.6-7.9 and 13-14.8 µm, excluding the 30/35 µm anchors; (b) a two-blackbody disk continuum; and (c) the existing polynomial without the non-negativity offset. After the same binning and 8-13 µm normalization, recompute the Sequencer distance matrix and Ward dendrogram. If HD 113766 leaves the Fe-rich forsterite clade in any of these variants, the central debris-disk validation is an artifact of the P1 continuum assumption rather than a property of CLUES.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing weakness is in P1, which feeds every CLUES distance. Eq. 3 defines average emissivity as Fdisk/Fcont with Fcont a third-order polynomial anchored at 5.61-7.94, 13.02-13.50, 14.32-14.83, 30.16-32.19, and 35.07-35.92 µm (Sec. 3.3), then shifted by a non-negativity offset. Section 3.5 concedes that this continuum treatment is unreliable for the broad 20/30 µm complexes because a single temperature cannot represent the underlying continuum. This is not only a long-wavelength nuisance: because the polynomial is global, errors from the 30-35 µm anchors change the polynomial level under the 10 µm region, altering the 10 µm line-to-continuum ratios and therefore the features normally considered reliable. The distance matrix (Sec. 4.1-4.3) is computed on the full emissivity vector, so continuum errors propagate into the MST, the Ward dendrogram, and the HD 113766 grouping that anchors the real-disk validation. The forsterite and meteorite experiments use library emissivity/reflectance inputs and cannot test this debris-disk preprocessing; they validate the clustering stack, not P1. The paper defers sensitivity tests to Paper II, but the current central claim that CLUES works as advertised on IRS debris disk spectra is only as strong as this untested continuum assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces CLUES, an unsupervised machine-learning workflow for classifying mid-infrared spectra, built on the Sequencer algorithm. The workflow computes multi-scale distance matrices (using a distance scale and EMD-like metric), then combines a minimum spanning tree, Ward hierarchical clustering, silhouette-score cluster selection, and MDS visualization. It is validated on three benchmark tasks: a forsterite emissivity library with known Fo numbers, a mixed mineral library into which the well-studied debris disk HD 113766 is embedded, and a set of 59 meteorite reflectance spectra. The paper's central claim is that CLUES non-parametrically and interpretably recovers known mineralogical groupings and can therefore be applied to Spitzer IRS debris disk spectra, with a full demographic analysis deferred to Paper II.","tokens_in":25163,"tokens_out":8316,"duration_ms":84420,"significance":"If the method holds up, it would fill a practical gap: a systematic, reproducible, and interpretable way to classify the large Spitzer IRS debris disk sample and similar MIR datasets. The strongest concrete evidence is the forsterite experiment (Figure 6/7), where the MST and hierarchical clustering recover the Fo-number ordering and pair clean/noisy copies of the same composition, and the HD 113766 experiment (Figure 14), where the disk is grouped with Fe-rich forsterite, consistent with earlier detailed modeling. The meteorite experiment (Figure 16) additionally shows that broad compositional classes separate cleanly. These benchmarks use external, pre-existing libraries and are therefore meaningful validation of the clustering stack, not circular demonstrations. The main weakness is that the preprocessing step P1, specifically the continuum normalization used for real IRS spectra, is not validated end-to-end and is acknowledged by the authors to be unreliable in parts of the wavelength range used.","major_comments":[{"comment":"The P1 continuum normalization is load-bearing for the real-disk demonstration but is not validated by the benchmark experiments. The average emissivity in Eq. (3) is Fdisk/Fcont, where Fcont is a third-order polynomial anchored at 5.61–7.94, 13.02–13.50, 14.32–14.83, 30.16–32.19, and 35.07–35.92 µm, and Section 3.5 concedes that this treatment is unreliable for the 20/30 µm complexes because the broad features cannot be represented by a single blackbody-like continuum. Because the polynomial is global, errors at the 30–35 µm anchors can change the polynomial level under the 10 µm region and therefore alter the 10 µm line-to-continuum ratios that drive the HD 113766 grouping in Fig. 14(d). The forsterite, mineral-library, and meteorite experiments use laboratory emissivity/reflectance inputs and thus validate the clustering stack rather than the P1 preprocessing of IRS spectra. I request a sensitivity analysis (anchor-set choice, polynomial degree, offset procedure, binning factor, and wavelength truncation) on at least HD 113766 and a few representative disk spectra, or an explicit narrowing of the real-disk claims to Paper II.","section":"3.3, 3.5, Eq. (3), 5.1.2"},{"comment":"The distance scale and metric are tuned on the same data used for the reported demonstrations, and no stability check is shown. The scale l=5 is selected because it maximizes MST elongation for the forsterite library (Table 1: elongation 21.49 versus 15.11 at l=1), while the mineral library, HD 113766 experiment, and meteorite experiment use l=10, l=15, and l=1, respectively (Table 3). The 'non-parametric' label therefore needs qualification: the workflow has hyperparameters whose optimal values differ by dataset, and the resulting clusters (e.g., the optimal cluster number 9 and the grouping of HD 113766 with Fe-rich forsterite) may depend on these choices. I ask for a sensitivity test over l and metric for at least the forsterite and HD 113766 experiments, reporting cluster memberships or silhouette stability, so that the reader can see whether the qualitative conclusions are robust rather than selected.","section":"4.3, Table 1, Table 3"},{"comment":"The robustness claim that 'SNR > 5' is needed for reasonable clustering results is based solely on uncorrelated Gaussian noise added to library spectra. This does not cover the correlated artifacts (fringing, point-to-point calibration residuals) that motivate the binning step in Section 3.4, nor does it cover continuum-fitting errors from P1. The 20%-noise degradation test is useful, but as written the SNR>5 statement is likely too strong for real IRS disk spectra. I recommend either extending the simulations to include correlated/fringe-like noise and a small set of continuum-mismatch scenarios, or restricting the robustness statement to the ideal-library case.","section":"6.2"}],"minor_comments":[{"comment":"The distance measure called EMD in Section 4.2 appears to be the energy distance used in Baron & Ménard (2021), not the standard Earth Mover's/Wasserstein distance; please align the terminology and citation with the original implementation.","section":"4.2"},{"comment":"Equation (6) has a zero denominator on the diagonal of the distance matrix; please state explicitly that diagonal entries are excluded from the percentage-difference histogram in Figure 10, or redefine the statistic for those entries.","section":"4.5, Eq. (6)"},{"comment":"The statement that CLUES can distinguish grain sizes for HD 113766 and places its spectrum among library spectra with a 2–5 µm grain size distribution is presented without a figure, table, or quantitative criterion; since this is a nontrivial claim, it should be documented or removed.","section":"6.3"},{"comment":"The Ward linkage update in Eq. (4) is written as a distance between pairwise distances, which is not the usual Ward criterion; please provide a reference or clarify how this recursion defines the increase in within-cluster variance.","section":"4.4, Eq. (4)"},{"comment":"A code/data availability statement is missing. Since the paper introduces a named tool (CLUES) and a methodology is the primary deliverable, providing a public repository would substantially improve reproducibility and practical uptake.","section":"7 and elsewhere"},{"comment":"There are small typographical errors, including 'over scalesl' in Section 6.1 and a stray space in 'Dan M. W atson' in the author list; these should be corrected in proof.","section":"6.1 and byline"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a methods-focused journal. The blocking issue is end-to-end validation of P1 on real IRS spectra, not the clustering benchmarks themselves; the requested sensitivity analysis is substantive but local. I would not require the full Paper II catalog demographics in this revision. A code release would materially increase the value of a methodology paper, but its absence is not a correctness issue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nCLUES is a well-assembled pipeline that applies Baron & Ménard's Sequencer, EMD distances, hierarchical clustering, and MDS to mid-IR mineral spectra. The forsterite library experiment, where the distance matrix and clustering recover the Fo-number ordering, is a real result and gives confidence the clustering stack works. The meteorite separation and the HD 113766 grouping with Fe-rich forsterite are also genuine external validations. The authors are unusually honest: Section 3.5 states that the 20/30 µm normalization is unreliable with a single blackbody continuum, and they flag that Paper II will explore normalization choices. That transparency is worth crediting.\n\nThe main soft spot is exactly the one the stress-test flags: the real-disk validation goes through P1, and P1's continuum is a global third-order polynomial anchored at 5.6–7.9, 13–14.8, and 30–35 µm. The 30–35 µm anchors are noisy, and because the polynomial is global, error or noise there can change the polynomial level under the 10 µm region. That means the HD 113766 result—the one that supposedly shows CLUES works on IRS debris disk spectra—is only as strong as that continuum assumption. The library and meteorite tests use emissivity/reflectance inputs and cannot test P1. The paper defers sensitivity tests to Paper II, but the current central claim for IRS data is conditional on an untested preprocessing step.\n\nTwo additional, smaller issues. The distance scale is tuned per dataset (l=5 for forsterite, l=10 for the mineral library, l=15 with the disk, l=1 for meteorites). That is not a flaw by itself, but the non-parametric label sits oddly with that much tuning, and the paper doesn't show how robust the clusters are to l. And there's no direct comparison to the spectral-indices method they criticize; that would strengthen the case that CLUES adds information over the existing approach.\n\nI would send this to peer review. It's a solid methods contribution with honest benchmarks and no red flags. The right ask is to either add a sensitivity analysis for the continuum anchor choice and the offset, or to clearly scope the claims away from the 20–30 µm features for the IRS catalog until that analysis exists. Worth citing once the revisions land.","headline":"CLUES is a useful, honestly assembled clustering pipeline whose benchmark results validate the clustering stack but not the underlying continuum preprocessing; the P1 sensitivity is the main soft spot.","tokens_in":25809,"tokens_out":2943,"would_cite":true,"duration_ms":29290,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"CLUES, a fully non-parametric clustering workflow, sorts debris disk spectra by mineral content without fitting any spectral model.","keywords":["Debris disks","Silicate grains","Mid-infrared spectroscopy","Unsupervised clustering","Sequencer algorithm","Earth Mover's Distance","Hierarchical clustering","Minimum spanning tree"],"falsifier":"Rerun CLUES on a well-modeled disk such as HD 113766 after deliberately changing the continuum fit—for example, adding a cold blackbody component or moving the anchor regions—and check whether the disk still clusters with Fe-rich forsterite; if the cluster membership or recovered Fo-number ranking shifts with the continuum choice, the mineralogical readout is dominated by the normalization the paper itself flags as unreliable.","tokens_in":24643,"feed_emoji":"🪐","tokens_out":8017,"duration_ms":71170,"temperature":0.7,"pith_summary":"This paper introduces CLUES, an unsupervised machine-learning workflow that classifies mid-infrared spectra of debris disks without assuming a spectral model or fitting parameters. The authors aim to replace the traditional handful of spectral band ratios with a full-spectrum, multi-scale distance measurement, so that hundreds of Spitzer IRS disk spectra can be sorted by mineralogical similarity rather than by a few hand-picked features. They validate the method on three test cases: a pure mineral library, a well-studied debris disk (HD 113766), and a set of 59 meteorite spectra. In each case the unsupervised groupings recover known compositional structure, including the forsterite iron-magnesium ordering and the Fe-rich forsterite content of HD 113766. If this works on the full catalog, it would give a data-driven demographic map of silicate mineralogy across several hundred debris disks without human bias.","feed_headline":"New tool sorts debris disk spectra by mineral content","feed_subtitle":"CLUES recovers known silicate groupings and matches detailed modeling of HD 113766 without any spectral fitting.","key_machinery":"The load-bearing object is the 'average emissivity' spectrum, defined in Eq. (3) as the disk flux divided by a fitted continuum; this is the quantity CLUES compares. The load-bearing mechanism is the Sequencer-based distance matrix: the spectrum is divided into chunks of width set by a 'distance scale' $l$, and each chunk pair is compared with the Earth Mover's Distance, which measures how much 'work' is needed to reshape one spectral chunk into another. A scale list $[1,2,5,10,20,50]$ is combined into an elongation-weighted matrix, so both narrow and broad features contribute. Ward-linkage hierarchical clustering then groups spectra by minimizing within-cluster variance, the silhouette score chooses the cluster count, and the minimum spanning tree plus metric multi-dimensional scaling provide the 1D and low-dimensional views. This combination is what carries the argument: it lets the full 7–33 micron wavelength range, including the broad 20 and 30 micron complexes, inform the classification without any parametric assumption about line-to-continuum ratios.","core_discovery":"The central claim is that a fully non-parametric pipeline—Sequencer's multi-scale Earth Mover's Distance distance matrix, Ward-linkage hierarchical clustering, and silhouette-score cluster selection—recovers physically meaningful mineral groupings from mid-infrared spectra, without any parametric spectral fitting. The strongest demonstration is the forsterite library experiment: with the library spectra shuffled, CLUES orders them along a minimum spanning tree by increasing Fo number (the Mg/(Mg+Fe) ratio in olivine), pairs clean and noisy copies of each composition, and clusters high-Fe from high-Mg endmembers. In the HD 113766 experiment, the disk spectrum is grouped with Fe-rich forsterite, matching the result of previous detailed parametric modeling, and the method independently flags anorthite as a previously unconsidered candidate component. The same workflow cleanly separates achondritic, carbonaceous, and ordinary chondrite meteorite spectra. The paper presents CLUES not as a replacement for detailed modeling but as a first-pass, bias-reducing engine for narrowing a vast compositional parameter space down to exemplar spectra and candidate minerals.","pith_inferences":["A natural extension not explored in the paper is to embed theoretical or laboratory spectra of amorphous silicates, carbonaceous grains, and ices into the library; the distance matrix would then double as a quantitative mineralogical-similarity scale for ranking candidates before MCMC fitting.","The anorthite flag for HD 113766 could be tested with JWST MIRI spectroscopy, which extends to shorter wavelengths; if the 10 micron match disappears when the 5–7 micron region is included, the flag is an artifact of the truncated wavelength range rather than a real composition.","The paper's noise experiment implies a practical screening rule: spectra with SNR below about 5 in the 10 micron complex will likely produce unreliable clusters, so the follow-up catalog analysis should report cluster membership uncertainties marginalized over the SNR distribution rather than point assignments."],"forward_implications":["The full 571-disk Spitzer IRS catalog can be processed through CLUES to yield a global, data-driven taxonomy of debris disk silicate mineralogy; Paper II is set up to do exactly this.","Because CLUES recovered the Fo-number ordering in forsterite without any parametric fitting, the same distance matrix can be used to rank any library spectrum (mineral, grain size, temperature) against a disk spectrum and narrow the parameter space for detailed fitting.","The HD 113766 grouping with Fe-rich forsterite reproduces prior modeling, so CLUES can serve as a consistency check for detailed spectral fits of individual disks.","The meteorite experiment shows that CLUES can separate mixed-composition samples (achondrites, carbonaceous chondrites, ordinary chondrites), so the tool generalizes beyond debris disks to any spectral library with a common wavelength grid.","CLUES is designed to scale to thousands of spectra and data cubes such as JWST MIRI IFU observations, where traditional parametric component fitting becomes computationally prohibitive."],"supporting_citations":[{"why":"Defines the Sequencer algorithm and the multi-scale distance-matrix construction that CLUES builds on.","marker":"Baron & Ménard 2021"},{"why":"Provides the laboratory forsterite emissivity library used in the Fo-number validation experiment.","marker":"Chihara et al. 2002"},{"why":"Supplies laboratory emissivity data and the association between 10, 20, and 30 micron features used to justify full-wavelength analysis.","marker":"Koike et al. 2003"},{"why":"The spectral-index method CLUES is designed to extend; also the normalization scheme that CLUES adopts.","marker":"Morlok et al. 2014"},{"why":"Source of the Spitzer IRS debris disk catalog, stellar photosphere parameters, and data reduction details.","marker":"Chen et al. 2014"},{"why":"Identifies the 120 disks with solid-state features that define the sample and provides the polynomial continuum fitting approach.","marker":"Mittal et al. 2015"},{"why":"Provides the prior detailed modeling of HD 113766 that found Fe-rich forsterite, the ground truth for the validation experiment.","marker":"Olofsson et al. 2012"},{"why":"Earlier detailed MIR modeling of HD 113766 used to cross-check CLUES's mineralogical grouping.","marker":"Lisse et al. 2008"},{"why":"Supplies the ECOSTRESS library spectra used for the mineral and meteorite test datasets.","marker":"Meerdink et al. 2019"}],"fun_headline_variants":["Sequencing silicates in debris disks with unsupervised clustering","CLUES tool sorts disk spectra by mineral content automatically","Unsupervised ML orders forsterite by Mg/Fe ratio in disks","New clustering recovers disk mineral groups without spectral fitting","Spectral clustering reveals silicate order in debris disk dust"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything downstream assumes that the continuum-subtracted, 8–13 micron normalized 'average emissivity' spectrum carries the mineralogical information, even though the authors state that this normalization cannot reliably fix the amplitudes of the 20 and 30 micron features because the underlying continuum is not a single blackbody.","fun_headline_variants_meta":{"raw":{"variants":["Sequencing silicates in debris disks with unsupervised clustering","CLUES tool sorts disk spectra by mineral content automatically","Unsupervised ML orders forsterite by Mg/Fe ratio in disks","New clustering recovers disk mineral groups without spectral fitting","Spectral clustering reveals silicate order in debris disk dust"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000344,"raw_usage":{"total_tokens":1908,"prompt_tokens":979,"completion_tokens":929,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":595,"completion_tokens_details":{"reasoning_tokens":847}},"tokens_in":595,"tokens_out":929,"duration_ms":9692,"temperature":1.0,"reasoning_tokens":847,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:26:35.740960+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun CLUES on a well-modeled disk such as HD 113766 after deliberately changing the continuum fit—for example, adding a cold blackbody component or moving the anchor regions—and check whether the disk still clusters with Fe-rich forsterite; if the cluster membership or recovered Fo-number ranking shifts with the continuum choice, the mineralogical readout is dominated by the normalization the paper itself flags as unreliable.","supporting_citations":[{"cited_title":"2021, The Astrophysical Journal, 916, 91","cited_arxiv_id":null,"evidence_quote":"Defines the Sequencer algorithm and the multi-scale distance-matrix construction that CLUES builds on."},{"cited_title":"2008, The Astrophysical Journal, 673, 1106","cited_arxiv_id":null,"evidence_quote":"Earlier detailed MIR modeling of HD 113766 used to cross-check CLUES's mineralogical grouping."},{"cited_title":"K., Hook, S","cited_arxiv_id":null,"evidence_quote":"Supplies the ECOSTRESS library spectra used for the mineral and meteorite test datasets."}],"review_version":1}