{"id":"5bc09392-0577-4230-b142-3c5367751d86","arxiv_id":"2501.11657","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A machine learning pipeline using clustering and CNNs classifies HI galaxy profiles, with a 2D image representation reported to improve accuracy by 13% over 1D methods.","lead":"This paper applies a mix of machine learning methods, including clustering and neural networks, to sort the radio profiles of hydrogen gas in galaxies into shape classes. It introduces a trick of turning the one-dimensional profiles into two-dimensional images, which the authors say improves classification accuracy, and shares all code publicly.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 13% improvement over 1D methods is unsupported: the manuscript reports only the 2D accuracy (63%) and never presents the 1D baseline it compares against.","rationale":"The reader's verdict is CONDITIONAL and the stated rationale already mentions 'no baseline 1D number', but the formally identified weakest assumption is the validity of the Espada et al. (2011) analytical labels as ground truth. My stress-test focuses on a different, more directly load-bearing gap for the central claim: the missing 1D baseline that makes the 13% improvement unverifiable. Even if the analytical labels are accepted as the target classification, the comparative claim cannot be evaluated without the 1D accuracy. This omission is internal to the paper's argument, not a matter of external consensus. The repository being public is a credit, and it makes the concrete test feasible. The verdict remains CONDITIONAL: the claim can be accepted only if the 1D baseline is supplied and the 13% gap is reproduced with a controlled comparison. No verdict change is therefore needed.","tokens_in":3983,"tokens_out":4969,"duration_ms":47753,"concrete_test":"Inspect the public GitHub repository (https://github.com/gabojaimesillanes/JAE-Intro-ICU-2024-Gabriel-Jaimes-IAA) for the 1D classification run on the 318 CIG profiles, then rerun both 1D and 2D pipelines with identical preprocessing, labels, train/test split, random state, and CNN architecture. Verify that the difference between the 2D accuracy (63%) and the 1D accuracy is exactly or approximately 13 percentage points. If the repository contains no 1D baseline or the gap differs, the headline claim requires revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim in Section 4, 'improving classification accuracy by 13% compared to traditional 1D methods', depends on a 1D baseline accuracy that never appears in the paper. Section 3 reports 'a classification success rate of 63% for classifying profiles with 2D images' and mentions 54 configurations, but gives no 1D accuracy, no error bars, and no specification of what 'traditional 1D methods' means (classifier, features, train/test split, or CNN architecture). Without that baseline, the 13% improvement cannot be computed or checked. This is not a matter of ground-truth validity; even accepting the Espada et al. (2011) labels as target classes, the comparative claim is unsubstantiated in the manuscript. The open repository may contain the missing numbers, but the paper itself does not support its headline result.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a machine-learning framework for classifying integrated HI galaxy spectral profiles, applied to 318 CIG profiles and 30,780 ALFALFA profiles. The methodology combines unsupervised clustering (K-means, spectral clustering, DBSCAN, agglomerative) for feature extraction with supervised classifiers (KNN, SVM, Random Forest) and CNN optimization. The paper's central methodological novelty is the transformation of 1D spectral profiles into three 2D representations (rotated spectrum, subtracted asymmetry, normalized asymmetry) to improve asymmetry classification. The authors report a 63% classification success rate with the 2D images and claim a 13% accuracy improvement over traditional 1D methods. All code and models are publicly available in a GitHub repository following FAIR principles.","tokens_in":4190,"tokens_out":3317,"duration_ms":36087,"significance":"If substantiated, the claimed 13% improvement over 1D classification would be a useful practical contribution for classifying large numbers of HI profiles from upcoming surveys such as SKA. The paper's strengths are its open-source repository, reproducible code, and the attempt to compare multiple unsupervised and supervised ML methods on a well-defined astronomical dataset. However, the headline quantitative claim is currently unverifiable because the paper does not report the 1D baseline accuracy or the experimental protocol behind the comparison. The evaluation also relies entirely on an analytical classification from a prior study by the same group, whose reliability is not discussed. The significance of the contribution therefore depends on strengthening the empirical reporting.","major_comments":[{"comment":"The claim 'improving classification accuracy by 13% compared to traditional 1D methods' is not supported by the manuscript because no 1D baseline accuracy is reported anywhere. Section 3 reports a 63% success rate for the 2D models, but without the 1D accuracy (and a clear definition of what 'traditional 1D methods' means, including classifier, features, train/test split, and architecture), the 13% improvement cannot be computed or checked. This is a load-bearing number for the paper's main conclusion and must be provided, along with the exact comparison protocol.","section":"Section 4"},{"comment":"The evaluation uses the analytical profile classification of Espada et al. (2011) as the target labels, and this prior work has overlapping authorship with the AMIGA group. The paper does not discuss the authority or limitations of this analytical scheme, nor does it report any measure of inter-rater agreement or independent validation. The 63% accuracy should be explicitly framed as agreement with the Espada et al. classification rather than an absolute ground truth; otherwise, the accuracy claim is vulnerable to any biases in that reference scheme. Please provide justification for using this benchmark or add independent validation.","section":"Section 2.3 and Section 3"},{"comment":"The manuscript provides no details of the CNN architecture, hyperparameters, training/validation split, number of runs, or error bars. The sentence 'the results suggest that the methodology is robust and scalable' is not backed by any statistical evidence, such as standard deviations over random seeds or the 54 configurations mentioned. Report the mean and standard deviation of the classification accuracy, and specify the configuration search space, to allow the robustness claim to be assessed.","section":"Section 3"},{"comment":"The description of the ALFALFA classification pipeline is too incomplete for reproducibility. For example, 'the temporal shapelet transform' is mentioned but not defined or referenced, and the number of clusters, feature extraction steps, and how the 18 classifications were generated are not specified. The GitHub repository is helpful, but the paper itself should contain the essential methodological parameters so that the results can be reproduced without reverse-engineering the code.","section":"Section 2.2"}],"minor_comments":[{"comment":"There is a typographical error: 'in a efficient way' should be 'in an efficient way'.","section":"Abstract"},{"comment":"The text reads 'improve the e fficiency' with a stray space; it should be 'improve the efficiency'.","section":"Section 1"},{"comment":"Figure 3 is referenced in the text only by a caption; please add a sentence in the main text describing what the three panels show and how they relate to the models.","section":"Section 2.3"},{"comment":"The number '30.780' uses a period as a thousands separator; for international consistency, use '30,780' or '30780'.","section":"Section 3"},{"comment":"The definition of W50 (width at 50 percent of peak) is standard, but it would help to state explicitly that it refers to HI line width at half the peak flux.","section":"Section 3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is unusually short for a journal article and reads more like an extended abstract or project report. The authors should be encouraged to either substantially expand the methodological details and quantitative results, or the editor may consider whether this format fits the journal's standards. The GitHub repository is a useful resource, but it cannot substitute for the paper's own reporting of the baseline accuracy and experimental setup."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the 13% improvement over 1D methods is the paper's headline claim and it does not appear in the text in a checkable form. The authors report 63% accuracy for the 2D approach but never give the 1D accuracy they are comparing against. So the comparative claim is unsubstantiated, exactly as the stress-test note says. That is the main thing you need to know.\n\nWhat is genuinely useful: the paper is a concrete, reproducible benchmark of standard ML tools on real HI data. It tries four clustering methods (K-means, spectral, DBSCAN, agglomerative), three classical classifiers, and a CNN, and it introduces three simple 2D image representations of 1D line profiles intended to expose asymmetry. The three models—rotated spectrum, subtracted asymmetry, normalized asymmetry—are easy to describe and implement. The authors also put code, documentation, and notebooks in a public repository, which is real value for anyone wanting to prototype similar work. That is the part of the paper I would take seriously.\n\nThe soft spots are mostly on the quantitative side. No baseline 1D accuracy, no error bars, no confusion matrix, no CNN architecture or hyperparameter details. 'Robust and scalable' is an assertion, not a result. On top of that, the target labels come from Espada et al. (2011), a prior paper from the same group. The ML models are trained and tested against those labels, so the 63% is a measure of agreement with that analytical classification, not an independent validation. If the analytical labels are wrong or biased, the accuracy numbers track that bias. That does not make the framework circular, but it does mean the authority of the benchmark needs separate justification. There are also internal inconsistencies (18 classifications vs 54 configurations) and a vague reference to a 'temporal shapelet transform' that is never defined. These are fixable in revision.\n\nWhere does this leave it? The paper is a solid student-project-level methods comparison, not a breakthrough. A careful reader can use the repository and the 2D model ideas, but should not quote the 13% until the baseline appears. I would send it to peer review: it deserves an external referee to push on the numbers. My recommendation: major revision, requiring the 1D baseline, uncertainty estimates, model details, and a more guarded conclusion.","headline":"A useful but loosely reported ML pipeline for HI profiles; its headline 13% improvement claim is not backed by a reported 1D baseline.","tokens_in":4727,"tokens_out":3891,"would_cite":false,"duration_ms":36430,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Converting 1D HI spectral profiles into 2D images improves classification accuracy by 13% compared with traditional 1D methods.","keywords":["HI spectral profiles","galaxy classification","convolutional neural networks","unsupervised learning","profile asymmetry","2D image representation","ALFALFA","isolated galaxies"],"falsifier":"Re-run the released pipeline on the same 318 profiles with the analytical labels replaced by blind, independent expert classifications or by labels derived from resolved HI kinematics; if the 2D CNN no longer beats the 1D CNN by roughly 13%, the claimed advantage is an artifact of the reference labels rather than a property of the 2D representation.","tokens_in":3824,"feed_emoji":"📡","tokens_out":5290,"duration_ms":53356,"temperature":0.7,"pith_summary":"This paper claims that converting one-dimensional HI spectral profiles into two-dimensional images improves automated galaxy profile classification, raising accuracy by 13% over classifying the original 1D profiles. The authors build a machine-learning pipeline that pre-processes HI spectra, applies clustering and bootstrapped classifiers to large survey data, and uses convolutional neural networks on three different 2D representations designed to expose profile asymmetry. They test the pipeline on 318 isolated-galaxy profiles whose labels come from a previous analytical classification and on 30,780 profiles from a large survey. A sympathetic reader would care because the approach is a candidate for handling the millions of HI profiles expected from next-generation radio surveys.","feed_headline":"Turning 1D galaxy spectra into 2D images lifts classification by 13%","feed_subtitle":"CNN on rotated, subtracted HI profiles beats 1D methods and scales toward SKA-era surveys.","key_machinery":"The load-bearing mechanism is a set of three 2D transforms of each 1D profile: Model 1 rotates the original spectrum into an image; Model 2 splits the profile at its centre, reflects one side, subtracts it from the other, and rotates the difference; Model 3 is a normalized version of Model 2 with pixel intensities scaled. Feeding these images to a CNN is what produces the 13% improvement. The rest of the pipeline, including Busy-function fitting, iterative polynomial, Gaussian, and double-Lorentzian models, and clustering plus KNN, SVM, and Random Forest bootstrapping, supports the comparison but does not by itself carry the asymmetry gain.","core_discovery":"The paper's central claim is that representing each one-dimensional HI spectrum as a two-dimensional image exposes asymmetries that a CNN can exploit, and that this representation raises classification accuracy by 13% over the same task done on 1D profiles. On the 318-profile sample the best 2D pipeline reaches 63% agreement with the analytical classification used as reference. The claim is therefore not about physical truth of asymmetry in an absolute sense; it is that the 2D transformation makes machine classification track that reference scheme better than 1D input does.","pith_inferences":["The 13% gain is measured against one analytical labelling scheme; whether the 2D models capture physically real asymmetries could be tested by comparing their output with metrics from resolved HI velocity fields or simulated galaxies.","Because the transforms are generic, rotation and mirrored subtraction, the same recipe could be tried on other symmetric one-dimensional spectra, such as optical emission lines or molecular line profiles.","The per-iteration cost of about 1.13 hours on modest hardware suggests that scaling to surveys with millions of profiles will need lighter CNN architectures or pre-selection, a practical step the paper leaves open."],"forward_implications":["The 2D image representation can be applied to any future HI survey without waiting for new instruments, since it only transforms existing spectra.","The reported 63% accuracy sets a concrete baseline for classifying isolated-galaxy HI profiles against the analytical reference scheme.","The framework, including code and models, is openly documented so the 13% gain can be reproduced and stress-tested by other groups.","The same workflow is positioned to scale to millions of profiles expected from next-generation radio surveys, where manual or analytical classification will be impractical."],"supporting_citations":[{"why":"Supplies the analytical profile classification used as the ground-truth reference for the 318-profile comparison.","marker":"Espada et al. 2011"},{"why":"Provides the Busy function package used to fit and pre-process the HI spectral profiles.","marker":"Westmeier et al. 2014"},{"why":"Supplies the ALFALFA catalogue of 30,780 profiles used for the large-scale classification experiments.","marker":"Haynes et al. 2018"},{"why":"Defines the Catalogue of Isolated Galaxies from which the 318 profiles are drawn.","marker":"Karachentseva et al. 1986"}],"fun_headline_variants":["2D HI profile images boost galaxy classification by 13%","CNN on 2D spectra lifts HI classification accuracy 13%","Galaxy HI profiles: 2D images add 13% classification boost","Turning 1D HI spectra into 2D images ups classification 13%","2D image trick for HI profiles raises accuracy 13%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation assumes the earlier analytical classification of the 318 profiles is the correct answer; if those labels are biased or contain errors, the 63% accuracy and the 13% advantage of 2D over 1D are only measures of agreement with that scheme.","fun_headline_variants_meta":{"raw":{"variants":["2D HI profile images boost galaxy classification by 13%","CNN on 2D spectra lifts HI classification accuracy 13%","Galaxy HI profiles: 2D images add 13% classification boost","Turning 1D HI spectra into 2D images ups classification 13%","2D image trick for HI profiles raises accuracy 13%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000231,"raw_usage":{"total_tokens":1504,"prompt_tokens":980,"completion_tokens":524,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":596,"completion_tokens_details":{"reasoning_tokens":430}},"tokens_in":596,"tokens_out":524,"duration_ms":5715,"temperature":1.0,"reasoning_tokens":430,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T17:59:14.693271+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the released pipeline on the same 318 profiles with the analytical labels replaced by blind, independent expert classifications or by labels derived from resolved HI kinematics; if the 2D CNN no longer beats the 1D CNN by roughly 13%, the claimed advantage is an artifact of the reference labels rather than a property of the 2D representation.","supporting_citations":[{"cited_title":"The AMIGA sample of isolated galaxies: VIII. The rate of asymmetric HI profiles in spiral galaxies","cited_arxiv_id":"1107.0601","evidence_quote":"Supplies the analytical profile classification used as the ground-truth reference for the 318-profile comparison."},{"cited_title":"E., Lebedev , V","cited_arxiv_id":null,"evidence_quote":"Defines the Catalogue of Isolated Galaxies from which the 318 profiles are drawn."}],"review_version":1}