{"id":"b494d18f-a955-4eb4-a22c-bfd4174b459c","arxiv_id":"2505.14275","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A supervised neural network classifier for solar Stokes V profile shapes is introduced, tested on DKIST, Hinode, GREGOR, and simulations, and used for the first statistical DKIST/ViSP quiet Sun inversion analysis.","lead":"The authors trained a neural network to sort shapes of circularly polarized light signals from the Sun, then used it to compare quiet Sun observations from three telescopes and a simulated sunspot. The tool gives solar physicists a reproducible way to count profile types, and it surfaces a disagreement between the MURaM sunspot model and real observations about where reversed magnetic fields appear.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Training and test pools are drawn only from the 250 profiles nearest each k-means centroid, so the MLP is never validated on the peripheral profiles shown in Figure 1; if those profiles are classified differently across datasets, the reported cross-dataset statistics are biased.","rationale":"The paper's headline contribution is a supervised ML classifier that enables 'detailed statistical analyses' of Stokes V profiles across DKIST, Hinode, GREGOR, and simulations. The entire statistical edifice—class percentages in Figure 2, the DKIST vs Hinode similarity claim, the GREGOR differences, and the sunspot RPMF fractions in Figure 11—is computed by deploying the MLP to every pixel. But the model is trained and evaluated exclusively on profiles that are nearest neighbours of k-means centroids (Section 3.3). The test metrics in Table 3 (e.g., DKIST test f1 = 0.91) therefore only certify performance on these cluster cores. Figure 1 shows the population's periphery: the 50 furthest profiles per cluster frequently have extra or missing lobes and are not represented in the training pool. If such peripheral profiles are misclassified at a different rate in different datasets, every cross-dataset comparison in the paper inherits a dataset-dependent bias. For example, the surprising similarity between DKIST and Hinode class statistics, despite modelled resolution differences, could partly reflect both classifiers being accurate on cores but biased on peripheries in similar ways, or cancelling errors. Conversely, the GREGOR differences could be inflated if the NIR line produces a different periphery distribution. The authors note the possibility of poor generalisation in Section 5, but do not test it. A targeted periphery test would either confirm the claim or show that the reported statistics are pool-dependent. This is the most load-bearing concern because it affects the central claim and all downstream comparisons; the other limitations (single MURaM snapshot, no PSF, arbitrary thresholds) are acknowledged and would not invalidate the method itself if the classifier generalised.","tokens_in":23256,"tokens_out":4834,"duration_ms":46324,"concrete_test":"Take the 50 furthest profiles in each of the 35 DKIST clusters (the blue dotted profiles in Figure 1), have them manually labelled with the same threshold rules, and run the trained DKIST MLP on them; compare the resulting f1-score with the Table 3 test f1 of 0.91. If the periphery f1 drops below about 0.8, or if the same procedure for Hinode/GREGOR yields different drops, the reported cross-dataset class percentages are not reliable for peripheral profiles and the 'robustly classifies' claim needs qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the supervised classifier 'robustly classifies' solar spectropolarimetric data and enables statistical comparisons is only supported for cluster-core profiles. In Section 3.3, training, validation, and test sets are created by selecting the 250 closest Stokes V profiles to each of the 35 k-means++ centroids (8750 profiles for DKIST). The validation and test metrics in Table 3 therefore measure performance on this restricted pool, not on the full population. Figure 1 explicitly shows that the furthest 50 profiles in each cluster differ from the centroid, often with extra or missing lobes, and these are excluded from training. If the fraction or morphology of such peripheral profiles differs between DKIST, Hinode, and GREGOR (or between quiet Sun and sunspot data), the class fractions in Figures 2 and 11 and the cross-dataset comparisons would be biased in a way that the validation metrics cannot detect. The authors acknowledge 'lack of proper generalisation to unseen data' in Section 5, but no test quantifies this. Without a held-out set drawn from outside the cluster-core pool, the claimed robustness is unverified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper trains a multi-layer perceptron (MLP) to classify the morphology of Stokes V profiles, using four quiet-Sun classes (symmetric, asymmetric, single-lobed, Q-like) and five sunspot classes (positive, negative, double positive, mixed polarity, and a collapsed negative/double-negative class). The classifier is applied to quiet-Sun observations from DKIST/ViSP, Hinode/SP, and GREGOR/GRIS-IFU, as well as to synthetic observations from MANCHA and MURaM simulations. The authors report validation and test metrics for every model (Table 3), present the first statistical analysis of quiet-Sun DKIST/ViSP data using inversions and a supervised classifier, compare supervised and unsupervised (k-means) classification, and examine the occurrence of reverse-polarity magnetic fields in a simulated sunspot. The central claims are that supervised ML robustly classifies solar spectropolarimetric data, that k-means centroid labeling introduces systematic errors that can compromise statistical comparisons, and that in the MURaM sunspot simulation the 1564.85 nm line detects more reverse-polarity fields in the penumbra than the 630.25 nm line, in contradiction to observations.","tokens_in":23454,"tokens_out":6328,"duration_ms":57322,"significance":"If the central claims hold, the trained classifier is a reusable, open-source tool for large-scale Stokes-profile classification, and the sunspot line-dependent discrepancy gives modelers a concrete target. Strengths of the paper include explicit validation and test metrics for every model, controlled re-synthesis that isolates spectral-line effects from spatial-resolution effects, a transparent k-means versus MLP comparison, and use of the MURaM simulation's ground-truth magnetic field to test the physical interpretation of the classifier output. The main weakness is that the training, validation, and test pools are constructed only from profiles closest to k-means centroids, so the reported metrics do not establish generalization to the full data population. In addition, some validation checks are logically circular because they use the same amplitude thresholds that defined the labels. The paper is honest about several of these limitations, but they are not quantified or resolved.","major_comments":[{"comment":"The labelled pool for each quiet Sun dataset is built from the 250 profiles closest to each of the 35 k-means++ centroids, and the training, validation, and test splits in Table 3 are drawn exclusively from this pool. The test metrics therefore measure performance only on cluster-core profiles. Figure 1 shows that the furthest 50 profiles in each cluster often have additional or missing lobes, and these peripheral profiles are never seen by the classifier. If the fraction of such peripheral profiles differs between DKIST, Hinode, and GREGOR, or between quiet Sun and sunspot data, the class fractions in Figures 2 and 11 and the cross-dataset comparisons would be biased in a way that the reported accuracies cannot detect. The sentence in Section 5 acknowledging \"lack of proper generalisation to unseen data\" identifies this issue but does not quantify it. Please add a held-out test set drawn from the full population (e.g., random profiles outside the cluster-core pool) and report metrics on that set, or explicitly restrict the paper's robustness claims to the cluster-core population.","section":"§3.3, Fig. 1"},{"comment":"The amplitude-asymmetry analysis is presented as \"empirical evidence that the MLP classifier has successfully distinguished between these morphological types based on asymmetry.\" However, the class definitions in §4.3 already impose that a symmetric profile has a subordinate lobe reaching at least 0.9 of the dominant lobe, while an asymmetric profile has a subordinate lobe below that threshold. Consequently, δa is near zero for the symmetric class and non-zero for the asymmetric class essentially by construction, so Figure 7 does not provide independent validation. To make this point load-bearing, the authors should use an asymmetry measure not directly tied to the labelling thresholds (e.g., area asymmetry or lobe-separation asymmetry) or compare distributions on a held-out sample with human labels that were not used in setting the 0.9 and 0.25 thresholds.","section":"§4.6, Eq. (2)"},{"comment":"The sunspot classifier's \"negative\" class is defined to include both simple negative and negative double profiles because the MLP could not be trained to separate them. The central sunspot claim—that the 1564.85 nm line detects more reverse-polarity fields in the penumbra than the 630.25 nm line—rests on the relative counts of this class. If negative double profiles are more common in one line than the other, their inclusion in the same class biases the ratio, and because double profiles are associated at least in part with magneto-optical effects rather than genuine polarity reversals, the comparison with observed RPMF detection rates becomes ambiguous. Please report the proportion of manually labelled negative profiles that are actually double profiles for each spectral line, and discuss how this class impurity affects the comparison with Franz et al. (2016) and other observations.","section":"§4.7, Fig. 11"},{"comment":"The text states \"We find no statistically significant differences in magnetic flux density across classes; the distributions are broadly similar,\" but no statistical test is described or reported. Given that the later discussion in Section 5 contrasts the lack of differences in magnetic flux density with clear trends in inclination angle, the absence of any significance testing weakens the claim. Please either provide appropriate statistical tests (e.g., Kruskal-Wallis, bootstrap confidence intervals) or rephrase the statement to report the observed distributional overlap without asserting statistical significance.","section":"§4.4, Fig. 6"}],"minor_comments":[{"comment":"There are minor typographical issues in the title and abstract: the title contains \"F rom\" with an extra space, and the abstract uses an unmatched quotation mark in \"double’ profiles.\"","section":"Abstract and title"},{"comment":"The hyperparameter table lists a batch size of 125 for DKIST; if this is not a typo for 128, please clarify, since all other batch sizes are powers of two.","section":"§3.3, Table 3"},{"comment":"For the GREGOR dataset, the 60% and 23% fractions of circular and linear polarization are cited from a previous paper (Campbell et al. 2023a) rather than measured here; this should be stated more explicitly in the text so that the reader does not think all fractions come from the current analysis.","section":"§4.1"},{"comment":"The phrase \"lack of proper generalisation to unseen data\" in the Discussion is important, but it is not connected to any quantitative estimate of the generalization gap. This would be a natural place to mention the extent to which the held-out cluster-core test set overestimates performance on peripheral profiles.","section":"§5"}],"recommendation":"major_revision","confidential_remarks":"For the editor: The paper is a solid and transparent application of supervised ML to Stokes-profile classification, with open code and per-model validation metrics. The main concern is that the test data are not representative of the full population, which directly affects the central claim of 'robust' classification and the cross-dataset statistics. This is fixable by validating on a random sample of the full datasets. The amplitude-asymmetry validation is also circular as written. I recommend major revision rather than rejection because the core methodology is defensible and the identified issues can be addressed within the scope of the manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is the first supervised classification of Stokes V morphology, and it delivers the first DKIST/ViSP quiet-Sun inversion-based statistics. The MLP with code on GitHub is a genuinely reusable tool, and the k-means versus MLP comparison (Fig. 8) cleanly shows centroid labeling inflates symmetric counts. The MURaM sunspot result—1564.85 nm seeing more reverse polarity than 630.25 nm, opposite to observations—is a clear, testable target for modelers. Credit where due: validation/test metrics are reported for every model in Table 3, the sunspot magneto-optical analysis uses controlled re-synthesis, and the authors are transparent about limits in Section 5.\n\nThe soft spots are real but proportionate. The training pool is drawn only from the 250 profiles nearest each k-means centroid, and the validation/test sets come from the same pool. So the roughly 0.9 f1-scores do not measure performance on the peripheral profiles shown in Figure 1, which often have extra or missing lobes. If the fraction of such profiles differs between DKIST, Hinode, and GREGOR, the cross-dataset statistics in Figure 2 are biased in a way the metrics cannot detect. The authors admit \"lack of proper generalisation to unseen data\" in Section 5 but never quantify it. A held-out set from outside the cluster cores would settle it. Minor issues: the 0.25/0.9 label thresholds are arbitrary, no error bars appear on class fractions, and the sunspot conclusions rest on one MURaM snapshot with no PSF or noise. The test f1 for MURaM drops to 0.78, so \"robustly classifies\" in the abstract is stronger than what is shown.\n\nOverall, the central claim is defensible with caveats, the method is a contribution, and the flaws are limitations rather than fatal errors. This deserves serious peer review. If I were the editor, I would send it out with a request to add an out-of-pool test set or visibly soften the generalization claims.","headline":"First supervised Stokes-V classifier plus first DKIST/ViSP quiet-Sun statistics; useful and honest, but cluster-core training pools leave generalization claims unverified.","tokens_in":24038,"tokens_out":1970,"would_cite":true,"duration_ms":18650,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A supervised network classifies solar Stokes V profiles with accuracy close to or above 90%, enabling the first quiet-Sun DKIST/ViSP analysis and exposing a sunspot-simulation contradiction over which line detects reverse-polarity fields.","keywords":["Stokes V profiles","spectropolarimetry","multi-layer perceptron","k-means clustering","quiet Sun","sunspot penumbra","reverse polarity magnetic fields","supervised machine learning"],"falsifier":"Apply the paper's own labelling thresholds (subordinate lobe at least 0.9 of the main lobe for symmetric, at most 0.25 for single-lobed) automatically to every profile in each dataset, not just to cluster centroids, and compare the resulting class fractions with the MLP's output; if the disagreement approaches the size of the k-means-versus-MLP disagreement the paper documents, the claimed robustness of the supervised statistics is not established. A second check: retrain the classifier on pools that deliberately include the furthest profiles from each cluster and confirm that the reported cross-dataset differences (DKIST versus Hinode versus GREGOR) survive.","tokens_in":22975,"feed_emoji":"☀️","tokens_out":17598,"duration_ms":147874,"temperature":0.7,"pith_summary":"The paper sets out to establish that supervised machine learning—a multi-layer perceptron trained on manually labelled Stokes V profiles—can classify the shapes of solar circular-polarisation signals into morphological classes with validation and test metrics close to or above 90%, and that this is reliable enough to support statistical comparisons between telescopes and against simulations. It applies the classifier to quiet-Sun data from DKIST/ViSP, Hinode/SP, and GREGOR/GRIS-IFU, as well as synthetic observations from MANCHA granulation and MURaM sunspot simulations, and in doing so presents the first statistical analysis of quiet-Sun DKIST/ViSP data based on inversions and profile classification. The methodological core is the demonstration that k-means clustering, when used to label profiles by their centroid shape, introduces systematic errors—most notably too many 'symmetric' labels—that would compromise cross-dataset statistics, whereas the supervised approach keeps errors bounded and class-balanced. The notable results are that DKIST and Hinode quiet-Sun morphologies agree despite modelling that says spatial resolution should separate them; that the 1564.85 nm line produces more symmetric and far fewer single-lobed profiles, consistent with its narrower response functions; and that in the MURaM sunspot simulation the 1564.85 nm line detects more reverse-polarity penumbral fields than the 630.25 nm line, the opposite of observations.","feed_headline":"Sunspot simulation contradicts observations on reverse polarity fields","feed_subtitle":"A neural net sorts solar polarization profiles ~90% right and shows where the sunspot simulation disagrees with observations.","key_machinery":"The load-bearing object is the multi-layer perceptron classifier, a three-layer fully connected network with Swish activation, class-weighted cross-entropy loss, dropout, early stopping, and Xavier initialization, which maps each Stokes V profile to one of four quiet-Sun classes (asymmetric, symmetric, Q-like, single-lobed) or five sunspot classes (positive, negative, double positive, mixed polarity). Its training labels come from a two-stage pipeline: k-means++ clustering with 35 clusters selects the 250 profiles nearest each centroid, and a human labels those thousands of profiles using amplitude thresholds of 0.9 and 0.25 on subordinate lobes; the MLP then replaces the k-means centroids as the classifier. Two comparison devices carry the argument: Sankey diagrams that trace how profiles move between classes when the same MANCHA atmosphere is synthesised in 630.25 nm versus 1564.85 nm, and between k-means centroid labels and MLP labels; and the SIR inversion code, which supplies both the synthetic profiles and the physical parameters (field strength, inclination, velocity) that the paper uses to show the classes carry physical meaning.","core_discovery":"The central discovery, stated on the paper's own terms, is that supervised machine learning provides a robust, reproducible way to classify the morphology of solar circular-polarisation profiles where earlier work relied on manual inspection or on k-means clustering. A multi-layer perceptron with three fully connected layers, Swish activation, and class-weighted cross-entropy loss, trained on a few thousand manually labelled profiles drawn from the cores of k-means clusters, reaches validation and test accuracies and f1-scores typically close to or above 90% on quiet-Sun data from DKIST/ViSP, Hinode/SP, and GREGOR/GRIS-IFU, and on synthetic observations from MANCHA granulation and MURaM sunspot simulations. Deployed across these datasets, the classifier produces four principal findings: the two visible-line quiet-Sun datasets (DKIST and Hinode) classify similarly even though degraded MANCHA syntheses predict that spatial resolution should change profile shapes; the near-infrared 1564.85 nm line produces more symmetric and far fewer single-lobed profiles than the visible 630.25 nm line, in line with its narrower response functions; nearly a fifth of the simulated penumbra shows mixed-polarity profiles, of which 67–75% coincide with genuine line-of-sight polarity reversals while the 'double' profiles common in the 630.25 nm penumbra are mostly magneto-optical effects in strongly inclined fields; and the 1564.85 nm line detects more than twice as many reverse-polarity profiles as the 630.25 nm line in the penumbra, in direct contradiction to observations. The paper also shows that labelling profiles by k-means centroids alone systematically overestimates the symmetric class, an error that is not cured by increasing the number of clusters.","pith_inferences":["My inference: the DKIST/Hinode agreement is usable as a diagnostic — if the analysed DKIST scan was seeing-limited, a future good-seeing ViSP dataset run through the same classifier should drift toward the high-resolution MANCHA statistics (more single-lobed, fewer symmetric profiles), giving a morphology-based measure of DKIST's effective resolution.","My inference: because the cited response-function work gives the 525.02 nm line the strongest inclination sensitivity, synthesising the full MURaM snapshot in that line and classifying it with the same tool should maximise the double-profile fraction; if future high-resolution observations of that line do not show the predicted abundance, the magneto-optical interpretation would need revision.","My inference: the four quiet-Sun classes could be retrained on Stokes Q and U profiles as DKIST's polarimetric sensitivity grows, turning linear-polarisation morphology into a statistical map of horizontal fields — an extension the paper notes is not yet feasible at disk centre.","My inference: the Sankey-transfer analysis suggests a direct test of the line-difference interpretation — synthesising the same MANCHA cube at intermediate spectral lines such as 525.02 nm would show whether single-lobed and symmetric fractions vary monotonically with line-formation height, which would confirm that the response function, not the atmosphere, drives the differences."],"forward_implications":["The classifier is a reusable public tool: trained once and deployed to millions of pixels, it lets future DKIST observations be statistically compared with this quiet-Sun baseline without manual profile inspection.","The MURaM sunspot simulation fails a specific observational test: it predicts the 1564.85 nm line detects over twice as many reverse-polarity penumbral profiles as the 630.25 nm line, whereas observations show visible lines are far better at revealing three-lobed and reverse-polarity signatures; correcting this mismatch is a concrete target for simulation work.","Mixed-polarity profiles, about 18% of the simulated penumbra, can be read as tracers of genuine polarity reversals along the line of sight, while 'double' profiles in the 630.25 nm line trace magneto-optical effects in nearly horizontal fields; the paper shows inclination alone can create or erase the double shape.","For inversion work, the near-infrared line is better described by atmospheres with no gradients in optical depth, since its classification statistics and inferred parameters vary little with optical depth compared with the visible line.","Amplitude-asymmetry statistics become cleaner when computed only on the MLP-isolated symmetric and asymmetric classes, excluding Q-like and single-lobed profiles for which the asymmetry measure is not physically meaningful."],"supporting_citations":[{"why":"Pioneered k-means clustering of quiet-Sun Stokes V profiles; establishes the unsupervised approach whose centroid-labelling errors the paper demonstrates.","marker":"Khomenko et al. (2003)"},{"why":"Supplies the k-means recipe, the pre-processing scheme, and the label categories the paper adapts; the unsupervised baseline for the systematic-error comparison.","marker":"Viticchié & Sánchez Almeida (2011)"},{"why":"The MURaM sunspot simulation from which all sunspot Stokes profiles are synthesised; the model whose reverse-polarity line-dependence contradicts observations.","marker":"Rempel (2012)"},{"why":"Provides the manual three-lobed-profile detection method and the penumbral RPMF fractions (4%, up to 17% with three-lobed profiles) the sunspot statistics are compared with.","marker":"Franz & Schlichenmaier (2013)"},{"why":"The observational result that visible lines detect up to an order of magnitude more three-lobed profiles than the 1564.85 nm line — the finding the MURaM simulation contradicts.","marker":"Franz et al. (2016)"},{"why":"The SIR code used to synthesise the simulated Stokes vectors and to invert all three observed quiet-Sun datasets.","marker":"Ruiz Cobo & del Toro Iniesta (1992)"},{"why":"Provides the line response functions for 630.25, 1564.85, and 525.02 nm used to interpret the line-dependent classification statistics.","marker":"Quintero Noda et al. (2021)"},{"why":"The DKIST/ViSP dataset analysed; the paper presents the first statistical quiet-Sun analysis of this data.","marker":"Campbell et al. (2023b)"},{"why":"Argues spatial resolution and noise hide a significant fraction of RPMFs; used to frame why the simulation's NIR-line detection may not appear in real observations.","marker":"Bharti & Rempel (2019)"}],"fun_headline_variants":[],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole comparison rests on the assumption that the profiles nearest the cluster centres—the only ones used to train the classifier—represent the full population of profile shapes, even though the paper's own figures show that the most unusual profiles in each cluster look very different from the ones the classifier ever sees.","fun_headline_variants_meta":{"error":"Client error '402 Payment Required' for url 'https://api.deepseek.com/chat/completions'\nFor more information check: https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/402"},"cache_creation_input_tokens":0},"created_at":"2026-08-07T15:38:32.742922+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply the paper's own labelling thresholds (subordinate lobe at least 0.9 of the main lobe for symmetric, at most 0.25 for single-lobed) automatically to every profile in each dataset, not just to cluster centroids, and compare the resulting class fractions with the MLP's output; if the disagreement approaches the size of the k-means-versus-MLP disagreement the paper documents, the claimed robustness of the supervised statistics is not established. A second check: retrain the classifier on pools that deliberately include the furthest profiles from each cluster and confirm that the reported cross-dataset differences (DKIST versus Hinode versus GREGOR) survive.","supporting_citations":[{"cited_title":"2016, A&A, 596, A4, doi: 10.1051/0004-6361/201628407","cited_arxiv_id":null,"evidence_quote":"The observational result that visible lines detect up to an order of magnitude more three-lobed profiles than the 1564.85 nm line — the finding the MURaM simulation contradicts."},{"cited_title":"2019, ApJ, 884, 94, doi: 10.3847/1538-4357/ab3c6b","cited_arxiv_id":null,"evidence_quote":"Argues spatial resolution and noise hide a significant fraction of RPMFs; used to frame why the simulation's NIR-line detection may not appear in real observations."}],"review_version":1}