{"id":"ab38f8c9-e7bc-46a3-95d5-d1f43990d327","arxiv_id":"1908.07542","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Self-organizing maps applied to nonparametric variability estimators identify variable AGN light curves with 86% purity and 66% completeness, comparable to deep learning.","lead":"The authors test whether an unsupervised machine-learning method, self-organizing maps, can find galaxies whose brightness varies over time due to active galactic nuclei. On roughly 8,300 WISE-selected AGN candidates, it recovers visually classified variable light curves with 86% purity and 66% completeness, similar to supervised deep learning.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Visual-classification ground truth is unvalidated, so reported purity/completeness may measure agreement with human labels rather than true AGN variability; an inter-rater reliability check is needed.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the visual classification is used as ground truth for both training and evaluation, and no quantitative reliability check is provided. I agree that this is the most important gap because every performance number in the paper is a comparison to those labels. The SOM, the deep-learning baseline, and the reported purity/completeness all inherit the properties of the human labels; if the labels are noisy or biased, the central claim about finding variable AGN is correspondingly weakened, even though the abstract carefully says \"variable classified AGN.\" The simulation test provides useful evidence that the method can separate flat and sinusoidal light curves in the presence of noise, but it does not establish that the human labels on real WISE light curves are trustworthy. The paper is otherwise transparent about its methods and the comparison to deep learning is informative, so the appropriate verdict remains conditional rather than a rejection. No additional concern is strong enough to change the reader's verdict; the proposed inter-rater reliability check would settle whether the visual-classification assumption actually holds.","tokens_in":9754,"tokens_out":5533,"duration_ms":564352,"concrete_test":"Have at least two additional human classifiers independently re-label a random subset of ~300 Stripe 82 light curves (stratified by SOM cell and variability class) using a pre-specified written protocol, then compute Cohen's kappa between raters and re-evaluate the SOM predictions against the consensus (majority) labels on that subset. If kappa < 0.6, or if purity and completeness against consensus differ from 86% and 66% by more than ~5 percentage points, the headline numbers are not robust to label noise and should be reported as measuring agreement with visual labels only. A complementary check is to compare the variable class against an objective variability threshold (e.g., chi-squared significance or excess variance) to see whether disagreements concentrate near the classification boundary.","verdict_should_be":"UNCHANGED","load_bearing_attack":"All headline performance numbers (Eq. 8: purity 86%, completeness 66%) are evaluated against a single human visual classification of 8,309 light curves into 7,558 nonvariable and 751 variable objects. The paper gives no inter-rater reliability, no independent confirmation, and no quantitative definition of what counts as \"visually variable.\" If the human labels are noisy or systematically biased (e.g., toward smooth, monotonic, or large-amplitude changes and against low-amplitude stochastic variability), then the measured purity and completeness describe agreement with those labels, not detection of physical AGN variability. This is not merely a philosophical limitation: the SOM is trained, thresholded, and evaluated on estimators derived from the same light curves, so any systematic label error propagates directly into C_SOM. The simulation test (Eq. 7) does not close this gap, because it uses idealized sinusoidal/flat light curves with known truth rather than realistic AGN power spectra and cannot quantify label reliability on real data. A secondary concern is that the 50% cell-fraction threshold is described as chosen to maximize metrics, and the clipping of 100 training objects is retained because it \"improves performance slightly\"; without a separate validation set it is unclear whether the reported numbers are unbiased generalization estimates or are partially tuned on the evaluation data.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an unsupervised machine-learning pipeline for identifying variable AGN light curves. The authors define eight nonparametric and two parametric variability estimators, train a self-organizing map (SOM) on an 80% subsample of 8,309 WISE-selected AGN from Stripe 82, and evaluate on the remaining 20%, using a single human visual classification (7,558 nonvariable, 751 variable) as ground truth. On simulated light curves with realistic noise and sampling, the SOM achieves purity 91% and completeness 79%; on the observed sample it is reported to achieve purity 86% and completeness 66% (Eq. 8) with ACC = 0.94, MCC = 0.72, F1 = 0.75. The SOM is compared with a supervised multilayer perceptron, which yields purity 79% and completeness 58%, and the authors conclude that the SOM is comparable while additionally providing visualization of estimator correlations.","tokens_in":10034,"tokens_out":9940,"duration_ms":90260,"significance":"If the reported performance is unbiased, the paper provides a useful, domain-aware alternative to supervised deep learning for variability selection, with fast classification and the ability to visualize correlations on the SOM. The train/test split, the noise-matched simulations, and the comparison to a deep network are strengths, and the use of standard public libraries improves reproducibility. However, the headline numbers are only as good as the visual labels used for training and evaluation, so the current contribution is best read as reproducing a human visual classification; the paper's broader physical claim requires independent validation of those labels.","major_comments":[{"comment":"The quoted metrics do not follow from the printed confusion matrix. With C_SOM = ((0.85, 0.01), (0.04, 0.09)), purity is TP/(TP+FP) = 0.09/(0.09+0.01) = 0.90 and completeness is TP/(TP+FN) = 0.09/0.13 = 0.69, not 86% and 66%; MCC also evaluates to about 0.76, not 0.72. The same discrepancy occurs in Eq. (7), Eq. (9), and Eq. (10). Since these numbers are the central quantitative claim, the matrices and quoted metrics must be reconciled (or exact, unrounded values and a rounding policy provided) before the performance can be assessed.","section":"Section 3.2.2, Eq. (8)"},{"comment":"The ground truth is a single human visual classification of 7,558 nonvariable and 751 variable light curves, with no inter-rater reliability check, no quantitative definition of 'visually variable,' and no independent confirmation. Because the SOM is trained, thresholded, and evaluated against these labels, Eq. (8) measures agreement with this one human classification rather than necessarily detecting physical AGN variability. The simulation test in Section 3.2.1 uses sinusoidal and flat light curves and therefore cannot diagnose systematic biases in the human labels. The authors should either provide an inter-rater reliability assessment, validate against an independent variability indicator, or explicitly restrict the claim to reproducing the visual classification.","section":"Section 2.1 and Section 3.2.2"},{"comment":"The 50% per-cell fraction is described as chosen to 'maximize the metrics,' and the clipping of 100 training objects is retained because it 'improves the performance slightly.' If these choices are made using the held-out test-set metrics, the reported ACC, MCC, purity, and completeness in Eqs. (7) and (8) are not unbiased generalization estimates. A separate validation set should be used for threshold and preprocessing choices, or the sensitivity of the metrics to these choices should be quantified and reported.","section":"Section 3.2.1 and Section 3.2.2"}],"minor_comments":[{"comment":"In the second paragraph, 'enables the study the formation' should be 'enables the study of the formation,' and 'occurance' should be 'occurrence.'","section":"Introduction"},{"comment":"The optimizer name 'adams' should be 'Adam.'","section":"Section 3.3"},{"comment":"No uncertainty is quoted for the observed purity and completeness; the authors report ±0.01 variation for simulations and should state whether the observed values come from a single SOM run and provide a variance estimate.","section":"Section 3.2.2"},{"comment":"The parametric GP estimators are stated to be non-repeatable in about 5% of cases, but the paper does not say whether the variance-estimator results in Section 3.2.2 are from single runs or averaged, which limits reproducibility.","section":"Section 2.2"}],"recommendation":"major_revision","confidential_remarks":"The internal inconsistency between the printed confusion matrices and the quoted purity/completeness values is the most serious issue; if it is a typographical error, it can be fixed, but the threshold-selection and label-reliability concerns remain. The core methodology is sound and the simulations are valuable, so I would not reject the manuscript. The revision should be manageable within the scope of a Letters paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I read Faisst et al. 1908.07542. It's a competent, clearly written proof of concept for using self-organizing maps on nonparametric variability estimators to flag variable AGN in WISE/NEOWISE light curves. The headline numbers—86% purity, 66% completeness on ~8,300 real AGN, with a plain MLP giving 79% and 58% under the same conditions—are plausible, and the direct comparison to supervised deep learning is genuinely informative. The SOM also gives you a map of how the estimators behave, which is a nice interpretability bonus, and the simulation test with flat and sinusoidal curves under realistic noise and time sampling is a reasonable sanity check.\n\nThe main soft spot is exactly what the stress test says: the ground truth is one human's visual classification of the light curves, with no inter-rater check, no quantitative definition of \"variable,\" and no independent confirmation. So the observed purity and completeness should be read as agreement with those labels, not as a validated measurement of physical AGN variability. The simulation doesn't close that gap, because sinusoids and flat curves are not realistic AGN power spectra. On top of that, the 50% per-cell threshold is chosen to maximize the metrics, and 100 training objects are clipped because it \"improves performance slightly\"; without a separate validation set, the real-data numbers likely carry a bit of tuning optimism. These are real caveats, but they don't sink the paper. The central claim—SOMs can separate variable from nonvariable light curves with these estimators about as well as a small neural net while keeping the visualization—holds up. The paper is honest about being a proof of concept, and the methods are described well enough to reproduce.\n\nI'd send this to a serious referee. The issues are about validation and framing, not a load-bearing flaw. If I worked on time-domain surveys, I'd cite it as a useful example of unsupervised variability selection.","headline":"A solid proof of concept for SOM-based AGN variability selection, but the headline numbers rest on unvalidated visual labels and threshold tuning, so they should be read as agreement with human classification.","tokens_in":10522,"tokens_out":2663,"would_cite":true,"duration_ms":206592,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An unsupervised self-organizing map identifies visually classified variable AGN in WISE light curves with 86% purity and 66% completeness, matching supervised deep learning.","keywords":["active galactic nuclei","AGN variability","self-organizing maps","machine learning","light curves","time-domain astronomy","WISE","unsupervised classification"],"falsifier":"Have two or more independent observers, blind to the original labels, reclassify a random subset of the 8,309 WISE light curves, then retrain and evaluate the SOM against the consensus labels; if purity and completeness fall substantially below 86% and 66%, part of the reported performance is agreement with label noise. A complementary check would inject simulated variable light curves with known amplitudes into the real sample and measure how recovery depends on amplitude.","tokens_in":9599,"feed_emoji":"🔭","tokens_out":12031,"duration_ms":101646,"temperature":0.7,"pith_summary":"This paper argues that an unsupervised machine-learning method, the self-organizing map, can pick variable active galactic nuclei — galaxies whose bright cores flicker as their central black holes accrete matter — out of roughly 8,300 infrared light curves about as well as a supervised deep-learning network. The map is trained on eight nonparametric measures of variability and, on the observed WISE sample, flags objects with 86% purity (most of what it flags really is variable) while recovering 66% of the visually identified variables. The authors first validate the pipeline on simulated light curves with realistic noise and sampling, reaching 91% purity and 79% completeness. Their broader point is that the map doubles as a visualization of the data, showing which variability statistics separate variable from nonvariable objects and where noise mimics variability, in a form that scales to the large time-domain datasets of upcoming surveys.","feed_headline":"Self-organizing maps find variable AGN at 86% purity","feed_subtitle":"A 30x30 map of eight variability statistics matches deep-learning results and shows where variability lives in the data.","key_machinery":"The engine is the self-organizing map, an unsupervised algorithm that folds an N-dimensional feature space into a 30×30 grid of cells while preserving neighborhoods: nearby cells contain light curves with similar variability statistics. The features are eight nonparametric variability estimators — χ², standard deviation, median absolute deviation, interquartile range, robust median statistics, normalized excess variance, peak-to-peak amplitude, and the inverse von Neumann ratio (the ratio that uses correlations between consecutive points, high for smooth trends and low for short-timescale jitter) — plus optional Gaussian-process variance and length-scale parameters. Each cell is labeled variable if more than half of the training light curves mapped into it are variable, and any new light curve is classified instantly by landing in a cell. The cell structure carries the argument: it separates variable from nonvariable objects spatially, and its per-cell median estimator values reveal which statistics track variability and where low signal-to-noise degeneracies hide.","core_discovery":"The central claim is that a self-organizing map trained on nonparametric variability estimators identifies visually classified variable AGN light curves in a WISE-selected Stripe 82 sample about as well as supervised deep learning, while keeping the data structure visible. Applied to 8,309 AGN candidates, 751 of which were visually flagged as variable, the SOM yields 86% purity and 66% completeness (accuracy 0.94, F1 0.75, Matthews correlation coefficient 0.72); the multilayer-perceptron comparison gives 79% purity and 58% completeness (MCC 0.65). The authors first test the method on simulated light curves with realistic noise and time sampling, where it reaches 91% purity and 79% completeness, then show that on real data the variable curves cluster in a compact region of the map. They also report that separating the three visual subclasses is not robust, a difficulty they attribute to the small size of the monotonic subgroups (66 increasing and 98 decreasing out of 8,309 objects), and that χ², robust median statistics, and the inverse von Neumann ratio carry most of the discriminative signal, while median absolute deviation, interquartile range, and normalized excess variance show degeneracies at low signal-to-noise.","pith_inferences":["Because the evaluation labels come from a single visual inspection, the reported 86% and 66% are best read as agreement with those human labels; an independent relabeling study would reveal how much of the apparent performance is shared label noise.","The estimator maps suggest that χ², RoMS, and the inverse von Neumann ratio dominate the signal while MAD, IQR, and excess variance are degenerate at low signal-to-noise; replacing them with uncertainty-aware estimators could plausibly raise completeness above 66% without sacrificing much purity.","The per-cell variable fraction is effectively a continuous score, so a survey could threshold it at values other than 50% to trade purity against completeness depending on whether it needs a clean sample or a complete census.","The same machinery could be extended to multi-band light curves, where the map topology would show whether variability in different wavelengths traces the same cells or separates by physical mechanism."],"forward_implications":["The same trained map can classify new light curves instantly without retraining, making it practical for real-time filtering in large time-domain surveys.","Because it consumes generic variability estimators rather than AGN-specific features, the method is claimed to transfer to supernovae, exoplanet transits, pulsars, and other time-sampled transients.","The map's cell layout is a built-in diagnostic: per-cell estimator maps expose which variability indicators are reliable and where photometric noise mimics variability, information a deep network does not expose.","The SOM organizes data without requiring complete labels, so sparse visual classifications can be layered on top of the map rather than used to train a supervised model from scratch.","On the WISE sample, a variable-AGN catalog selected at the SOM's 86% purity would contain roughly one nonvariable interloper for every six variable objects, a contamination level the authors treat as acceptable for statistical studies."],"supporting_citations":[{"why":"Defines the self-organizing map algorithm that the paper trains on variability estimators.","marker":"Kohonen 1982, 1990"},{"why":"Supplies the definitions of the nonparametric variability estimators used as features.","marker":"Sokolovsky et al. (2017)"},{"why":"Describes construction of the WISE AGN candidate sample and the light curves used here.","marker":"Prakash et al. (2019)"},{"why":"Gives the W1−W2 color cut that defines the AGN candidates.","marker":"Stern et al. (2012)"},{"why":"Provides the Stripe 82 AGN catalog from which the color-selected candidates are drawn.","marker":"Jiang et al. (2014)"},{"why":"The SOM training setup follows this earlier photometric-redshift application of self-organizing maps.","marker":"Masters et al. (2015)"},{"why":"Supplies the self-organizing-map implementation used for training and classification.","marker":"Hanke et al. (2009)"},{"why":"Defines WISE and its W1 band, the photometry source for the light curves.","marker":"Wright et al. (2010)"}],"fun_headline_variants":["Self-organizing map matches deep learning for variable AGN","Unsupervised SOM hunts variable AGN with 86% purity","Mapping variable AGN: SOM rivals neural nets at 86% purity","Self-organizing maps reveal variable AGN in Stripe 82"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the visual inspection that sorted 7,558 light curves as nonvariable and 751 as variable provides reliable ground truth; if that labeling is noisy or biased, the purity and completeness numbers measure agreement with the labels rather than real variability detection.","fun_headline_variants_meta":{"raw":{"variants":["Self-organizing map matches deep learning for variable AGN","Unsupervised SOM hunts variable AGN with 86% purity","Mapping variable AGN: SOM rivals neural nets at 86% purity","Self-organizing maps reveal variable AGN in Stripe 82"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000196,"raw_usage":{"total_tokens":1428,"prompt_tokens":1079,"completion_tokens":349,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":695,"completion_tokens_details":{"reasoning_tokens":274}},"tokens_in":695,"tokens_out":349,"duration_ms":3900,"temperature":1.0,"reasoning_tokens":274,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:03:52.197548+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have two or more independent observers, blind to the original labels, reclassify a random subset of the 8,309 WISE light curves, then retrain and evaluate the SOM against the consensus labels; if purity and completeness fall substantially below 86% and 66%, part of the reported performance is agreement with label noise. A complementary check would inject simulated variable light curves with known amplitudes into the real sample and measure how recovery depends on amplitude.","supporting_citations":[{"cited_title":"1982, Biological Cybernetics, 43, 59","cited_arxiv_id":null,"evidence_quote":"Defines the self-organizing map algorithm that the paper trains on variability estimators."},{"cited_title":"V., Gavras, P., Karampelas, A., et al","cited_arxiv_id":null,"evidence_quote":"Supplies the definitions of the nonparametric variability estimators used as features."},{"cited_title":"A Flaring AGN In a ULIRG candidate in Stripe 82","cited_arxiv_id":"1908.04280","evidence_quote":"Describes construction of the WISE AGN candidate sample and the light curves used here."},{"cited_title":"O., Sederberg, P","cited_arxiv_id":null,"evidence_quote":"Supplies the self-organizing-map implementation used for training and classification."}],"review_version":1}