{"id":"df85a414-ba88-41ac-ba04-2a10117411b6","arxiv_id":"2607.13233","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A controlled cross-validation decomposition attributes the apparent X-ray-induced accuracy drop in AGN/SFG classification to sample selection and label noise, not to the X-ray feature itself.","lead":"This paper re-examines a puzzling drop in Random Forest accuracy when X-ray features were added to an AGN/star-forming galaxy classifier, and finds the drop is mostly a sample-selection artifact: once the sample is held fixed, the X-ray feature barely changes accuracy and is used in a slightly harmful way. It also shows that BPT-derived training labels are noisy near the classification boundary, exactly where X-rays might have helped.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Sample-selection attribution is unsupported: the paper's own decomposition accounts for only 3.52 of the 8.25 pp drop, leaving ~4.7 pp unexplained.","rationale":"Identified the most load-bearing concern as the quantitative inconsistency in the decomposition. The paper's central claim is that the previously reported decrease reflects predominantly sample selection. To support that, the decomposition must reproduce the original drop. It does not: 3.52 + (-0.01) = 3.51 pp vs. 8.25 pp. This is not a subtle rounding issue; four of eight percentage points are missing. The paper acknowledges the two accuracies were measured on different samples, but the controlled comparison uses samples that are themselves different from the originals (25,668 vs 28,708; 277 vs 312) and a different resampling scheme (5-fold vs 80-20). Without reconciling these, the attribution is an overclaim. The reader's weakest_assumption focused on whether the 35 excluded objects are boundary-concentrated; that is a related but narrower concern. Even if those objects are arbitrarily hard, the 312-to-277 exclusion occurred after sample selection from the full sample; the missing 4.7 pp likely arises from the full-sample side or the methodology difference. Our proposed test directly checks whether the decomposition reproduces the original drop under the original protocol. If it does, the concern is invalidated; if not, the paper should be revised to describe the previous decrease as not fully explained, with the feature effect shown to be negligible only on the fixed subsample. The controlled (b vs c) finding is independent and valuable, so the verdict should remain CONDITIONAL rather than REJECT.","tokens_in":11327,"tokens_out":4672,"duration_ms":44063,"concrete_test":"Recompute the three configurations using the exact sample definitions and preprocessing of Ding & Rodriguez (2024): all 28,708 full-sample objects and all 312 X-ray-detected objects, with the original missing-data handling, using the same 80-20 split. Compare the resulting sample effect (full optical-only vs X-ray subsample optical-only) and feature effect (X-ray subsample with/without X-ray) to the original 8.25 pp drop. If the sum of the two effects still deviates from 8.25 pp by more than 1 pp (beyond CV noise), the claim that sample selection 'predominantly' drives the drop is not supported. Additionally, report accuracy and BPT distances of the 35 excluded X-ray objects when scored by the (b) model to check whether their exclusion underestimates the sample effect.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"In §4, the paper reports a three-way decomposition: (a) full optical sample (25,668 objects, optical features) 97.37%, (b) X-ray subsample (277 objects, optical features) 93.85%, (c) same + X-ray flux 93.86%. The sample effect (a→b) is 3.52 pp and the feature effect (b→c) is -0.01 pp. Yet the prior accuracy drop it claims to explain is 8.25 pp (97.51→89.26). These numbers leave ~4.74 pp unaccounted for. The paper states 'the apparent accuracy decrease reported in Ding & Rodriguez (2024) therefore reflects predominantly sample selection,' but its own measured sample effect is less than half the total drop. Possible sources of the residual include the different sample definitions (28,708 vs 25,668; 312 vs 277), the switch from an 80-20 split to 5-fold CV, and unstated preprocessing changes. Because the residual is large and unreconciled, the central attribution is not quantitatively established. The controlled comparison (b vs c) remains credible, but the explanation of the original decrease is not. The reader's concern about the 35 excluded objects is related, but even substantial boundary concentration among those objects would not necessarily close a 4.7 pp gap given the 312-object sample size; the gap must be located explicitly.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper investigates a previously reported counterintuitive result from Ding & Rodriguez (2024), in which adding an X-ray flux feature to a Random Forest classifier of AGN and star-forming galaxies decreased accuracy from 97.51% to 89.26%. The authors propose two explanations: an astrophysical one based on luminosity overlap between AGN and SFGs around the canonical 10^42 erg/s threshold, and a machine-learning one based on instance-dependent label noise in BPT-derived training labels. Using a three-way cross-validation decomposition, they report that restricting to the X-ray-detected subsample lowers accuracy by 3.52 percentage points before any X-ray feature is added, while adding the X-ray feature changes accuracy by only -0.01 percentage points. They also report that misclassified objects cluster near the Kewley demarcation line and that the X-ray flux feature has a small but significantly negative permutation importance. The paper concludes that the apparent accuracy decrease reflects predominantly sample selection rather than the X-ray feature itself, and recommends independent labels or noise-robust training for future classifiers.","tokens_in":11620,"tokens_out":2649,"duration_ms":29916,"significance":"If the central claim were quantitatively supported, the paper would be a useful cautionary study for multi-wavelength machine-learning classification in astrophysics, showing that apparent feature-induced performance drops can be dominated by sample composition rather than feature content. The controlled fixed-sample comparison (configurations b vs c) is clean and credible, and the boundary-concentration analysis (KS D=0.745, p=1.67e-9) is a specific, falsifiable diagnostic that supports the label-noise narrative. The paper also makes a fair point that X-ray luminosity overlap is real in optically selected moderate-luminosity samples. However, the central attribution of the original 8.25-point drop to sample selection is not quantitatively established by the paper's own decomposition, which leaves a large residual unexplained. This is a load-bearing issue that must be resolved before the paper's headline conclusion can be accepted.","major_comments":[{"comment":"The paper's own numbers do not support the statement that the reported accuracy decrease 'reflects predominantly sample selection.' The original drop is 8.25 percentage points (97.51 to 89.26). The decomposition in Section 4 gives a sample effect of 3.52 pp (97.37 to 93.85 from configuration a to b) and a feature effect of -0.01 pp (b to c). This leaves roughly 4.7 pp of the original drop unaccounted for. Section 5 nevertheless concludes that the drop is 'driven predominantly by sample selection.' That conclusion is not supported by the measured sample effect, which is less than half of the total drop. Please locate the residual explicitly: possible sources include differences in sample construction (28,708 vs 25,668; 312 vs 277), the change from an 80-20 split to 5-fold CV, and any other preprocessing changes. Without reconciling this residual, the central attribution is not established","section":"Section 4, Section 5"},{"comment":"The sample effect is measured on the 277-object subsample with complete photometry, but the original X-ray model used 312 objects. The 35 excluded objects are dropped to avoid the classifier learning missingness patterns, but the paper never tests whether those objects are disproportionately concentrated near the BPT boundary or have systematically different X-ray properties. If the excluded objects are preferentially hard boundary cases, the measured sample effect could be underestimated. This concern is directly relevant to the decomposition's validity. Even if this does not fully close the ~4.7 pp residual identified above, it should be addressed with a sensitivity test or at least an explicit discussion of the excluded objects' properties.","section":"Section 4"},{"comment":"The Discussion draws a strong conclusion from the decomposition, stating that 'the apparent 8.25% accuracy drop reported in Ding & Rodriguez (2024) is driven predominantly by sample selection.' Given the residual gap quantified above, this is premature. The conclusion should be softened to what the controlled comparison actually shows: on a fixed sample, the X-ray feature has negligible effect on accuracy. The paper can still argue that sample selection plays a role, but the quantitative claim of predominance requires accounting for the missing ~4.7 pp.","section":"Section 5, 'It is worth noting...'"}],"minor_comments":[{"comment":"The statement 'k-corrections are small at these redshifts and are neglected' is unsupported. Since the X-ray band is 0.2-12 keV and the sample spans a range of redshifts, a brief justification or a quantification of the expected k-correction magnitude would strengthen the luminosity-overlap analysis.","section":"Section 3"},{"comment":"The text says the permutation importance confidence interval lies 'entirely below zero across 100 permutations,' but the error bars are described as one standard deviation. Please specify how the confidence interval is constructed (e.g., 1-sigma interval from the permutation distribution) so the reader can interpret 'significantly negative' precisely.","section":"Section 4, Figure 2"},{"comment":"The abstract states the decrease 'reflects predominantly sample selection,' but the Introduction already notes the two accuracies were measured on different samples. The wording in the abstract should be qualified to reflect the decomposition's actual coverage (3.52 pp of the 8.25 pp drop) until the residual is explained.","section":"Abstract and Section 1"}],"recommendation":"major_revision","confidential_remarks":"The controlled comparison (b vs c) is a solid contribution and could be publishable on its own. The paper's broader claim about the cause of the original accuracy drop needs to be reworked to match the quantitative decomposition. The referee report focuses on the numerical inconsistency and the excluded-object sensitivity, which are fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe useful core of this paper is the controlled decomposition in §4: on a fixed 277-object X-ray subsample, adding the X-ray flux changes accuracy by -0.01 pp (93.85 to 93.86). That is a clean, reproducible comparison and it convincingly kills the idea that X-rays actively hurt the classifier. The luminosity overlap numbers (41.6% of AGN below 10^42, 13.5% of SFGs above) are useful quantification for the SDSS regime, and the boundary-concentration diagnostics (median distance 0.123 vs 0.743 dex, KS D=0.745) are striking and well presented.\n\nBut the paper overreaches when it says the original 8.25 pp drop 'reflects predominantly sample selection.' Its own decomposition accounts for only 3.52 pp of that drop (a→b). The remaining ~4.7 pp is never reconciled. The full-sample accuracy in the 2024 paper was 97.51% on 28,708 objects with an 80-20 split; here configuration (a) is 97.37% on 25,668 objects with 5-fold CV. The difference between 28,708 and 25,668, plus the split change, plus the 312→277 object cut, could easily account for the rest, but the paper doesn't say. As written, the attribution is not quantitatively established.\n\nThe stress-test note is right: even if the excluded 35 objects are boundary-concentrated, that alone won't close a 4.7 pp gap on a 312-object sample. The gap needs to be located explicitly.\n\nThe label-noise interpretation is plausible and well motivated by the cited literature (Agostino & Salim, Birchall et al.), and the negative permutation importance is an interesting signal. But the paper itself notes the effect on accuracy is negligible, so the practical impact is modest: the prior result was mostly a sample artifact, and X-rays are not net harmful, just underused.\n\nWho gets value from this: anyone who uses the Ding & Rodriguez classifier, and anyone building multi-wavelength classifiers on optically-selected samples with sparse X-ray coverage. It's a careful correction to a published result, not a new technique.\n\nRecommendation: worth sending to peer review, with a request that the author reconcile the full 8.25 pp drop or soften the 'predominantly' claim. The controlled b vs c comparison stands regardless. The paper would be stronger with a direct account of the missing ~4.7 pp.","headline":"A clean controlled comparison shows X-rays don't hurt accuracy, but the paper's claim that sample selection explains the original 8.25-point drop is unsupported by its own numbers.","tokens_in":12079,"tokens_out":1911,"would_cite":true,"duration_ms":18837,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"X-ray data were not the cause of a galaxy classifier's accuracy drop; the harder X-ray-selected sample was.","keywords":["active galactic nuclei","star-forming galaxies","machine learning classification","label noise","X-ray luminosity","BPT diagram","random forest","permutation importance"],"falsifier":"Inspect the BPT positions of the 35 objects excluded for incomplete photometry: if their median distance to the Kewley demarcation curve is closer to 0.123 dex than to 0.743 dex, the sample-selection attribution would be weakened because the excluded objects would be disproportionately boundary-concentrated.","tokens_in":11189,"feed_emoji":"🔭","tokens_out":5562,"duration_ms":51212,"temperature":0.7,"pith_summary":"A previous experiment found that adding an X-ray flux feature to a Random Forest classifier of active galaxies and star-forming galaxies lowered accuracy from 97.51% to 89.26%, contradicting the idea that X-rays are a reliable AGN diagnostic. This paper argues that the decrease is mostly an artifact of comparing two different samples: the X-ray-detected subsample is harder to classify, and the feature itself changes accuracy by only -0.01 percentage points on a fixed sample. The deeper reason X-rays fail to help is that the training labels, derived from optical emission-line ratios on the BPT diagram (the standard optical line-ratio classification), are least reliable exactly at the boundary where X-rays would be most informative, so the model learns to ignore the feature. The paper also shows that AGN and star-forming galaxies overlap heavily in X-ray luminosity in this moderate-luminosity sample, so the physical signal is weaker than the prevailing theory assumes.","feed_headline":"Sample, not X-rays, caused the classifier's accuracy drop","feed_subtitle":"A controlled test shows the X-ray feature changes accuracy by -0.01 points; noisy labels near the BPT boundary explain why it can't help.","key_machinery":"The load-bearing analytical tool is a controlled three-way cross-validation decomposition that isolates the effect of changing the sample (full optical sample vs X-ray-detected subsample) from the effect of adding the X-ray feature, together with permutation importance to probe the feature's net contribution within a trained model. The interpretative mechanism is instance-dependent label noise: the probability of a wrong training label is highest near the BPT demarcation line, exactly where X-ray data would be most informative, so the random forest learns to avoid splitting on the X-ray feature, using it sparingly and to slight net cost.","core_discovery":"The central claim is that the previously reported accuracy decrease reflects predominantly sample selection rather than the X-ray feature itself. A controlled three-way cross-validation decomposition shows that restricting to the X-ray-detected subsample lowers accuracy by 3.52 percentage points before any X-ray feature is added, while adding the feature on the fixed subsample changes accuracy by only -0.01 points. The X-ray feature nevertheless has a small but significantly negative permutation importance: shuffling its values slightly improves held-out predictions, meaning the model's limited use of it is net-detrimental. The explanation offered is instance-dependent label noise: BPT-deriv","pith_inferences":["The negative permutation importance, though small, suggests a testable prediction: in a larger or deeper X-ray sample where the feature is less sparse, the X-ray signal might become positively important if the label-noise mechanism is the main obstacle.","The boundary-concentration result implies that retraining the same classifier with labels from an X-ray-excess selection (or another independent diagnostic) should substantially increase X-ray feature importance if the paper's explanation is correct.","The decomposition highlights a general methodological hazard: comparing classifier accuracy across different feature-availability regimes conflates sample and feature effects, so future multi-wavelength studies should always run a sample-holding-fixed feature ablation.","The luminosity overlap may be partly an artifact of the optical selection function; samples selected by X-ray brightness or deeper X-ray coverage would likely show more separation between AGN and star-forming galaxies."],"forward_implications":["Future classifiers on optically-selected samples should use independent label sources or noise-robust training methods, since BPT-derived labels are least reliable near the boundary.","The canonical 10^42 erg/s X-ray luminosity threshold is not a clean separator for moderate-luminosity samples: 41.6% of AGN fall below it and 13.5% of star-forming galaxies fall above it.","Adding a physically informative feature does not guarantee improved accuracy when training labels are noisy; the model may down-weight the feature even if it would help on true classes.","The apparent 8.25% accuracy drop in the previous paper is reinterpreted as a 3.52-point sample-selection effect and a negligible -0.01-point feature effect, so X-ray data should not be abandoned as a diagnostic on this evidence."],"fun_headline_variants":["Sample selection, not X-rays, tanked the classifier","X-ray feature exonerated: sample noise did the damage","The X-ray paradox solved: labels, not luminosity","Why adding X-rays hurt? It didn't—the sample did","Label noise, not X-rays, explains the AGN classifier dip"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The 277 objects with complete photometry are assumed to represent the full 312-object X-ray-detected sample; if the 35 excluded objects cluster near the BPT boundary, the attribution of the accuracy drop to sample selection rather than the feature could be wrong.","fun_headline_variants_meta":{"raw":{"variants":["Sample selection, not X-rays, tanked the classifier","X-ray feature exonerated: sample noise did the damage","The X-ray paradox solved: labels, not luminosity","Why adding X-rays hurt? It didn't—the sample did","Label noise, not X-rays, explains the AGN classifier dip"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000272,"raw_usage":{"total_tokens":1533,"prompt_tokens":870,"completion_tokens":663,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":614,"completion_tokens_details":{"reasoning_tokens":578}},"tokens_in":614,"tokens_out":663,"duration_ms":7521,"temperature":1.0,"reasoning_tokens":578,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T05:48:08.427826+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inspect the BPT positions of the 35 objects excluded for incomplete photometry: if their median distance to the Kewley demarcation curve is closer to 0.123 dex than to 0.743 dex, the sample-selection attribution would be weakened because the excluded objects would be disproportionately boundary-concentrated.","supporting_citations":[],"review_version":1}