{"id":"103d5160-270e-4f04-b6c7-1143536dc9c6","arxiv_id":"2507.05901","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"The paper releases SoS DR2: a PASTEL-recalibrated spectroscopic reference and a neural-network catalog giving stellar parameters for about 23 million stars, with improved low-metallicity [Fe/H] accuracy validated on globular clusters.","lead":"Astronomers combined five spectroscopic surveys and a neural network to produce a new catalog of temperature, gravity, and metal content for about 23 million stars. The catalog is especially accurate for old, metal-poor stars, a regime where earlier machine-learning catalogs perform poorly, which matters for studying the Milky Way's early history.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The low-metallicity accuracy claim is anchored to PASTEL as ground truth, but PASTEL also calibrates the training labels and likely shares sources with the Harris cluster metallicities used for validation; without a demonstrated independent benchmark, the improvement may be partly inherited.","rationale":"The reader's weakest assumption correctly identifies PASTEL label reliability as the soft spot. I agree, but the more load-bearing formulation is the lack of demonstrable independence between the PASTEL-calibrated training set and the globular-cluster validation benchmark. If Harris (1996) metallicities are built from the same high-resolution literature as PASTEL, then the cluster validation is not an external check of 'accuracy' but a check of consistency with a shared literature scale. The paper's own admission that PASTEL and the surveys disagree for metal-poor giants, without resolving which side is wrong, means the decision to treat PASTEL as truth is exactly the assumption the headline claim requires. The concrete retraining and cross-match test would settle this. The paper has real strengths: it is transparent, releases the catalog, provides multiple validation routes, and does not overstate the very metal-poor improvement. These support a CONDITIONAL verdict, which the reader already reached; my concern reinforces that verdict rather than changing it, so I recommend UNCHANGED.","tokens_in":28987,"tokens_out":7167,"duration_ms":85549,"concrete_test":"Retrain the SoS-ML pipeline twice: once with the PASTEL recalibration (Eq. 2) and the 406 PASTEL ultra-metal-poor training stars removed, and once as published; then recompute Fig. 8 for the same 20 globular clusters. In parallel, cross-match the Vasiliev & Baumgardt cluster members used for validation to PASTEL and to the SoS-Spectro training set, and compute the Harris offsets restricted to stars with no PASTEL or training overlap. If the +0.08 offset and 0.10 scatter survive the leave-PASTEL-out retraining and the non-overlap subset, the claim is independent; if they degrade or vanish, the improvement is largely inherited from PASTEL.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline improvement in metal-poor [Fe/H] accuracy is validated almost entirely on globular clusters using Harris (1996) metallicities (Sect. 6.4, Fig. 8). The training labels, however, were recalibrated to PASTEL via Eq. 2 (Sect. 3.1), and the very metal-poor training set was augmented with 406 PASTEL stars (Sect. 4.3, App. C.1). PASTEL and Harris are both literature compilations of high-resolution spectroscopy; cluster stars and the papers underlying Harris cluster metallicities are plausibly represented in PASTEL, and the paper does not quantify this overlap. The authors themselves state that for metal-poor giants, PASTEL and every spectroscopic survey disagree by up to 2 dex (in logg) and that 'it is difficult to decide whether the problem lies in PASTEL or in the surveys' (Sect. 3.1), yet PASTEL is adopted as ground truth for calibration. If PASTEL carries a systematic offset in the metal-poor regime, SoS-ML will inherit it, and part of the '+0.08 vs +0.14/+0.43/+0.78' improvement is a realignment to the PASTEL scale rather than an independent accuracy gain. Additionally, the Gu et al. globular-cluster offset is based on only two clusters (App. D.1), weakening the cross-method comparison. The claim may still be true, but the decisive evidence is not yet independent.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents the second data release of the Survey of Surveys project: a recalibrated spectroscopic catalog (SoS-Spectro) and a new photometric machine-learning catalog (SoS-ML) containing Teff, log g, and [Fe/H] for about 23 million stars. The ML method is a multilayer perceptron trained on SoS-Spectro labels, using Gaia astrometry and photometry, SDSS or SkyMapper photometry, and Gaia/Andrae parameter estimates as input features. The paper reports test-set errors of about 50 K, 0.07 dex, and 0.08 dex for Teff, log g, and [Fe/H], and validates the results against other ML catalogs, globular and open clusters, APOKASC asteroseismic gravities, and spectral types. The central claim is that SoS-ML improves precision and accuracy relative to other ML catalogs, especially in the metal-poor range, based on the globular cluster comparison.","tokens_in":29212,"tokens_out":4543,"duration_ms":49431,"significance":"If the low-metallicity accuracy claim holds, the catalog is a valuable community resource: it is public, covers about 19 million stars with ML parameters and an additional 3.75 million with spectroscopic parameters, and it includes a careful error budget that combines a training error and a repeatability error derived from ten independent trainings. The authors are also appropriately cautious in several places, explicitly acknowledging that the reference catalog has its own biases and that the APOKASC log g comparison shows a residual 0.14 dex offset. However, the decisive evidence for the headline metal-poor improvement is not yet independent: PASTEL is used both to calibrate the training labels and to augment the metal-poor training sample, and the overlap between PASTEL and the Harris (1996) globular cluster metallicities used for validation is not quantified. The paper would be materially strengthened by an external validation sample that was not used in any calibration or training step.","major_comments":[{"comment":"The low-metallicity accuracy claim rests on a benchmark that is not independent of the training labels. SoS-Spectro is recalibrated against PASTEL via Eq. (2) in Sect. 3.1, and the metal-poor training sample is augmented with 406 PASTEL stars in Sect. 4.3; the headline validation in Sect. 6.4 then compares ML [Fe/H] with Harris (1996) globular cluster metallicities. The paper does not quantify how many cluster stars or underlying literature measurements are shared between PASTEL and Harris, and the authors themselves note up to 2 dex disagreement between PASTEL and all spectroscopic surveys for metal-poor giants (Sect. 3.1). As a result, the +0.08 dex offset and 0.10 dex scatter reported in Fig. 8 may represent a realignment to the PASTEL scale rather than an independent accuracy gain. Please quantify the PASTEL-Harris overlap and re-validate on an external sample not used in any calibration step, or explicitly downgrade the claim to consistency with PASTEL.","section":"Sect. 3.1, 4.3, 6.4"},{"comment":"Andrae et al. (2023) [Fe/H] is used as an input feature (Tables 2 and 3), and the same catalog is then used as a comparison benchmark in Fig. 7 (bottom-right panel). The excellent median agreement (-0.02 dex, MAD 0.08) is therefore partly by construction and cannot serve as independent evidence for accuracy. The text in Sect. 6.3 even uses this agreement, combined with the cluster comparison, to argue that the method improves on Andrae et al.; please separate the input-dependence discussion from the validation and rely on the cluster, asteroseismic, and spectral-type comparisons for claims of improvement over Andrae et al.","section":"Sect. 6.3 and Tables 2-3"},{"comment":"The log g validation against APOKASC-3 shows a median offset of 0.14 dex for giants, with error bars barely touching the 1:1 line, and the authors attribute this to the PASTEL-based calibration. Since the catalog reports formal log g errors of about 0.07-0.08 dex (Table 5), the systematic offset is larger than the quoted precision. The broad abstract statement of 'substantial improvements ... in terms of precision and accuracy' should be restricted to [Fe/H] or supported by a log g validation that reaches the claimed accuracy; otherwise the error bars in Table 5 should be recalibrated to include this systematic component.","section":"Appendix C.3"},{"comment":"The comparison with Gu et al. (2025) for globular clusters is based on only two clusters (NGC 7078 and NGC 7089), as stated in Appendix D.1; the mean offset +0.78 and scatter 0.25 shown in Fig. 8 therefore rest on very limited data and should not be given equal weight in the cross-catalog comparison. Either compute the Gu et al. statistics on a larger cluster sample or explicitly mark this panel as preliminary.","section":"App. D.1 and Fig. 8"}],"minor_comments":[{"comment":"There are several typographical errors, including 'di fferent' in the Abstract and Introduction, 'Suveys' in the Introduction, and 'Trainig' in the caption of Fig. 4; these should be corrected.","section":"General"},{"comment":"The text describes Eq. (2) as a 'three-parameter linear fit', but the equation has four fitted coefficients (a, b, c, d) corresponding to three predictors plus an intercept; please rephrase as a linear fit with three predictors plus an intercept.","section":"Sect. 3.1"},{"comment":"The acronym for the SAGES catalog appears as 'SAGE' in Sect. 6.2 and as 'SAGA' in Appendix C.1; please use a single consistent name.","section":"Sect. 6.2 and App. C.1"},{"comment":"The quality cuts on the ML errors (eML on log g > 1.0 dex, eML on [Fe/H] > 1.0 dex) are very loose, and only the total number of removed stars (1557) is reported; please also report how many stars are removed by each individual cut so readers can judge the impact of these thresholds.","section":"Sect. 5.2"}],"recommendation":"major_revision","confidential_remarks":"The paper describes a useful and well-executed catalog effort, and the public release is a clear asset. My main concern is that the headline metal-poor accuracy claim is not yet supported by an independent benchmark: PASTEL is used both to calibrate the training labels and as part of the validation chain, and the overlap with the Harris cluster metallicities is unquantified. This is fixable with additional validation or with a more measured claim, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a solid, useful catalog paper with a real twist—they train a small MLP to refine existing Gaia DR3 and Andrae et al. parameter estimates rather than predict from photometry from scratch—and they release both the recalibrated SoS-Spectro and the ~19-million-star SoS-ML catalog. The validation is unusually thorough: internal test errors (median ~50 K, 0.08 dex), comparisons with Zhang, Gu, and Andrae, globular and open clusters, APOKASC-3, and spectral types for cool stars. Credit is due: the paper is transparent about limitations, including the very metal-poor failure mode, and ships the catalog publicly.\n\nThe soft spot is the low-metallicity headline. The training labels are recalibrated to PASTEL (Eq. 2, Sect. 3.1) and the very metal-poor training set is augmented with PASTEL stars (Sect. 4.3). PASTEL is then used as a benchmark in Appendix C.1, and the globular cluster metallicities from Harris (1996) are a literature compilation that likely shares sources with PASTEL. So the +0.08 dex offset versus Harris, compared to +0.43 for Andrae and +0.78 for Gu, is partly a realignment to the PASTEL scale, not a fully independent accuracy gain. The authors acknowledge the ambiguity—they say it is \"difficult to decide whether the problem lies in PASTEL or in the surveys\"—but they proceed with PASTEL as ground truth anyway. That does not invalidate the catalog, but the abstract's \"substantial improvements\" should be read as \"improvements relative to the PASTEL scale\" until an independent benchmark for metal-poor giants is shown.\n\nOther issues are minor. The APOKASC-3 comparison shows a 0.14 dex logg offset for giants, discussed in App. C.3. The Andrae comparison is partially circular (Andrae [Fe/H] is an input feature), acknowledged in Sect. 6.3. Very metal-poor stars still misbehave despite the PASTEL enrichment, as they candidly report. The per-star errors are ML test plus repeatability, excluding reference systematics, exactly as stated.\n\nWho gets value: anyone needing large photometric stellar parameter catalogs with published error estimates, especially for metal-poor work. The strict quality cuts limit it to relatively clean stars, but that is stated. I would send this to a serious referee; the referee should ask for a quantitative assessment of the PASTEL/Harris source overlap and a clear statement of which validation results are truly independent.","headline":"A genuinely useful catalog with a novel 'refine, don't predict from scratch' approach, but the headline metal-poor accuracy gain is partly inherited from PASTEL, and the validation independence is weaker than the abstract suggests.","tokens_in":30034,"tokens_out":4769,"would_cite":true,"duration_ms":46274,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A photometric neural network, seeded with previous estimates and trained on a PASTEL-recalibrated spectroscopic reference, predicts stellar parameters for tens of millions of stars and claims its largest accuracy gain exactly where…","keywords":["Survey of Surveys","stellar parameters","machine learning","multi-layer perceptron","metal-poor stars","globular clusters","photometric surveys","Gaia DR3"],"falsifier":"Measure NLTE-corrected high-resolution metallicities for stars in the 20 globular clusters used here (e.g., NGC 7078/M15, NGC 6397) and compare them with both PASTEL and the SoS-ML predictions; if the NLTE values fall systematically below PASTEL by more than about 0.1 dex, the claimed metal-poor validation is an artifact of the reference. A second, cheaper check: the paper already shows median [Fe/H] residuals of roughly 0.6 dex versus PASTEL and SAGA for [Fe/H] below -2, so verifying whether those extreme stars' true metallicities match the catalog would delimit the claim's validity range.","tokens_in":28628,"feed_emoji":"⭐","tokens_out":9768,"duration_ms":90757,"temperature":0.7,"pith_summary":"The paper sets out to show that photometric surveys can yield stellar parameters accurate enough for large-scale Galactic archaeology, provided the machine-learning model does not start from scratch. Its recipe is to take existing estimates (Gaia DR3 Teff and logg, Andrae et al. (2023) [Fe/H]) and refine them with a multi-layer perceptron trained on SoS-Spectro, a homogenized spectroscopic reference that was itself recalibrated on the high-resolution PASTEL database. The result is a catalog of around 19 million ML-derived parameters, validated on globular clusters where the mean [Fe/H] offset drops to +0.08 dex with 0.10 scatter, compared with +0.43 (Andrae et al.), +0.78 (Gu et al.), and +0.14 (Zhang et al.). A reader should care because metal-poor stars are precisely where photometric ML methods have historically been least trustworthy, and the claimed improvement is concentrated there.","feed_headline":"Machine learning sharpens stellar parameters for 23 million stars","feed_subtitle":"Seeding with Gaia and Andrae et al. estimates cuts low-metallicity bias against globular clusters to +0.08 dex.","key_machinery":"The load-bearing mechanism is the two-stage calibration chain. First, a linear correction (Eq. 2) rescales SoS-Spectro's homogenized survey parameters to the PASTEL high-resolution system, reducing the low-gravity, low-metallicity deviations by more than a factor of two. Second, a multi-layer perceptron (a stacked neural network with 18 hidden layers, batch normalization, dropout, and Leaky ReLU activations) learns to map photometric and astrometric features—absolute magnitudes in both reddened and dereddened forms, distance, Gaia DR3 gspphot values, Andrae et al. (2023) [Fe/H], and error estimates—onto those recalibrated labels. The seeding with previous parameter estimates and the missing-value flags are what allow the network to act as a refiner rather than a from-scratch predictor. Errors are computed per star by summing in quadrature the training error (MLP versus SoS-Spectro on the test set) and a repeatability error from ten differently initialized trainings.","core_discovery":"The central claim is that two ingredients, a high-resolution-calibrated reference catalog and the use of pre-existing parameter estimates as input features, are what make photometric stellar parameters accurate. SoS-Spectro is built from five spectroscopic surveys (APOGEE, GALAH, Gaia-ESO, RAVE, LAMOST) and recalibrated to PASTEL through the linear correction $\\Delta f = a + b\\,T_{\\rm eff}^{\\rm SoS} + c\\,\\log g^{\\rm SoS} + d\\,[{\\rm Fe/H}]^{\\rm SoS}$ (Eq. 2), which more than halves the deviations at low log g and low [Fe/H]. A multi-layer perceptron then predicts each parameter separately from features that include distance, reddened and dereddened absolute magnitudes, Gaia gspphot temperatures and gravities, Andrae et al. (2023) metallicities, and their errors, with flags to handle missing values. On the test sample the predictions reproduce SoS-Spectro to medians of about 50 K in Teff, 0.08 dex in logg, and 0.07 dex in [Fe/H]; merged with SoS-Spectro, the released catalog covers about 23 million stars. The paper argues that the improved low-metallicity behavior is visible in the globular cluster comparison, where its [Fe/H] residuals are much smaller and better centered than those of the other ML catalogs, and that this stems from the PASTEL recalibration and from refining rather than replacing previous estimates.","pith_inferences":["If PASTEL's metal-poor giant metallicities carry systematics (the paper explicitly cannot decide whether PASTEL or the surveys are at fault), part of the globular-cluster agreement is inherited from the reference rather than produced by the ML method; a comparison against independent NLTE abundances of the same clusters would separate the two.","The same seeding recipe with asteroseismic logg as an additional input feature could plausibly fix the residual 0.14 dex gravity offset seen against APOKASC-3, since the paper's own validation shows the offset mirrors the SoS-Spectro versus PASTEL logg disagreement for giants.","The persistent ~0.6 dex median residual for [Fe/H] below -2 versus PASTEL and SAGA, even after training enrichment, suggests the low-metallicity improvement is real in the globular-cluster range (-2.3 to -0.7) but should not be extrapolated to the extremely metal-poor regime.","The method's dependence on distance and reddening input quality implies the accuracy claims hold where those are reliable; applying the catalog to heavily reddened or crowded fields should be done with the provided train-area flags."],"forward_implications":["If the claim holds, photometric surveys alone can supply usable stellar parameters for tens of millions of stars, extending spectroscopic-quality metallicity measurements to full-sky Gaia and multi-band samples.","The metal-poor regime, where previous ML catalogs show offsets of +0.14 to +0.78 dex against globular clusters, becomes accessible for studies of the halo and accreted populations.","The refine-don't-reinvent strategy, using existing parameter estimates as inputs, can be applied to any future parameter set as survey pipelines improve.","The catalog provides a homogeneous 23-million-star reference spanning the SDSS and SkyMapper footprints, useful for cluster studies and asteroseismic comparisons.","Future releases can absorb newer survey reductions (APOGEE DR17, GALAH DR4, LAMOST DR10), since the current SoS-Spectro exhibits wavy biases against them at the 10-150 K, 0.1-0.3 dex level."],"supporting_citations":[{"why":"Provides the first SoS release and the preliminary SoS-Spectro parameters that are recalibrated and used as training labels.","marker":"Tsantaki et al. (2022)"},{"why":"The PASTEL high-resolution compilation used for the external recalibration in Eq. 2 and the reference defining metal-poor behavior.","marker":"Soubiran et al. (2016)"},{"why":"Supplies the [Fe/H] input feature and is the main baseline catalog the method claims to improve upon.","marker":"Andrae et al. (2023)"},{"why":"Gaia DR3 provides astrometry, photometry, gspphot parameters, and the coordinate system for cross-matching.","marker":"Gaia Collaboration et al. (2023)"},{"why":"Geometric distances used to compute absolute magnitudes and as a network feature.","marker":"Bailer-Jones et al. (2021)"},{"why":"The globular cluster metallicity compilation used as ground truth in the validation supporting the metal-poor claim.","marker":"Harris (1996)"},{"why":"The XP-spectra ML catalog compared throughout; its +0.14/0.28 globular-cluster residual is the baseline for the paper's +0.08/0.10.","marker":"Zhang et al. (2023)"},{"why":"The SAGES random-forest catalog, the closest method-and-scale comparison, showing +0.78/0.25 cluster residuals.","marker":"Gu et al. (2025)"},{"why":"APOKASC-3 asteroseismic surface gravities used for the independent logg validation (0.14 dex offset).","marker":"Pinsonneault et al. (2025)"},{"why":"The cross-match software used to combine Gaia with SDSS and SkyMapper catalogs.","marker":"Marrese et al. (2017, 2019)"}],"fun_headline_variants":["AI refines stellar parameters for 23 million stars, boosting low-metallicity accuracy","New catalog improves stellar parameters for 23 million stars using machine learning","Machine learning plus prior estimates yields sharper stellar parameters for 23M stars","Survey of Surveys II: ML-trained catalog for 23 million stars","Improved stellar parameters for 23M stars via ML and recalibrated references"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The low-metallicity training labels are assumed trustworthy because SoS-Spectro was recalibrated against the PASTEL database; if PASTEL's metallicities for metal-poor giants are systematically biased, the improved globular-cluster agreement is partly inherited rather than newly established.","fun_headline_variants_meta":{"raw":{"variants":["AI refines stellar parameters for 23 million stars, boosting low-metallicity accuracy","New catalog improves stellar parameters for 23 million stars using machine learning","Machine learning plus prior estimates yields sharper stellar parameters for 23M stars","Survey of Surveys II: ML-trained catalog for 23 million stars","Improved stellar parameters for 23M stars via ML and recalibrated references"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000529,"raw_usage":{"total_tokens":2693,"prompt_tokens":1230,"completion_tokens":1463,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":846,"completion_tokens_details":{"reasoning_tokens":1365}},"tokens_in":846,"tokens_out":1463,"duration_ms":11228,"temperature":1.0,"reasoning_tokens":1365,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:16:57.742263+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure NLTE-corrected high-resolution metallicities for stars in the 20 globular clusters used here (e.g., NGC 7078/M15, NGC 6397) and compare them with both PASTEL and the SoS-ML predictions; if the NLTE values fall systematically below PASTEL by more than about 0.1 dex, the claimed metal-poor validation is an artifact of the reference. A second, cheaper check: the paper already shows median [Fe/H] residuals of roughly 0.6 dex versus PASTEL and SAGA for [Fe/H] below -2, so verifying whether those extreme stars' true metallicities match the catalog would delimit the claim's validity range.","supporting_citations":[{"cited_title":"2016, , 591, A118","cited_arxiv_id":null,"evidence_quote":"The PASTEL high-resolution compilation used for the external recalibration in Eq. 2 and the reference defining metal-poor behavior."},{"cited_title":"2023, , 267, 8","cited_arxiv_id":null,"evidence_quote":"Supplies the [Fe/H] input feature and is the main baseline catalog the method claims to improve upon."},{"cited_title":"M., & Rix , H.-W","cited_arxiv_id":null,"evidence_quote":"The XP-spectra ML catalog compared throughout; its +0.14/0.28 globular-cluster residual is the baseline for the paper's +0.08/0.10."},{"cited_title":"M., Marinoni , S., Fabrizio , M., & Giuffrida , G","cited_arxiv_id":null,"evidence_quote":"The cross-match software used to combine Gaia with SDSS and SkyMapper catalogs."}],"review_version":1}