{"id":"ef174eb5-d99c-4dce-b77c-d016a55da8bf","arxiv_id":"2411.19105","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A neural network trained on APOGEE data removes systematic ripples from Gaia Bp/Rp spectra, and model fitting of the corrected spectra yields a new 68-million-star catalog of temperature, gravity, and metallicity.","lead":"This paper builds a new star catalog with temperature, gravity, and metal content for 68 million Milky Way stars using Gaia's low-resolution spectra. It also releases a correction that removes systematic brightness ripples from those spectra, making the data more consistent with models and with Hubble Space Telescope standards.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline accuracy and precision figures are partly in-sample: Section 5.1 validates on IAS, which contains the TAS training set, so the -38 +/- 167 K, 0.05 +/- 0.40, -0.12 +/- 0.19 and 1.2% RMS are not established as out-of-sample for the 68M catalog.","rationale":"I read the paper as claiming two things: a flux-correction that improves XP spectrophotometry, and a 68M-star catalog with the quoted APOGEE-relative systematics. The most load-bearing condition for both is that the NN trained on TAS generalizes to all stars to which the correction is applied. The reader's verdict flags this, and the manuscript's own Section 4 caveat and Section 4.4 exceptions give the concern additional weight. Independent support exists: the CALSPEC comparison and star-cluster metallicity validation are genuinely external, and the catalog and code are public. However, the APOGEE-based numbers, the only place the headline systematics are established, are computed on a sample containing the training set, with unequal quality cuts, so they cannot by themselves support the full-catalog claim. The metal-poor subset is explicitly built from uncorrected spectra, so the EMP-search claim is not covered by the quoted validation. None of this implies the correction is wrong; it means the current paper has not cleanly separated in-sample performance from out-of-sample generalization. A holdout re-analysis is cheap and would settle it. Verdict stays CONDITIONAL: accept the catalog as a useful resource, but require the out-of-sample validation before the quoted systematics are used for high-precision science.","tokens_in":22295,"tokens_out":7463,"duration_ms":67580,"concrete_test":"Recompute the Section 5.1 and Section 5.2 validation after explicitly removing every TAS star from IAS, and report the same metrics separately for the held-out 20% TAS split and for the APOGEE stars that failed one or more of the |Delta Teff| < 200 K, |Delta log g| < 0.5, |Delta[M/H]| < 0.5 cuts, using identical quality cuts before and after correction. If the out-of-sample means/scatters of Delta Teff, Delta log g, Delta[M/H] and the RMS-vs-model peak exceed the quoted -38 +/- 167 K, 0.05 +/- 0.40, -0.12 +/- 0.19 and 1.2% by more than the stated uncertainties, the headline generalization claim is unsupported; if they match, the conditional objection is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative validation is contaminated by training data. In Section 4.1 the TAS is constructed by keeping only APOGEE stars for which a first-pass FERRE fit of the XP spectra already agrees with APOGEE (|Delta Teff| < 200 K, |Delta log g| < 0.5, |Delta[M/H]| < 0.5). The NN is then trained on the residual patterns of these 157,478 stars (Section 4.3). Section 5.1 reports before/after improvements on IAS, the parent APOGEE sample, which contains TAS; the text never states that training stars are excluded. The before/after comparison is therefore partly a regression-to-training-labels effect. The same applies to the Section 5.2 RMS-vs-model improvement (3.7% to 1.2%), which is minimized by construction. The only clean external check is CALSPEC, which shows a smaller improvement (3.2% to 2.4%, with one exception) on 109 mostly bright stars. The quality cuts are also unequal: dflux_per < 20% before correction versus < 8% after (Section 5.1), so selection effects contribute to the apparent gain. Finally, Section 4.4 explicitly does not correct stars with first-pass [M/H] <= -2.5 or missing u/v photometry, so the 124,188-star metal-poor catalog and part of the 68M catalog rest on uncorrected spectra, outside the APOGEE validation; the paper's own caveat in Section 4 states the correction is not trusted below [M/H] = -2.5, yet the EMP-search claim depends on that subset.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a neural-network-based correction for the systematic residuals (\"wiggles\") in Gaia XP spectra, using an APOGEE-selected training sample (TAS) of stars whose first-pass FERRE fits already agree with APOGEE labels. The corrected spectra are refit with FERRE against a new Synple/Kurucz model grid to derive Teff, log g, and [M/H] for 68,394,431 stars with 4000 <= Teff <= 7000 K plus a 124,188-star metal-poor subset with [M/H] <= -2.5. The paper reports improved agreement with APOGEE parameters (-38 +/- 167 K, 0.05 +/- 0.40 dex, -0.12 +/- 0.19 dex), reduced RMS between XP spectra and models (3.7% to 1.2%), and improved agreement with CALSPEC standards (3.2% to 2.4%).","tokens_in":22669,"tokens_out":4371,"duration_ms":41029,"significance":"If the claimed accuracy is established out of sample, this is a valuable contribution: it provides a publicly available, model-linked atmospheric parameter catalog for tens of millions of stars, a public flux-correction code, and a new model grid, with an independent external check against CALSPEC that is a genuine strength. The comparison with external catalogs and star clusters is also useful. However, the headline APOGEE-based accuracy figures are weakened by the overlap between the training and validation samples and by unequal quality cuts, so the current manuscript overstates the strength of the validation relative to what is demonstrated.","major_comments":[{"comment":"The TAS training set is carved from the IAS by requiring first-pass XP fits to agree with APOGEE (|Delta Teff| < 200 K, |Delta log g| < 0.5, |Delta[M/H]| < 0.5), and Section 5.1 then validates the before/after atmospheric parameters on the IAS. The text never states that TAS stars are excluded from the IAS validation. The reported improvements (e.g., Delta log g from 0.11 +/- 0.53 to 0.05 +/- 0.40) are therefore partly a regression-to-training-labels effect. Please repeat the validation on IAS \\ TAS, or on an independent sample, and explicitly state whether any TAS source remains in the comparison.","section":"Sections 4.1 and 5.1"},{"comment":"The before/after comparison uses different quality cuts: dflux_per < 20% before correction versus < 8% after correction. Because the after-correction sample is selected to contain only spectra that already fit the model well, part of the apparent improvement in the dispersions can be a selection effect. Please report the comparison using identical cuts before and after, or show that the conclusions are unchanged under a common cut.","section":"Section 5.1"},{"comment":"Section 4.4 states that spectra with first-pass [M/H] <= -2.5 or with missing u/v photometry are not corrected, yet the released metal-poor catalog of 124,188 stars with [M/H] <= -2.5 is built from these uncorrected fluxes. The APOGEE-based validation in Section 5.1 cannot validate this regime because the training sample contains very few stars below [M/H] = -2.5. The paper's own caveat in Section 4 says the correction is not trusted below -2.5, which is in tension with the abstract's EMP-search claim. Please validate the EMP subset against the Section 2.2 metal-poor sample, or explicitly state that the EMP catalog rests on uncorrected spectra.","section":"Sections 4.4 and 5.4"},{"comment":"The RMS reduction from 3.7% to 1.2% is computed on the IAS, which contains the TAS training stars, and the NN is trained to predict the residuals that are then subtracted before the same FERRE model is refit; this number is therefore minimized by construction. The clean external spectrophotometric check is the CALSPEC comparison, which shows a smaller improvement from 3.2% to 2.4% on 109 bright stars, with one exception after quality cuts. The abstract and summary should attribute the 1.2% figure to the in-sample model comparison and the 2.4% figure to the CALSPEC external validation, rather than presenting the range as if both endpoints were externally supported.","section":"Sections 5.2 and 5.3"},{"comment":"The neural network is trained on 157,478 (S_const) APOGEE stars with S/N > 70 and Teff between 3500 and 8000 K, but it is applied to roughly 200 million XP spectra, including fainter stars, hotter stars, and regions of high extinction that are underrepresented in the APOGEE footprint. No demonstration is provided that the correction interpolates or extrapolates reliably across the input space of the final catalog. Please show distributions of the NN inputs (G, colors, E(B-V), Teff, [M/H]) for TAS versus the 68M catalog, and report residual diagnostics as a function of these inputs on a truly out-of-sample subset.","section":"Sections 4.1 and 4.4"}],"minor_comments":[{"comment":"There are numerous typographical and wording errors, including \"Janurary\" in the acceptance date, \"de-reddened\" for \"dereddened\", \"di fferent\" for \"different\", \"In addtion\" for \"In addition\", and \"deviation of several parameters\" where \"derivation\" is intended.","section":"Throughout"},{"comment":"The caption does not state which dflux_per threshold is used in the \"quality cuts\" panels before and after correction, making the single exception to the improvement difficult to evaluate; please specify the thresholds.","section":"Figure 7 caption"},{"comment":"The text says the results are \"more similar to those of Zhang et al. (2023) in Teff and log g, but align more closely with Andrae et al. (2023b) for [M/H]\" based on direct catalog comparisons, but direct catalog differences do not by themselves indicate accuracy; please clarify that these statements refer to agreement, not absolute accuracy.","section":"Appendix A"},{"comment":"The paper quotes APOGEE as the reference for atmospheric parameters but does not mention that APOGEE labels are themselves on the APOGEE/ASPCAP scale; a sentence noting this would help readers interpret the quoted systematic offsets.","section":"Section 2.3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the scope of A&A and the public release of code and catalogs is a significant asset. The main risk is that the headline validation numbers are partly in-sample; the requested out-of-sample analyses are feasible within the manuscript's scope and should be obtainable from the existing data products."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: the paper delivers a genuinely useful public catalog of Teff/logg/[M/H] for 68 million stars and a released flux-correction code, and the independent CALSPEC check gives some real support to the correction. But the headline accuracy numbers against APOGEE are not as clean as the abstract makes them look, because the validation sample contains the training sample. Use the CALSPEC numbers when quoting the correction's performance, not the APOGEE numbers.\n\nWhat's actually new: the specific 68,394,431-source catalog with model-driven parameters and the correction code. The components are mostly existing tools (GaiaXPy, FERRE, Kurucz/Synple, APOGEE-trained NN, wiggle corrections), and the authors cite that prior work. The modest innovation is using metallicity-sensitive colors as NN inputs instead of metallicity itself, which makes sense because metallicity is the thing you are trying to derive. The paper documents the wiggle dependence on color, magnitude, and extinction well, and the comparison with Huang et al.'s correction is fair and useful. The cluster validation is a nice extra, though it's only six clusters.\n\nThe soft spots are real but not fatal. The APOGEE validation is partly circular: the TAS training set is carved out of the IAS sample that is later used for comparison, keeping only stars whose first-pass fits already agreed with APOGEE. Section 5.1 never states that training stars are excluded, so the -38±167 K, 0.05±0.40 dex, -0.12±0.19 dex numbers and the 3.7%-to-1.2% RMS improvement are partly in-sample. The independent CALSPEC check shows a smaller gain (3.2%-to-2.4% with quality cuts), which is the more honest estimate. The quality cuts are also unequal: dflux_per <20% before correction versus <8% after, so selection effects inflate the apparent improvement. And the 124,188-star metal-poor catalog is built from uncorrected spectra, because Section 4.4 explicitly skips stars with first-pass [M/H] <= -2.5. The paper's own caveat says the correction is not trusted below -2.5, yet the EMP-search claim rests on that subset. This should be disclosed much more prominently.\n\nWho it's for: anyone building or using large stellar-parameter catalogs, and anyone trying to hunt for extremely metal-poor candidates. It deserves a serious referee. The catalog is a resource, the code is released, the method is reproducible, and the problems are fixable with an out-of-sample APOGEE validation and clearer caveats. Send it to review, with major comments on the validation.","headline":"A useful public catalog and correction code whose claimed APOGEE validation is partly in-sample; trust the independent CALSPEC numbers more.","tokens_in":23252,"tokens_out":3355,"would_cite":true,"duration_ms":29402,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A neural-network flux correction for Gaia BP/Rp spectra makes the relative spectrophotometry about twice as precise and yields atmospheric parameters for 68,394,431 stars.","keywords":["Gaia DR3","BP/Rp spectra","spectrophotometric calibration","stellar atmospheric parameters","neural network","metallicity","extremely metal-poor stars","synthetic spectra"],"falsifier":"Take a set of stars outside the training domain, for example with $[\\mathrm{M/H}]$ below $-2.5$ or $T_{\\mathrm{eff}}$ above $7000$ K, or fainter than the training magnitudes, get their parameters from independent high-resolution spectra, and check whether applying the paper's neural-net correction to their BP/Rp spectra improves or degrades the agreement with both the model fits and the independent parameters. If the correction systematically worsens the fit or biases the parameters outside the training range, the generalization claim behind the 68-million-star catalog fails.","tokens_in":22113,"feed_emoji":"🔭","tokens_out":8216,"duration_ms":66122,"temperature":0.7,"pith_summary":"This paper tries to turn the low-resolution Gaia BP/Rp spectra of more than two hundred million stars into a reliable all-sky stellar catalog by first removing systematic flux errors and then fitting the cleaned spectra with synthetic stellar spectra. The authors train a neural network on stars with well-measured parameters from a large high-resolution infrared survey to predict the wavelength-dependent “wiggle” pattern in the BP/Rp fluxes, and show that the correction cuts the error in relative spectrophotometry from about $3.2\\%$--$3.7\\%$ to $1.2\\%$--$2.4\\%$. On the corrected spectra they measure effective temperature, surface gravity, and metallicity for $68{,}394{,}431$ stars in the range $4000$ to $7000$ K, including a subset of $124{,}188$ stars with $[\\mathrm{M/H}]$ at or below $-2.5$. If correct, this provides a physical-model-based map of stellar parameters across the whole sky that can be used to trace the Milky Way's formation and to find extremely metal-poor stars for follow-up.","feed_headline":"Flux fix yields parameters for 68 million Gaia stars","feed_subtitle":"A neural net removes systematic wiggles in BP/Rp spectra, sharpening relative fluxes from ~3.5% to ~2%.","key_machinery":"The load-bearing object is the residual pattern between each normalized BP/Rp spectrum and its best-fitting synthetic spectrum, called the wiggle. The paper models this pattern with a feed-forward neural network whose inputs are 14 photometric quantities, including metallicity-sensitive colors, Gaia magnitudes, and reddening $E(B-V)$, and whose outputs are flux corrections as a function of wavelength. The corrected spectra are then passed to FERRE, a chi-squared fitting engine, which matches them against a new grid of Kurucz model-atmosphere spectra computed at constant and variable resolution. The wiggle model is what lets the paper separate instrument-calibration systematics from the astrophysical signal.","core_discovery":"The paper's central claim is that the systematic residuals between the absolute-calibrated Gaia BP/Rp spectra and synthetic spectra are not random but depend on stellar color, brightness, extinction, and metallicity, and can therefore be predicted and removed. A neural network with seven hidden layers maps 14 photometric and reddening inputs to a 330-point flux correction; applying it brings the BP/Rp spectra into closer agreement with both model atmospheres and independent space-based flux standards. The resulting parameters agree with the high-resolution survey's values to $-38 \\pm 167$ K in $T_{\\mathrm{eff}}$, $0.05 \\pm 0.40$ dex in $\\log g$, and $-0.12 \\pm 0.19$ dex in $[\\mathrm{M/H}]$ for stars between $4000$ and $7000$ K. The authors therefore claim that the corrected spectra are accurate enough to build a catalog of $68{,}394{,}431$ stars and to support a targeted search for extremely metal-poor stars.","pith_inferences":["The paper only applies the neural-net correction to stars with initial $[\\mathrm{M/H}] > -2.5$ and with complete u/v photometry; the global catalog's reliability for the faintest, most reddened, and most metal-poor stars is therefore an extrapolation, and targeted comparisons there would be the first test.","Since the training labels come from one high-resolution survey, any systematic offset in that survey's temperatures, gravities, or metallicities would be inherited by this catalog; an independent comparison on metal-poor stars could reveal it.","The method's success suggests a natural extension: use the same residual-prediction idea to search for further wavelength-dependent systematics in other low-resolution surveys, or to extract additional labels such as alpha-enhancement if training data allow.","The catalog's $\\log g$ dispersion of about $0.40$ dex is the weakest of the three parameters; adding Gaia parallaxes or asteroseismic constraints would likely tighten it without changing the flux-correction machinery."],"forward_implications":["The corrected BP/Rp spectra can support all-sky metallicity mapping of the Milky Way with a catalog of 68 million stars.","The 124,188-star metal-poor subset provides a large candidate pool for finding extremely metal-poor stars, which can then be confirmed by high-resolution spectroscopy.","The flux corrections reduce the relative spectrophotometric error from $3.2\\%$--$3.7\\%$ to $1.2\\%$--$2.4\\%$, making the corrected spectra a more reliable reference for synthetic photometry and external calibrations.","Because parameters are obtained by fitting model atmospheres, the catalog offers an independent, physically grounded cross-check for purely data-driven parameter catalogs.","The same correction-plus-fitting pipeline can be rerun on future data releases of the same spectra with updated calibrations."],"supporting_citations":[{"why":"Defines the Gaia BP/Rp spectra calibration and documents the residual patterns that the paper's wiggle correction targets.","marker":"Montegriffo et al. (2023)"},{"why":"Supplies the BP/Rp instrument model and wavelength-dependent resolving power used to build the synthetic model grids.","marker":"Carrasco et al. (2021)"},{"why":"Provides the APOGEE DR17 catalog of stellar parameters used to build the neural-network training sample.","marker":"Abdurro’uf et al. (2022)"},{"why":"Provides the FERRE fitting engine used to match corrected BP/Rp spectra to model atmospheres.","marker":"Allende Prieto et al. (2006)"},{"why":"Defines the nsc library of Kurucz model-atmosphere spectra that the synthetic grid is based on.","marker":"Allende Prieto et al. (2018)"},{"why":"Supplies the CALSPEC flux standards used as the independent check on the corrected spectrophotometry.","marker":"Bohlin et al. (2014)"},{"why":"Provides an earlier BP/Rp correction package that the paper compares against when validating its own corrections.","marker":"Huang et al. (2024a)"},{"why":"Demonstrates with mock spectra that BP/Rp data can constrain metallicity down to the extremely metal-poor regime, motivating the metal-poor search.","marker":"Witten et al. (2022)"}],"fun_headline_variants":["Neural net fixes Gaia spectra for 68 million stars","AI-corrected Gaia spectra yield 68M star parameters","Systematic flux correction sharpens Gaia BP/RP data","68M stars get better parameters from corrected Gaia spectra","Flux fix for Gaia BP/RP: 68M stellar parameters"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The neural network is trained on about 157,000 bright stars whose model fits already agree with a high-resolution survey, and the whole catalog depends on assuming that the wiggle pattern it learned applies to all 200-plus million BP/Rp stars, including fainter, hotter, and more metal-poor objects not represented in training.","fun_headline_variants_meta":{"raw":{"variants":["Neural net fixes Gaia spectra for 68 million stars","AI-corrected Gaia spectra yield 68M star parameters","Systematic flux correction sharpens Gaia BP/RP data","68M stars get better parameters from corrected Gaia spectra","Flux fix for Gaia BP/RP: 68M stellar parameters"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000817,"raw_usage":{"total_tokens":3693,"prompt_tokens":1176,"completion_tokens":2517,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":792,"completion_tokens_details":{"reasoning_tokens":2433}},"tokens_in":792,"tokens_out":2517,"duration_ms":17877,"temperature":1.0,"reasoning_tokens":2433,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T10:31:52.796646+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of stars outside the training domain, for example with $[\\mathrm{M/H}]$ below $-2.5$ or $T_{\\mathrm{eff}}$ above $7000$ K, or fainter than the training magnitudes, get their parameters from independent high-resolution spectra, and check whether applying the paper's neural-net correction to their BP/Rp spectra improves or degrades the agreement with both the model fits and the independent parameters. If the correction systematically worsens the fit or biases the parameters outside the training range, the generalization claim behind the 68-million-star catalog fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Demonstrates with mock spectra that BP/Rp data can constrain metallicity down to the extremely metal-poor regime, motivating the metal-poor search."}],"review_version":1}