REVIEW 4 major objections 5 minor 16 references
Metallicities of 20 Million Giant Stars Based on Gaia XP spectra
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read An uncertainty-aware neural network trained on just 1,012 bright giants derives iron abundances for roughly 20 million Gaia giant stars, reaching $\mathrm{[Fe/H]} \approx -4$ without the carbon-enhancement bias of earlier XP-based catalogs.
desk verdict A genuinely useful Gaia XP metallicity catalog with solid validation, but the headline counts don't add up and the reliability claim runs ahead of the validated parameter space. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the UA-CSNet, a two-branch multilayer perceptron. One branch maps the corrected, dereddened, re-sampled 60-pixel XP spectrum into a latent feature vector $Z$ and a predicted mean $\mathrm{[Fe/H]}$; the second branch concatenates $Z$ with the spectral error vector and predicts the heteroscedastic variance of the estimate. Training minimizes a cost-weighted Gaussian negative log-likelihood, $$\mathcal{L} = c(x_i)\sum_i\left(\frac{(Y_i-\Phi_{FC}(Z))^2}{2\$sigma_i^{2}$}+\log\sigma_i\right),$$ with bin-based costs $c(x_i)=\left(f_n(x_i)/\max f_n(X)\right)^{-\gamma}$ for $\gamma=0.5$ that inflate the penalty for rare metal-poor stars, an error floor of $\epsilon=0.02$ dex, and $L_2$ regularization. This combination is what keeps the sparse, noisy metal-poor tail of the training distribution from being dominated by the populous solar-metallicity stars.
What would settle it
Cross-match the released catalog with high-resolution spectroscopic metallicities (for example from APOGEE or SAGA) for a few hundred giants inside the published sample but outside the training domain — say $14 < G < 17$ or $0.2 < E(B-V) < 0.5$ — and measure the residual scatter and bias. If the scatter exceeds roughly 0.4 dex or a systematic offset above 0.3 dex appears in that regime, the claim that all 20 million stars carry reliable metallicities fails for the faint, extincted majority; if the residuals stay near the 0.22 dex scatter measured on the bright SAGA test set, the generalization holds.
Extended reading notes
Core claim
The central claim is that the UA-CSNet — an uncertainty-aware, cost-sensitive neural network — produces reliable and precise $\mathrm{[Fe/H]}$ estimates for approximately 20 million giant stars selected from Gaia DR3, including (by the abstract's count) 360,000 very metal-poor and 50,000 extremely metal-poor stars, and reaches the metal-poor tail down to $\mathrm{[Fe/H]} \sim -4$. The network outputs both a mean $\mathrm{[Fe/H]}$ and a per-star Gaussian variance, partitioned into aleatory uncertainty from spectral noise and epistemic uncertainty from sparse training coverage, and it is trained with a loss that up-weights rare metal-poor stars so that the imbalanced $\mathrm{[Fe/H]}$ distribution does not wash out the extremes. The paper validates the estimates against the independent SAGA catalog (0.22 dex scatter), against SDSS/SEGUE and LAMOST spectroscopy down to $\mathrm{[Fe/H]} \approx -3.5$, against the bright very-metal-poor sample of Viswanathan et al. (2024), and on five star clusters, and it shows that carbon-enhanced very metal-poor and extremely metal-poor stars are not overestimated, in contrast to earlier XP-based catalogs. On that basis, the paper presents the released catalog as reliable for studying the formation and chemo-dynamical evolution of the Milky Way.
Load-bearing premise
The network is trained on 1,012 bright, high-latitude, lightly extincted giants from PASTEL ($G \approx 6$–14, $|b|>20^\circ$, $E(B-V)<0.2$) and then applied to roughly 31 million stars selected with only the first three training cuts, with the 20-million-star 'reliable' sample allowing $E(B-V)$ up to 0.5, so the central claim assumes the learned spectrum-to-metallicity mapping still holds for fainter, more heavily extincted stars the training set never sampled.
Editorial extensions
If this is right
- A public catalog of about 20 million giants with $\mathrm{[Fe/H]}$ and per-star uncertainties (validated for $G<15$) becomes available for Galactic structure studies; the paper itself displays vertical metallicity gradients across the disk in the $R$–$Z$ plane.
- The catalog's very metal-poor and extremely metal-poor component gives follow-up spectroscopic surveys a concrete target list for the oldest Milky Way populations.
- Carbon-enhanced metal-poor stars are not systematically overestimated, so the metal-poor tail of this catalog is cleaner than earlier XP-based photometric metallicities.
- The uncertainty-aware, cost-weighted training scheme shows a route to abundance estimation from low-resolution spectra when labels are sparse and the label distribution is strongly imbalanced, without requiring new observations.
Reading between the lines
- If the metal-poor tail is as clean as claimed, the very metal-poor subset is large enough to map the spatial structure of the early-accretion halo — for example, searching for dwarf-galaxy debris by clustering these giants in kinematics and chemistry, a use the paper does not develop.
- The two-branch architecture should transfer to other low-resolution surveys and to other labels such as $\mathrm{[\alpha/Fe]}$ or carbon abundance; the paper makes no such claim.
- The abstract and Section 4.3 give different counts for the same final sample (360,000 VMP and 50,000 EMP stars versus 1,089,712 VMP and 369,269 EMP candidates), so a reader should check the released catalog, not the text, for the definitive census.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents UA-CSNet, an uncertainty-aware cost-sensitive neural network that estimates metallicity and an associated uncertainty from dereddened, corrected Gaia BP/RP (XP) spectra of giant stars. The model is trained on 1,012 PASTEL giants selected with six quality and giant-star cuts, then applied to 31,360,788 giants selected with the first three cuts; a subset with E(B−V)<0.5 is designated as the final reliable sample of about 20 million stars. The authors validate against SAGA, SDSS/SEGUE, LAMOST DD-Payne, and two Gaia-based catalogs, test on five star clusters, and report that carbon enhancement does not bias their very metal-poor estimates. They quote approximately 360,000 VMP and 50,000 EMP stars in the Abstract, while Section 4.3 gives different candidate counts, and the paper releases the catalog through a public DOI.
Significance. If the 20-million-star reliability claim is established, this would be a substantial community resource for Galactic archaeology, particularly because the method reaches [Fe/H] ~ −4 and is claimed to be insensitive to carbon enhancement. The paper's strengths include independent validation on SAGA, SDSS/SEGUE, and LAMOST, cluster tests, a specific carbon-enhancement check, and public release of the catalog at https://doi.org/10.12149/101604. However, the current analysis does not fully demonstrate reliability over the full magnitude/extinction range of the released sample, and the internal inconsistency in the VMP/EMP counts makes the exact definition of the final sample ambiguous. These issues are fixable and should be addressed before publication.
major comments (4)
- [Abstract; §4.3] The reported numbers of very metal-poor and extremely metal-poor stars are mutually inconsistent: the Abstract and Summary state approximately 360,000 VMP and 50,000 EMP stars, whereas §4.3 reports 1,089,712 VMP and 369,269 EMP candidate stars. Because these counts are headline results and the “final reliable sample” is defined only by E(B−V)<0.5 after the first three cuts, the reader cannot determine which selection produces the abstract numbers. Please specify the exact cuts (including any G-magnitude or error cut) used for the 20 million-star sample and make all quoted counts consistent.
- [§4.2–§4.3] The reliability claim for the 20 million-star sample rests on extrapolation beyond the training domain. The PASTEL training set is limited to roughly G=6–14, |b|>20°, and E(B−V)<0.2 (§2.3), and the final sample is selected with only the first three cuts, while the “reliable” subset relaxes the extinction cut to E(B−V)<0.5. Section 4.2 restricts application to the (BP−RP)_0 training range but not to G, and §4.3 recommends predicted errors only for G<15 while releasing [Fe/H] values for the full sample. The external comparisons in Figs. 7–8 reach G≈17 and are not stratified by E(B−V), and the adopted XP correction (§2.1) is calibrated only for G<17.5 and E(B−V)<0.8. The authors should either provide validation stratified by G and E(B−V) or explicitly flag restricted reliability in the released catalog and revise the “20 million reliable” claim accordingly.
- [§3.2, Eq. (12)] The cost-sensitive loss in Eq. (12) is written as L = c(x_i) Σ_i (...), which multiplies the entire sum by a single sample's cost and is not a valid training objective as written. It should presumably read L = Σ_i c(x_i) [(Y_i−Φ(Z))^2/(2σ_i^2)+log σ_i]. Please correct the equation and state explicitly how the histogram weights, M and N, and the exponent γ enter the per-sample cost.
- [§4.1] The hyperparameters defining the architecture and the loss (s, epoch, γ, ε, M, N) are described as chosen by experimentation, and no cross-validation or learning-curve analysis is reported for the 1,012-star training set. Because the quoted precision and uncertainty calibration depend on these settings, please add a cross-validation or a repeated hold-out analysis on PASTEL, or otherwise demonstrate that the reported 0.09 dex training-set scatter and 0.22 dex SAGA scatter are not specific to the single manual configuration.
minor comments (5)
- [§4.2.2, Fig. 6] Figure 6's caption labels the comparison catalog as “Huang et al. (2024)”, while the text and reference list discuss the same catalog as Huang et al. (2025). Please make the year consistent.
- [§3.1, Eqs. (6)–(8)] The uncertainty notation is inconsistent: Eq. (6) treats δY_i as the standard deviation of the Gaussian, Eq. (7) sets δY_i = max(ε, δY_{i,model}^2), and Eq. (8) introduces a new symbol σ_i without relating it to δY_i. Please define these symbols consistently and check the exponent in Eq. (7).
- [Throughout] The model name is written both as “CSnet” and “CSNet”; please use a single spelling.
- [§4.3] The paper states that the catalog is publicly available and describes the table columns, but it does not provide a formal data access statement in the main text beyond the DOI in the Abstract. A short sentence on the catalog access and format would improve usability.
- [References] “Zhang M. et al. (submitted)” is cited for LAMOST DD-Payne parameters but no preprint or DOI is given; please provide a complete reference or an arXiv identifier.
Circularity Check
No significant circularity: the [Fe/H] target is external PASTEL spectroscopy, and the model is tested on disjoint SAGA, SDSS, LAMOST, and cluster benchmarks; self-citations are pipeline tools, not target-defining.
full rationale
The derivation chain is supervised regression from Gaia XP spectra to PASTEL [Fe/H] labels. The target labels are external high-resolution spectroscopic measurements, not outputs of the model or of the same-group correction catalogs. SAGA is explicitly excluded from training and used as a disjoint test set; SDSS/SEGUE, LAMOST DD-Payne, and APOGEE cluster members provide independent external comparisons. The self-citations (CSNet architecture from Yang et al. 2022; XP systematic correction from Huang et al. 2024; extinction curve from Zhang et al. 2024; comparison sample from Huang et al. 2025) are pipeline components or validation benchmarks, and none of them defines the PASTEL metallicity labels or the VMP/EMP counts. No equation in Section 3 reduces the predicted [Fe/H] to an input statistic by construction, and no fitted parameter is renamed as a prediction. Two non-circular caveats are worth stating: (1) the abstract/summary VMP/EMP numbers (360,000/50,000) differ from Section 4.3's 1,089,712/369,269, which is an internal consistency problem rather than circularity; (2) Section 4.3's limitation that predicted errors should only be used for G<15 while releasing [Fe/H] for the full sample is a generalization caveat, not a circular step.
Assumptions & free parameters
free parameters (6)
- Network weights and biases =
Not released; optimized on 1012 PASTEL training stars
- Cost exponent gamma =
0.5
- Error floor epsilon =
0.02 dex
- Histogram bins (M, N) =
(10, 15)
- Optimization hyperparameters =
learning rate 0.001, batch size 512, 30000 epochs
- Final extinction threshold =
E(B-V) < 0.5 mag
assumptions (5)
- domain assumption Gaussian likelihood for predicted [Fe/H] (Eqs. 8-10)
- domain assumption Corrected Gaia XP spectra and the Zhang et al. (2024) extinction curve accurately remove systematics
- domain assumption PASTEL catalog metallicities are on a consistent scale down to [Fe/H] ~ -4
- domain assumption Giant selection cuts define a population shared by the training and target samples
- standard math Standard MLP capacity and backpropagation assumptions
Cite this review
Pith. "Pith review of Metallicities of 20 Million Giant Stars Based on Gaia XP spectra." pith.science (2026). https://pith.science/paper/3XEQQJRD
@misc{pith2026250505281,
author = {Pith},
title = {Pith review of: Metallicities of 20 Million Giant Stars Based on Gaia XP spectra},
year = {2026},
howpublished = {\url{https://pith.science/paper/3XEQQJRD}},
note = {Machine review of arXiv:2505.05281}
}
abstract
We design an uncertainty-aware cost-sensitive neural network (UA-CSNet) to estimate metallicities from dereddened and corrected Gaia BP/RP (XP) spectra for giant stars. This method accounts for both stochastic errors in the input spectra and the imbalanced density distribution in [Fe/H] values. With a specialized architecture and training strategy, the UA-CSNet improves the precision of the predicted metallicities, especially for very metal-poor (VMP; $\rm [Fe/H] \leq -2.0$) stars. With the PASTEL catalog as the training sample, our model can estimate metallicities down to $\rm [Fe/H] \sim -4$. We compare our estimates with a number of external catalogs and conduct tests using star clusters, finding overall good agreement. We also confirm that our estimates for VMP stars are unaffected by carbon enhancement. Applying the UA-CSNet, we obtain reliable and precise metallicity estimates for approximately 20 million giant stars, including 360,000 VMP stars and 50,000 extremely metal-poor (EMP; $\rm [Fe/H] \leq -3.0$) stars. The resulting catalog is publicly available at https://doi.org/10.12149/101604. This work highlights the potential of low-resolution spectra for metallicity estimation and provides a valuable dataset for studying the formation and chemo-dynamical evolution of our Galaxy.
Figures
Figures from the paper (10 more)
Reference graph
Works this paper leans on
-
[12]
doi:10.3847/1538-4365/ad4a6f 18Yang et al. Huang, Y., Beers, T. C., Wolf, C., et al. 2022, ApJ, 925,
-
[16]
M., Almeida-Fernandes, F., Arentsen, A., et al
doi:10.1088/0004-637X/808/1/16 Placco, V. M., Almeida-Fernandes, F., Arentsen, A., et al. 2022, ApJS, 262, 8. doi:10.3847/1538-4365/ac7ab0 Psaros, A. F., Meng, X., Zou, Z., et al. 2023, Journal of Computational Physics, 477, 111902. doi:10.1016/j.jcp.2022.111902 Recio-Blanco, A., de Laverny, P., Palicio, P. A., et al. 2023, A&A, 674, A29. doi:10.1051/0004...
arXiv 2022
-
[20]
doi:10.3847/0004-637X/833/1/20 York, D. G., Adelman, J., Anderson, J. E., et al. 2000, AJ, 120, 1579. doi:10.1086/301513 Zepeda, J., Beers, T. C., Placco, V. M., et al. 2023, ApJ, 947, 23. doi:10.3847/1538-4357/acbbcc Zhang, X., Green, G. M., & Rix, H.-W. 2023, MNRAS, 524,
-
[34]
doi:10.3847/1538-4365/ab5364 Xu, S., Yuan, H., Niu, Z., et al. 2022, ApJS, 258, 44. doi:10.3847/1538-4365/ac3df6 Xylakis-Dornbusch, T., Christlieb, N., Hansen, T. T., et al. 2024, A&A, 687, A177. doi:10.1051/0004-6361/202348885 Yamada, S., Suda, T., Komiya, Y., et al. 2013, MNRAS, 436, 1362. doi:10.1093/mnras/stt1652 Yao, Y., Ji, A. P., Koposov, S. E., et...
work page Pith review arXiv doi:10.48550/arxiv.2411.19105 2022
-
[43]
doi:10.3847/1538-4357/ad67db Bailer-Jones, C. A. L., Rybizki, J., Fouesneau, M., et al. 2021, AJ, 161, 147. doi:10.3847/1538-3881/abd806 Baumgardt, H. & Vasiliev, E. 2021, MNRAS, 505, 5957. doi:10.1093/mnras/stab1474 Buder, S., Asplund, M., Duong, L., et al. 2018, MNRAS, 478, 4513. doi:10.1093/mnras/sty1281 Buder, S., Sharma, S., Kos, J., et al. 2021, MNRAS, 506,
-
[45]
doi:10.3847/1538-4357/ac9e01 Rockosi, C. M., Lee, Y. S., Morrison, H. L., et al. 2022, ApJS, 259, 60. doi:10.3847/1538-4365/ac5323 Ruz-Mieres, D. 2022, Zenodo Schlegel, D. J., Finkbeiner, D. P., & Davis, M. 1998, ApJ, 500, 525. doi:10.1086/305772 Soubiran, C., Le Campion, J.-F., Cayrel de Strobel, G., et al. 2010, A&A, 515, A111. doi:10.1051/0004-6361/201...
-
[65]
doi:10.3847/1538-4357/ace628 Huang, Y., Beers, T. C., Xiao, K., et al. 2024, ApJ, 974,
-
[132]
doi:10.1088/0004-6256/146/5/132 Leung, H. W. & Bovy, J. 2024, MNRAS, 527, 1494. doi:10.1093/mnras/stad3015 Li, H., Aoki, W., Matsuno, T., et al. 2022, ApJ, 931, 147. doi:10.3847/1538-4357/ac6514 Li, J., Wong, K. W. K., Hogg, D. W., et al. 2024, ApJS, 272, 2. doi:10.3847/1538-4365/ad2b4 Liu, C., Bailer-Jones, C. A. L., Sordo, R., et al. 2012, MNRAS, 426, 2...
arXiv 2024
Show all 16 references
-
[150]
M., Weiler, M., Jordi, C., et al
doi:10.1093/mnras/stab1242 Carrasco, J. M., Weiler, M., Jordi, C., et al. 2021, A&A, 652, A86. doi:10.1051/0004-6361/202141249 Cenarro, A. J., Moles, M., Crist´ obal-Hornillos, D., et al. 2019, A&A, 622, A176. doi:10.1051/0004-6361/201833036 Chiba, M. & Beers, T. C. 2000, AJ, ...
2021 doi
-
[164]
C., Yuan, H., et al
doi:10.3847/1538-4357/ac21cb Huang, Y., Beers, T. C., Yuan, H., et al. 2023, ApJ, 957,
2023 doi
-
[192]
2024, ApJS, 271, 13
doi:10.3847/1538-4357/ad6b94 Huang, B., Yuan, H., Xiang, M., et al. 2024, ApJS, 271, 13. doi:10.3847/1538-4365/ad18b1 Huang, B., Yuan, H., Xu, S., et al. 2025, ApJS, 277, 7. doi:10.3847/1538-4365/ada9e6 Jaehnig, K., Bird, J., & Holley-Bockelmann, K. 2021, ApJ, 923, 129. doi:10...
2024
-
[1159]
2011, MNRAS, 412, 843
doi:10.1093/pasj/60.5.1159 Suda, T., Yamada, S., Katsuta, Y., et al. 2011, MNRAS, 412, 843. doi:10.1111/j.1365-2966.2011.17943.x Suda, T., Hidaka, J., Aoki, W., et al. 2017, PASJ, 69, 76. doi:10.1093/pasj/psx059 Steinmetz, M., Zwitter, T., Siebert, A., et al. 2006, AJ, 132,
2011
-
[1645]
F., Annau, N., McConnachie, A., et al
doi:10.1086/506564 Thomas, G. F., Annau, N., McConnachie, A., et al. 2019, ApJ, 886, 10. doi:10.3847/1538-4357/ab4a77 Viswanathan, A., Starkenburg, E., Matsuno, T., et al. 2024, A&A, 683, L11. doi:10.1051/0004-6361/202347944 Whitten, D. D., Placco, V. M., Beers, T. C., et al. ...
2019 doi
-
[1855]
2024, ApJ, 971, 127
doi:10.1093/mnras/stad1941 Zhang, R., Yuan, H., Huang, B., et al. 2024, ApJ, 971, 127. doi:10.3847/1538-4357/ad613e
2024 doi
-
[2022]
S., Beers, T
doi:10.1088/0004-6256/136/5/2022 Lee, Y. S., Beers, T. C., Sivarani, T., et al. 2008, AJ, 136,
2022 doi
-
[2050]
S., Beers, T
doi:10.1088/0004-6256/136/5/2050 Lee, Y. S., Beers, T. C., An, D., et al. 2011, ApJ, 738, 187. doi:10.1088/0004-637X/738/2/187 Lee, Y. S., Beers, T. C., Masseron, T., et al. 2013, AJ, 146,
2011 doi
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.