REVIEW 3 major objections 6 minor 1 cited by
This paper shows that an incomplete, deliberately selected spectroscopic survey can be converted into unbiased population statistics for the 100-parsec solar neighborhood by explicitly modeling the selection function and a new subpopulation
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 12:34 UTC pith:5QYHYQH7
load-bearing objection Useful selection-function and forward-modeling toolkit for SDSS-V SNC, but a real prior-specification bug over-predicts rare subpopulations in sparse HR bins. the 3 major comments →
A Probabilistic Framework for Population Studies of the Solar Neighborhood: Application to SDSS-V and Gaia
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that the effective selection function of the SDSS-V Solar Neighborhood Census relative to the Gaia Catalog of Nearby Stars can be decomposed into layers — carton targeting, fiber assignment, actual observation, and successful recovery of the spectroscopic quantity — and that the compound probability of surviving all layers is well described by Beta-distributed counts in bins of sky position, G magnitude, and BP−RP color. Against this backdrop, the authors propose that every spectroscopically defined subpopulation can be represented by a 'subpopulation probability' grid over the HR diagram, p_sub,k, and that the expected number of SDSS-V stars in a bin is the prod
What carries the argument
The central object is the subpopulation probability p_sub,k — a cell-wise posterior estimate of the chance that a GCNS star in HR-diagram bin k belongs to the spectroscopic class of interest. It enters a Poisson point-process intensity via Λ_obs,j = Σ_k A_{j,k} p_sub,k, where A_{j,k}, the 'effective selection factor', is the sum of detection probabilities (the analysis-dependent selection function) over all GCNS stars in selection-bin j and HR-bin k. The selection function itself is built as a Beta posterior from counts (k out of n GCNS stars observed in a bin), with bins in sky-position pixels, Gaia G magnitude, and BP−RP color; the framework enforces zero probability in unvisited bins. The
Load-bearing premise
The load-bearing premise is that a star's position on the color–magnitude diagram fully encodes the probability that it belongs to a given spectroscopically defined subpopulation; unresolved binaries and overlapping stellar populations (e.g., young pre-main-sequence versus old metal-rich stars) can break that link and skew the inferred subpopulation probabilities.
What would settle it
Run the released forward model on a mock Gaia sample in which a known fraction of stars are unresolved binaries whose combined light shifts them redder and overluminous, and check whether the recovered subpopulation probabilities in the affected color–magnitude bins match the true input fractions within the posterior width; a systematic mismatch would show the HR-position-only mapping is insufficient.
If this is right
- Any science case built on SDSS-V SNC data can define its own effective sample — targeted, assigned, observed, successfully reduced — and the same selection-function machinery corrects for all four layers of incompleteness.
- H-alpha emission mapping across the HR diagram shows the emission fraction rises sharply at the lowest masses, shifting the peak of the luminosity function from M_G ≈ 11 to ≈ 12, a population-level statement now available from a deliberately incomplete survey.
- Number-density measurements in metallicity bins reveal that low-mass stellar density increases with metallicity and that the mass-function slope changes with metallicity, setting up direct IMF-metallicity tests.
- The mock validation demonstrates that arbitrary subpopulation patterns across the HR diagram can be recovered from incomplete spectroscopy, with the caveat that low-count regions remain prior-dominated.
- Because the framework resamples distance posteriors and incorporates the GCNS completeness and 1/Vmax corrections, it yields volume-complete luminosity and mass functions for the 100-pc sample.
Where Pith is reading between the lines
- We infer the same machinery transfers to other fiber-fed spectroscopic surveys with algorithmic target assignment: once the parent catalog's completeness is known, the layered selection (carton → fiber assignment → observation → measurement success) can be characterized with the same Beta-count prescription.
- We expect the framework's most productive extension is a hierarchical version that treats unresolved binaries as a distinct latent population rather than folding them into the HR-diagram mapping; that would directly test the binary-bias concern the paper flags in its limitations section.
- The prior-dominance flags (KL divergence, variance shift, mean shift) double as a survey-design tool: they reveal HR-diagram cells where additional spectroscopic observations would add the most information about subpopulation membership.
- We speculate the selection function could be validated epoch-by-epoch: recomputing it on successive data releases should show detection probabilities converging field-by-field as scheduling gaps fill, and any jumps would indicate plan revisions the current static model cannot absorb.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a probabilistic framework for using the incomplete SDSS-V Solar Neighborhood Census (SNC) sample together with the Gaia Catalog of Nearby Stars (GCNS) 100 pc sample. It models an analysis-dependent selection function as a Beta posterior in bins of HEALPix sky position, Gaia G magnitude, and BP−RP color; introduces a 'subpopulation probability' grid p_sub,k across the HR diagram; and fits this grid with a Poisson point-process likelihood using JAX/numpyro MCMC. The fitted probabilities are then used to resample GCNS stars, enabling volume-corrected densities. The framework is validated on a mock subpopulation and illustrated with DR19 data on Hα emission and on stellar number density versus mass and metallicity. The paper also releases public code and reproducible analysis scripts.
Significance. If the framework performs as claimed, it would be a valuable community tool: it directly addresses the severe selection effects of SDSS-V's SNC and offers a practical path from an incomplete spectroscopic sample to population-level inferences tied to the nearly complete GCNS. The public code, reproducible scripts, and the self-consistency mock test are real strengths, and the selection function reproduces the main qualitative features of the survey planning logic. The central promise of unbiased, statistically robust population inference is, however, undermined by an internally inconsistent prior specification in Eqs. (9)–(10) that biases exactly the rare-subpopulation regime the method is designed for. The issue is local and fixable, but until it is addressed the paper's headline claims are not supported.
major comments (3)
- [§3.2 (Eqs. 9–10)] The Beta prior is claimed to have mean equal to the observed fraction f_o, but the cap β=2 breaks this for f_o < 1/3. For example, f_o = 0.01 yields Beta(1,2), whose mean is 1/3 and whose 95% interval is [0.013, 0.84]. Thus in HR bins with zero observed subpopulation stars, the posterior is essentially the prior, over-predicting counts by an order of magnitude relative to f_o. The mock validation in §4.1 shows an 'overestimation in low number regions' but attributes it to small-number statistics rather than to the prior. The rare-population applications (e.g., the metal-poor bins in §4.3) are exactly in this regime. The prior should be reparameterized so that its mean is f_o (for instance β=(1−f_o)/f_o, or a Beta(α=c f_o, β=c(1−f_o)) form) and the validation rerun.
- [§3.1 (k=0 rule) and §4.2–4.3] Setting S=0 whenever k=0 conflates unvisited survey fields with bins that are visited but contain no members of a rare subpopulation. For analysis-dependent subsamples (Hα emitters, metallicity cuts), a bin can contain GCNS stars and SDSS-V spectra but zero subpopulation members purely because the subpopulation is rare. The strict-zero rule then removes those bins from the Poisson likelihood, so p_sub,k in them is unconstrained and collapses to the prior, compounding the bias in the previous comment. Please distinguish unvisited bins from zero-detection bins and, in the latter case, use a nonzero Beta posterior (e.g., Beta(1, n+1)) or justify why S=0 is appropriate for rare subpopulations.
- [§4.1 and §4.3] The mock validation uses a sinusoidal subpopulation that is likely not rare, so it does not stress the prior-dominated regime that is central to the paper's motivating examples. The paper should include a rare-subpopulation mock (global fraction at the few-percent level) and report the bias in low-count bins separately from the width of the 95% envelope. The current statement that 'the forward model correctly recovers the true distribution when the errors are considered' is too weak to support the unbiasedness claim in the presence of a prior that is known to be biased.
minor comments (6)
- [§3.2] The sentence 'a minimum of β=2 so the result is never a uniform distribution' is incorrect: for f_o=0.5, β=(1−f_o)/f_o=1, and Beta(1,1) is the uniform distribution.
- [Eq. (18) / Appendix B] The calibration relation is described as a 'second degree polynomial' but contains a T^3 term; it is a cubic polynomial.
- [§4.3] The binary-removal criterion 'ipd_frac_multi_peak≥0' is a no-op because this quantity is non-negative by definition; the intended cut was presumably '>0'.
- [Appendix A] Minor typos: 'evalutate_Ajk' should be 'evaluate_Ajk', and the column name 'e_fe_h' appears inconsistently in the code example.
- [Figures 4–6] '95% confidence' should be '95% credible interval' or 'posterior percentile range' to match the Bayesian framework.
- [§4.1 validation scope] The mock test is a posterior-predictive check; it validates the statistical machinery only under the assumption that the HR-diagram mapping is correct. The paper should state explicitly that the mock does not independently validate that mapping, a limitation already acknowledged in §5.
Circularity Check
No significant circularity: selection function and subpopulation inference are empirically fitted, and the mock validation is a genuine recovery test; the data-dependent prior is a bias concern, not a circular reduction.
full rationale
The central derivation chain is not circular. The subsample selection function is empirically estimated as a Beta posterior over GCNS/SDSS-V counts (Eqs. 1–3), following the external Castro-Ginard et al. framework; the statement that the detection probabilities reproduce survey planning logic is a consistency check against independent operational knowledge, not a fitted constraint. The subpopulation probabilities p_sub,k are free parameters in a Poisson forward model (Eqs. 4–8); the GCNS draws made from the posterior are posterior predictive summaries of the fit, and the mock validation compares these against an independently defined mock truth, which is a genuine recovery test. The data-dependent prior in Eqs. 9–10 does mean that under-sampled HR bins remain prior-dominated, and the β cap makes the prior mean 1/3 for f_o<1/3 rather than f_o; this is an internal-consistency/bias issue that the paper itself acknowledges (Section 5 and the prior-flagging masks), but it is not a reduction of a prediction to its own input. The only overlapping-author citations (Way et al. 2026; Medan et al. 2025) enter as data-filtering and survey-context references, not as load-bearing derivation. Thus no quoted equation reduces to itself by construction.
Axiom & Free-Parameter Ledger
free parameters (4)
- p_sub,k (HR-diagram subpopulation probability grid) =
posterior samples per bin; not tabulated in paper
- Prior hyperparameters f_o and beta =
f_o = max(N_subpop/N_data, 0.01); beta = min((1-f_o)/f_o, 2)
- ASPCAP metallicity correction coefficients =
-6.095e-11, 9.164e-7, -4.564e-3, 7.538
- Binning scheme =
HEALPix order 3; dG=0.5 mag; d(BP-RP)=0.15 mag; dM_G=0.25 mag
axioms (7)
- domain assumption GCNS is complete enough that its remaining incompleteness (bright end, late M/L/T dwarfs, close binaries) can be ignored or corrected in post-processing
- domain assumption HR-diagram position is a probabilistic predictor of spectroscopic subpopulation membership
- domain assumption Selection probability is constant within (HEALPix, G, BP-RP) bins and follows Beta-Binomial counting
- domain assumption Observed subpopulation stars can be modeled as a Poisson process on the GCNS catalog with per-star intensity p_sub,k x S
- domain assumption Volume corrections use a thin-disk exponential density law with H=365 pc and z_sun=17 pc
- domain assumption GCNS distance posterior samples and 80th-percentile magnitude limits are accurate inputs
- domain assumption Bins with k=0 have strictly zero selection probability, interpreted as unvisited fields
Cite this review
Pith. "Pith review of A Probabilistic Framework for Population Studies of the Solar Neighborhood: Application to SDSS-V and Gaia." pith.science (2026). https://pith.science/paper/5QYHYQH7
@misc{pith2026260719501,
author = {Pith},
title = {Pith review of: A Probabilistic Framework for Population Studies of the Solar Neighborhood: Application to SDSS-V and Gaia},
year = {2026},
howpublished = {\url{https://pith.science/paper/5QYHYQH7}},
note = {Machine review of arXiv:2607.19501}
}
read the original abstract
Studies of the Solar Neighborhood require spectroscopic follow-up of stars identified in astrometric surveys to fully characterize their physical properties. The SDSS-V Solar Neighborhood Census (SNC) is a dedicated program to observe stars within 100~pc. However, due to competing observing programs and fiber assignment constraints, the resulting sample carries severe and complex selection effects. A framework is presented for characterizing the selection function of the SDSS-V SNC relative to the Gaia Catalog of Nearby Stars (GCNS), along with a forward modeling method to infer the properties of stellar subpopulations across the GCNS-defined 100 pc sample. The selection function is based on a method that models the selection probability as a function of sky position, Gaia G magnitude, and BP-RP. The resulting detection probabilities faithfully reproduce the known survey planning logic. This work further introduces the concept of a "subpopulation probability" -- a grid of posterior estimates across the Hertzsprung-Russell (HR) diagram representing the likelihood that a GCNS member belongs to a given SDSS-V defined subpopulation. The framework is validated with a mock dataset and its scientific utility is demonstrated through two applications using data from the Data Release 19: mapping H$\alpha$ emission across the HR diagram and measuring the variation of stellar density with mass and metallicity. These results illustrate how statistically robust population studies can be conducted with an incomplete spectroscopic survey when the selection function is well characterized. The code is made publicly available with this work, which will serve as an important tool for future studies.
Figures
Forward citations
Cited by 1 Pith paper
-
The Twentieth Data Release of the Sloan Digital Sky Survey: First All-Sky BOSS Spectra, eROSITA-SDSS-V Mapper Coordinated Observations, and a Preview of the Local Volume Mapper
DR20 releases over three million BOSS spectra (first southern-hemisphere SDSS-V optical data), 169 LVM integral-field tiles over six targets, and eighteen value-added catalogs.
Reference graph
Works this paper leans on
-
[1]
F., Argudo-Fernández, M., et al
Almeida, A., Anderson, S. F., Argudo-Fernández, M., et al. 2023, ApJS, 267, 44, doi: 10.3847/1538-4365/acda98
-
[2]
Blanton, M. R., Carlberg, J. K., Dwelly, T., et al. 2025, arXiv e-prints, arXiv:2505.21328, doi: 10.48550/arXiv.2505.21328
-
[3]
Bowen, I. S., & Vaughan, A. H., J. 1973, ApOpt, 12, 1430, doi: 10.1364/AO.12.001430
-
[4]
2018, JAX: composable transformations of Python+NumPy programs, 0.3.13
Bradbury, J., Frostig, R., Hawkins, P., et al. 2018, JAX: composable transformations of Python+NumPy programs, 0.3.13. http://github.com/jax-ml/jax
2018
-
[5]
2012, Monthly Notices of the Royal Astronomical Society, 427, 127
Bressan, A., Marigo, P., Girardi, L., & et al. 2012, Monthly Notices of the Royal Astronomical Society, 427, 127
2012
-
[6]
2023, A&A, 669, A55, doi: 10.1051/0004-6361/202244784
Cantat-Gaudin, T., Fouesneau, M., Rix, H.-W., et al. 2023, A&A, 669, A55, doi: 10.1051/0004-6361/202244784
-
[7]
Castro-Ginard, A., Brown, A. G. A., Kostrzewa-Rutkowska, Z., et al. 2023, A&A, 677, A37, doi: 10.1051/0004-6361/202346547
-
[8]
2016, The Astrophysical Journal, 823, 102 18
Choi, J., Dotter, A., Conroy, C., & et al. 2016, The Astrophysical Journal, 823, 102 18
2016
-
[9]
El-Badry, K., Rix, H.-W., & Heintz, T. M. 2021, MNRAS, 506, 2269, doi: 10.1093/mnras/stab323 ESA, ed. 1997, ESA Special Publication, Vol. 1200, The HIPPARCOS and TYCHO catalogues. Astrometric and photometric star catalogues derived from the ESA HIPPARCOS Space Astrometry Mission
-
[10]
2021, A&A, 649, A5, doi: 10.1051/0004-6361/202039834
Fabricius, C., Luri, X., Arenou, F., et al. 2021, A&A, 649, A5, doi: 10.1051/0004-6361/202039834
-
[11]
Felten, J. E. 1976, ApJ, 207, 700, doi: 10.1086/154538 Gaia Collaboration, Smart, R. L., Sarro, L. M., et al. 2021, A&A, 649, A6, doi: 10.1051/0004-6361/202039498 Gaia Collaboration, Vallenari, A., Brown, A. G. A., et al. 2023, A&A, 674, A1, doi: 10.1051/0004-6361/202243940
-
[12]
2023, A&A, 670, A19, doi: 10.1051/0004-6361/202244250
Golovin, A., Reffert, S., Just, A., et al. 2023, A&A, 670, A19, doi: 10.1051/0004-6361/202244250
-
[13]
Gunn, J. E., Siegmund, W. A., Mannery, E. J., et al. 2006, AJ, 131, 2332, doi: 10.1086/500975
doi:10.1086/500975 2006
-
[14]
Henry, T. J., Jao, W.-C., Subasavage, J. P., et al. 2006, AJ, 132, 2360, doi: 10.1086/508233
doi:10.1086/508233 2006
-
[15]
Henry, T. J., Jao, W.-C., Winters, J. G., et al. 2018, AJ, 155, 265, doi: 10.3847/1538-3881/aac262
-
[16]
Hoffman, M. D., & Gelman, A. 2011, arXiv e-prints, arXiv:1111.4246, doi: 10.48550/arXiv.1111.4246
-
[17]
Karim, T., & Mamajek, E. E. 2017, MNRAS, 465, 472, doi: 10.1093/mnras/stw2772
-
[18]
Kirkpatrick, J. D., Marocco, F., Gelino, C. R., et al. 2024, ApJS, 271, 55, doi: 10.3847/1538-4365/ad24e2
-
[19]
A., Rix, H.-W., Aerts, C., et al
Kollmeier, J. A., Rix, H.-W., Aerts, C., et al. 2026, AJ, 171, 52, doi: 10.3847/1538-3881/ae0576
-
[20]
E., Allende Prieto, C., Cooper, A
Koposov, S. E., Allende Prieto, C., Cooper, A. P., et al. 2024, MNRAS, 533, 1012, doi: 10.1093/mnras/stae1842
-
[21]
2001, MNRAS, 322, 231, doi: 10.1046/j.1365-8711.2001.04022.x
Kroupa, P. 2001, MNRAS, 322, 231, doi: 10.1046/j.1365-8711.2001.04022.x
arXiv 2001
-
[22]
Kullback, S., & Leibler, R. A. 1951, The annals of mathematical statistics, 22, 79
1951
-
[23]
2023, Nature, 613, 460, doi: 10.1038/s41586-022-05488-1
Li, J., Liu, C., Zhang, Z.-Y., et al. 2023, Nature, 613, 460, doi: 10.1038/s41586-022-05488-1
-
[24]
Lutz, T. E., & Kelker, D. H. 1973, PASP, 85, 573, doi: 10.1086/129506
doi:10.1086/129506 1973
-
[25]
Mann, A. W., Dupuy, T., Kraus, A. L., et al. 2019, ApJ, 871, 63, doi: 10.3847/1538-4357/aaf3bc
-
[26]
2026, imedan/snc_sf: 1.1.1, 1.1.1, Zenodo, doi: 10.5281/zenodo.21285596
Medan, I. 2026, imedan/snc_sf: 1.1.1, 1.1.1, Zenodo, doi: 10.5281/zenodo.21285596
-
[27]
Medan, I., Dwelly, T., Covey, K. R., et al. 2025, arXiv e-prints, arXiv:2506.15475, doi: 10.48550/arXiv.2506.15475
-
[28]
2019, arXiv preprint arXiv:1912.11554
Phan, D., Pradhan, N., & Jankowiak, M. 2019, arXiv preprint arXiv:1912.11554
Pith/arXiv arXiv 2019
-
[29]
W., Derwent, M
Pogge, R. W., Derwent, M. A., O’Brien, T. P., et al. 2020, in Society of Photo-Optical Instrumentation Engineers (SPIE) Conference Series, Vol. 11447, Ground-based and Airborne Instrumentation for Astronomy VIII, ed. C. J
2020
-
[30]
Evans, J. J. Bryant, & K. Motohara, 1144781, doi: 10.1117/12.2561113
-
[31]
Qiu, D., Johnson, J. A., Liu, C., et al. 2025, arXiv e-prints, arXiv:2511.20005, doi: 10.48550/arXiv.2511.20005
-
[32]
Riello, M., De Angeli, F., Evans, D. W., et al. 2021, A&A, 649, A3, doi: 10.1051/0004-6361/202039587
-
[33]
Rix, H.-W., Hogg, D. W., Boubert, D., et al. 2021, AJ, 162, 142, doi: 10.3847/1538-3881/ac0c13
-
[34]
2024, AJ, 167, 125, doi: 10.3847/1538-3881/ad2001
Saad, S., Lane, K., Kounkel, M., et al. 2024, AJ, 167, 125, doi: 10.3847/1538-3881/ad2001
-
[35]
1968, ApJ, 151, 393, doi: 10.1086/149446 SDSS Collaboration, Adamane Pallathadka, G.,
Schmidt, M. 1968, ApJ, 151, 393, doi: 10.1086/149446 SDSS Collaboration, Adamane Pallathadka, G.,
doi:10.1086/149446 1968
-
[36]
2025, arXiv e-prints, arXiv:2507.07093, doi: 10.48550/arXiv.2507.07093
Aghakhanloo, M., et al. 2025, arXiv e-prints, arXiv:2507.07093, doi: 10.48550/arXiv.2507.07093
-
[37]
A., Gunn, J
Smee, S. A., Gunn, J. E., Uomoto, A., et al. 2013, The Astronomical Journal, 146, 32
2013
-
[38]
Tinney, C. G., Reid, I. N., & Mould, J. R. 1993, ApJ, 414, 254, doi: 10.1086/173074
-
[39]
2026, The Astronomical Journal, 171, 252, doi: 10.3847/1538-3881/ae48ae
Way, Z., Lépine, S., Gagné, J., & Medan, I. 2026, The Astronomical Journal, 171, 252, doi: 10.3847/1538-3881/ae48ae
-
[40]
Wilson, J. C., Hearty, F. R., Skrutskie, M. F., et al. 2019, PASP, 131, 055001, doi: 10.1088/1538-3873/ab0075 19 APPENDIX A.SELECTION FUNCTION CODE EXAMPLE Below, we will go through a simple example of how to use the code associated with this work22 (Medan 2026). For this example, we will only consider the APOGEE SNC stars from DR19. The SDSS-V data neede...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.