{"id":"9cffbe42-96e4-462c-8f4f-d6a66aef18af","arxiv_id":"2411.15342","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A new genetic-algorithm spectral fitting code, tonalli, is validated on synthetic MARCS spectra and the APOGEE-2 Vesta solar spectrum, recovering stellar parameters within half a model-grid step for stars between roughly 3200 and 6250 K.","lead":"The paper introduces tonalli, a Python code that fits APOGEE-2 infrared spectra to synthetic stellar models using an asexual genetic algorithm, and validates it on simulated spectra and on the solar spectrum reflected by Vesta. It reports that the code recovers stellar temperatures, gravities, metallicities and alpha abundances for cool stars, with caveats at the hot end of its range.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The synthetic validation varies only Teff and log(g) while fixing [M/H]=[alpha/M]=0, so the claimed recovery of all four stellar parameters over 3200-6250 K is untested for the two abundance parameters.","rationale":"The reader's weakest_assumption correctly flags that Section 3 validates against synthetic spectra drawn from the same MARCS library and that continuum normalization is a risk. The sharper, more load-bearing point is that even the closed-box experiment never moves [M/H] or [alpha/M] away from zero. The stated central claim -- that tonalli can recover the four stellar parameters Teff, log(g), [M/H], and [alpha/M] with no appreciable bias for 3200-6250 K -- is therefore supported only for the solar-abundance slice of parameter space. The proposed test is a direct extension of the existing M2 setup and would settle whether the abundance-recovery claim holds even under ideal conditions; if it fails, the operating range must be restricted to solar abundances, and if it passes, model fidelity for real non-solar stars remains the next limitation but is outside the paper's closed-box scope. Because this is a gap in the evidence rather than a demonstrated internal inconsistency, it does not force a verdict change beyond the reader's CONDITIONAL assessment, so UNCHANGED is appropriate.","tokens_in":56475,"tokens_out":4644,"duration_ms":49052,"concrete_test":"Repeat the M2-50/M2-100 experiment of Section 3.2 with synthetic MARCS inputs at [M/H] in {-0.5, -0.25, +0.25, +0.5} dex and [alpha/M] in {-0.2, 0.0, +0.2, +0.4} dex, for representative (Teff, log(g)) = (3500, 4.5), (4500, 4.5), (5500, 4.0), (6000, 4.0) K/dex, using the same noise injection, N_rep=50, and the median bias defined by Equation (7). If any [M/H] or [alpha/M] bias exceeds half the MARCS grid step (0.125 dex, since the grid step is 0.25 dex), the claim that tonalli recovers all four stellar parameters over 3200-6250 K is not supported by the current experiments.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.2 describes the accuracy experiments as using 'synthetic spectra with zero metal and alpha-elements abundances' (Teff from 3000-7000 K, log(g)=3, 4, 5 dex, 66 models). Thus the M0/M1/M2 heat maps in Figure 6 and the operating ranges in Table 3 exercise only Teff and log(g) recovery; [M/H] and [alpha/M] are never varied from the solar value in the closed-box test. Yet the abstract and Section 3.2.3 claim that tonalli recovers 'the four stellar parameters' in this temperature range. For real cool stars, abundance errors are a principal concern: line depths scale with metallicity, and molecular opacity can trade off against Teff and log(g), so a pipeline can pass a solar-composition recovery test while being biased for [M/H]=-0.5 or [alpha/M]=+0.3. The Vesta solar spectrum is a single, solar-composition target, and the precise row of Table 5 is selected post hoc against IAU solar values (Equation 10), so it does not fill the gap. The most load-bearing weakness is therefore not just that the models are the same MARCS library, but that two of the four headline parameters are never varied in the validation that establishes the claimed operating range.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents tonalli, a Python code implementing an asexual genetic algorithm (Cantó et al. 2009) to fit APOGEE-2 spectra against a continuum-normalized, resolution-matched MARCS synthetic library. The fitted parameters are Teff, log(g), [M/H], [alpha/M], v sin(i), and RV, with limb darkening fixed at a chosen value. The paper describes the algorithm, the iterative BIC-based sigma-clipping continuum normalization, a k-nearest-neighbour spectral classifier, and a Monte Carlo repetition scheme for uncertainties. Validation is performed with 50 Monte Carlo realizations on 66 synthetic MARCS spectra at solar abundances, in three configurations (M0: already-normalized library spectra; M1: re-normalized spectra; M2: M1 plus noise at several S/N), and on the APOGEE-2 DR17 Vesta solar spectrum with RV fixed and optimized. The central claim is that tonalli can recover all four stellar parameters without appreciable bias for stars with effective temperatures between roughly 3200 and 6250 K.","tokens_in":1949,"tokens_out":2210,"duration_ms":72700,"significance":"If the stated operating range were fully established, tonalli would be a useful open tool for cool-dwarf and pre-main-sequence parameter estimation from APOGEE-2 spectra, a regime in which ASPCAP is known to be less reliable. The paper has real strengths: the algorithm is described in unusual detail with pseudo-code, the parameter-control exploration in Appendix B is systematic, the Monte Carlo uncertainty treatment is honest, and the continuum-normalized MARCS library is publicly released with a DOI. The Vesta analysis openly reports broad credible intervals, multimodal distributions, and model degeneracy. However, the headline accuracy claim is broader than the validation actually supports: the Section 3 synthetic experiment varies only Teff and log(g) while fixing [M/H]=[alpha/M]=0, and all targets are drawn from the same MARCS library used for fitting. These two issues make the four-parameter, 3200-6250 K claim load-bearing and currently under-supported.","major_comments":[{"comment":"The synthetic validation set is described as having \"zero metal and alpha-elements abundances\" with only Teff and log(g) varied (Section 3.2, second paragraph: 66 models with [M/H]=[alpha/M]=0). Consequently the heat maps in Figure 6 and the temperature ranges in Table 3 exercise recovery of Teff and log(g) at solar composition; they do not test recovery of [M/H] or [alpha/M] at non-solar values. The sentence in Section 3.2.3 stating that tonalli \"can recover with success the four stellar parameters in this experiment for stars with effective temperatures between ~3200 and ~6250 K\" is therefore not established for the two abundance parameters. Because abundance changes alter line depths and molecular opacities and can trade off against Teff and log(g), the absence of non-solar-abundance tests is load-bearing. I request adding M1/M2-style experiments at non-zero abundances, for example [M/H] = -0.5, -0.25, +0.25 and [alpha/M] = +0.2 or +0.4, or, failing that, explicitly restricting the claimed operating range to solar-composition recovery.","section":"Section 3.2, Figure 6, Table 3"},{"comment":"All three validation experiments (M0, M1, M2) generate the target spectra from the same MARCS library that tonalli fits against. These experiments therefore measure internal consistency of the interpolator, the optimizer, the continuum-normalization procedure, and the noise handling; they cannot detect systematic errors of the MARCS models relative to real stars. The paper acknowledges model dependence in Section 2.3.8, but the abstract and Section 3.2.3 present the M2 results as accuracy. The solar Vesta comparison is an external anchor, but its default credible intervals are broad (Table 5: Teff = 5779+372/-304 K, log(g) = 4.51+0.40/-0.33, [M/H] = -0.03+/-0.15, [alpha/M] = 0.00+0.16/-0.14). Please either reword the accuracy claims as \"internal recovery within the MARCS grid\" or add independent validation such as a second synthetic grid (PHOENIX, BT-Settl, or BOSZ) or a set of benchmark stars with interferometric or asteroseismic parameters.","section":"Section 3.2 and Section 2.3.8"},{"comment":"The second row of Table 5, with much tighter quoted intervals (Teff = 5780+55/-51 K, log(g) = 4.44+0.06/-0.03, [M/H] = -0.028+0.029/-0.033, [alpha/M] = -0.024+0.020/-0.025), is not the default tonalli result. It is the model selected by Equation (10) as minimizing a chi-square against the IAU solar values among the Appendix B2 models. This is a post hoc model-selection step that uses the same target values as the objective, so the resulting tight intervals quantify proximity to the chosen solar priors rather than independent precision. I advise presenting Equation (10) explicitly as a calibration or selection diagnostic and avoiding the selected row as evidence of nominal accuracy in the abstract or conclusions.","section":"Section 4.3.2, Table 5, Equation (10)"},{"comment":"The abstract states that tonalli efficiently predicts rotational and radial velocities in addition to the four stellar parameters, but Section 3 contains no synthetic recovery test for v sin(i) or RV. The Vesta run reports v sin(i) values near the APOGEE-2 resolution limit and an RV close to the header value, but these are not accuracy tests. Either add synthetic experiments that vary v sin(i) and RV and report their biases, or remove them from the abstract-level claim of predicted parameters.","section":"Section 3, Abstract"}],"minor_comments":[{"comment":"The text defines label 2 as emission-line stars, while the Figure 2 caption refers to \"label 3 (emission line) stars\"; this inconsistency should be corrected.","section":"Figure 2 caption and Section 2.3.3"},{"comment":"The introduction states the method is applicable for 3.0 <= log(g) <= 6.0, but the Section 3.2 synthetic grid uses log(g) = 3, 4, and 5 dex only. Either extend the validation to log(g) = 6 or adjust the stated gravity range.","section":"Section 1 and Section 3.2"},{"comment":"The adopted zero-generation population is given as N0 = 240 in Section 2.3.7 and Table 3, while Appendix B1 states N0 = 250 and refers to \"our adopted input parameters N0 = 240\"; the two values should be reconciled.","section":"Section 2.3.7 and Appendix B"},{"comment":"The text mentions the option \"weigthts\" in KNeighborsClassifier; this is a typo for \"weights\".","section":"Section 2.3.3"}],"recommendation":"major_revision","confidential_remarks":"This is a solid and unusually detailed methods paper, and the authors are transparent about many limitations. I agree with the stress-test assessment in substance: the closed-box, same-library synthetic test and the fixed-solar-abundance setup leave the four-parameter claim under-validated, and the tight solar row in Table 5 is a post hoc selection. These issues are fixable within the scope of the manuscript by extending the synthetic grid to non-solar abundances and by explicitly reframing the claims and Table 5. I do not see a novelty or citation-practice problem; the related work on IN-SYNC and APOGEENet is properly cited."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: tonalli is a genuine, carefully tested fitting tool for APOGEE-2 spectra, but the claimed 3200–6250 K, four-parameter operating range oversells what the validation actually exercises. The synthetic experiments vary only Teff and log(g) at fixed solar abundances, so the paper does not demonstrate unbiased recovery of [M/H] and [alpha/M] across that range.\n\nWhat's new: the concrete implementation of Cantó's asexual genetic algorithm for APOGEE-2 spectral fitting, with a BIC-based continuum normalisation, a k-NN classifier for early-type rejection, and a Monte Carlo uncertainty scheme. The synthetic validation design is thoughtful: baseline normalised library vs un-normalised vs noise-added, with 50 repetitions per model, and bias/precision heat maps in fractions of grid step. The solar Vesta test is an honest external anchor, and the median recovered solar parameters agree with DR17 within the interquartile ranges. The continuum-normalised MARCS library is released on Zenodo, which is useful.\n\nSoft spots, in proportion. The biggest is the abundance question. The stress-test note is right: Section 3.2 uses only zero-metallicity and zero-alpha spectra, so the four-parameter claim in the abstract and Section 3.2.3 is not supported by that experiment. For cool stars, metallicity and alpha abundance correlate with temperature and gravity through line and molecular opacities, so a solar-composition recovery test is weak evidence. The Vesta spectrum is a single solar-composition target and the post-hoc row in Table 5 (selected with Eq. 10 against IAU values) is a tuned example, not a prediction; it should be clearly labelled as such. The main Vesta result also has large IQRs (Teff ±~300 K, log(g) ±~0.4 dex), which is honest but limits practical precision. And because the synthetic input is drawn from the same MARCS library used for fitting, the validation is internal consistency; the only external check is the Sun. Finally, the tonalli source code is not released, only the library, which hampers reproducibility.\n\nBottom line: a solid methods paper for a specific niche — APOGEE-2 characterisation of cool dwarfs and pre-main-sequence stars where ASPCAP is weak. It deserves serious refereeing, but the operating-range wording must be aligned with the experiments, the abundance-parameter recovery needs either a new test or a softened claim, and the code should be released. I'd suggest major revision, not rejection.","headline":"A careful APOGEE-2 fitting pipeline whose headline four-parameter operating range is not backed by the synthetic tests, because only Teff and log(g) are varied at solar abundances.","tokens_in":57338,"tokens_out":2749,"would_cite":false,"duration_ms":26636,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An asexual genetic algorithm extracts four stellar parameters from APOGEE-2 spectra with no appreciable bias for stars from 3200 to 6250 K.","keywords":["asexual genetic algorithm","APOGEE-2 stellar spectra","stellar atmospheric parameters","MARCS synthetic spectra","pre-main-sequence stars","spectrum fitting","Monte Carlo uncertainties","continuum normalization"],"falsifier":"Apply tonalli to APOGEE-2 spectra of benchmark cool stars with independent effective temperatures from interferometry and surface gravities from asteroseismology, spanning 3200 to 6250 K, and check whether the recovered medians agree with the independent values within the claimed bias of less than half a MARCS grid step; a systematic offset that grows as $T_{\\rm eff}$ approaches 6250 K, or an offset on real stars that is absent in the synthetic recovery tests, would refute the working-range claim.","tokens_in":56280,"feed_emoji":"⭐","tokens_out":8745,"duration_ms":71893,"temperature":0.7,"pith_summary":"This paper presents tonalli, a Python code that pulls effective temperature ($T_{\\rm eff}$), surface gravity ($\\log(g)$), overall metallicity ($[{\\rm M/H}]$), and $\\alpha$-element abundance ($[\\alpha/{\\rm M}]$) out of APOGEE-2 near-infrared spectra by direct fitting against an interpolated MARCS synthetic library. The central claim under validation is that the code recovers all four stellar parameters with no appreciable bias for stars with $T_{\\rm eff}$ between roughly 3200 and 6250 K, demonstrated on noisy synthetic spectra and on the APOGEE-2 solar spectrum reflected by Vesta. If that claim holds, the code offers a model-based parameter route for cool dwarfs and subgiants, including pre-main-sequence stars, where the standard APOGEE-2 pipeline was not optimized. The paper also details a Monte Carlo repetition scheme that converts the stochastic optimizer output into per-star credible intervals, and it releases the continuum-normalized MARCS library it constructs.","feed_headline":"Asexual genetic search recovers cool-star parameters from APOGEE-2","feed_subtitle":"Direct synthetic fitting with little bias for 3200-6250 K stars, targeting cool dwarfs and pre-main-sequence stars standard pipelines miss.","key_machinery":"The load-bearing mechanism is the asexual genetic algorithm of Cantó et al. (2009), which differs from classical genetic algorithms by generating each new generation within a shrinking hyper-cube around each of the $N_p$ fittest parents, with side lengths decreasing as $(\\Delta x_i)_n = (\\Delta x_i)_0 \\, p^n$ for a convergence factor $p\\in(0,1)$. Each offspring's synthetic spectrum is formed by 4D linear interpolation over the nearest $N_{\\rm interpol}$ MARCS grid points, convolved to APOGEE-2 resolution, rotationally broadened, and Doppler-shifted, and its fitness is the $\\chi^2$ between that spectrum and the observed one. Before fitting, both observed and synthetic spectra are mapped onto a common pseudo-continuum by an iterative BIC-selected polynomial fit with asymmetric $\\sigma$-clipping, and the wavelength window for the figure of merit excludes the chip edges where the Brackett lines distort the normalization. A $k$-nearest-neighbours classifier on Mg i, Al i, and CO equivalent widths gates high-temperature and emission-line stars, and the whole fine search is repeated in a Monte Carlo loop so that the reported parameters are the median and interquartile range of the resulting distributions.","core_discovery":"On its own terms, the paper's contribution is the demonstration that an asexual genetic algorithm can reliably locate the best-fitting MARCS model for an APOGEE-2 spectrum without appreciable bias in a defined working range. In the noisy synthetic experiments, the reported bias stays below half the MARCS grid step for the four parameters over roughly 3200 to 6250 K, with the temperature bias growing at higher temperatures because the continuum normalization degrades near the Brackett lines at the chip edges. For the Vesta solar spectrum, the adopted univariate medians are $T_{\\rm eff}=5779$ K, $\\log(g)=4.51$ dex, $[{\\rm M/H}]=-0.03$ dex, and $[\\alpha/{\\rm M}]=0.00$ dex, with credible intervals on the order of a few hundred kelvin and a few tenths of a dex, and the paper finds these consistent with published spectroscopic determinations of the Sun.","pith_inferences":["Beyond the paper, the synthetic validation uses spectra drawn from the same MARCS library that tonalli fits, so the reported accuracy is primarily internal consistency; a cross-check against asteroseismic or interferometric benchmarks for cool dwarfs would test the model-fidelity part of the claim.","The paper's own Vesta experiment shows parameter degeneracies, including a shift in $\\log(g)$ when the radial velocity is optimized, so adding priors as the authors say they plan could shrink credible intervals substantially without changing the method.","The continuum-normalization failure above roughly 6250 K is tied to Brackett lines near chip edges, so restricting the figure-of-merit windows or adopting a different normalization for early-type stars might extend the same machinery to hotter stars.","The classifier gating is trained on young stars in the Pleiades and W3/4/5, so applying tonalli to field cool stars assumes those equivalent-width classes transfer; re-training on field stars would harden the pipeline for general surveys."],"forward_implications":["APOGEE-2 spectra of cool dwarfs and subgiants can be parameterized entirely by direct synthetic fitting, with no training labels and with per-star Monte Carlo uncertainties.","The working range from 3200 to 6250 K includes pre-main-sequence stars, giving a model-based route for studying young stellar populations where the standard APOGEE-2 pipeline is known to struggle.","The released continuum-normalized MARCS library and the user-selectable library interface allow the same pipeline to be re-run against other model grids, such as BT-NextGen, BOSZ, PHOENIX, or SpecModels, to expose model dependence.","The stopping criterion based on the temperature hyper-volume shrinking to 1 K, together with the recommended input parameters, provides a reproducible recipe for parameter recovery.","The median of the one-dimensional Monte Carlo distribution is argued to be a robust statistic when parameter degeneracies make the distribution multimodal, giving a concrete reporting convention."],"supporting_citations":[{"why":"Supplies the asexual genetic algorithm that carries the optimization.","marker":"Cantó et al. (2009)"},{"why":"Provides the MARCS model atmosphere grid that is the basis of the synthetic comparison library.","marker":"Gustafsson et al. (2008)"},{"why":"Provides the MARCS synthetic spectra and the solar reference parameters used for the library and comparisons.","marker":"Jönsson et al. (2020)"},{"why":"Provides the APOGEE-2 DR17 data, including the Vesta spectrum, and the ASPCAP parameters the solar result is compared with.","marker":"Abdurro'uf et al. (2022)"},{"why":"Defines the IN-SYNC direct-fitting methodology and the rotational broadening resolution limit that tonalli builds on.","marker":"Cottaar et al. (2014)"},{"why":"Presents the data-driven APOGEE Net approach whose young-star parameter estimates motivate the need for direct synthetic fitting.","marker":"Sprague et al. (2022)"},{"why":"Supplies the SciPy griddata routine used for the 4D interpolation of each offspring spectrum.","marker":"Virtanen et al. (2020)"},{"why":"Defines the APOGEE pixel masks and data quality flags used in cleaning the observed spectra.","marker":"Holtzman et al. (2015)"}],"fun_headline_variants":["Asexual GA fits APOGEE-2 cool stars with little bias","tonalli: GA decodes APOGEE-2 cool stars","Genetic search fits 3200-6250 K stars from APOGEE-2","Low-bias stellar parameters via asexual genetic algorithm","tonalli: asexual GA finds best MARCS fit for APOGEE-2"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the iterative continuum normalization places real and synthetic spectra on a common pseudo-continuum across all three APOGEE-2 chips; if it distorts line-to-continuum ratios, especially near the Brackett lines at the chip edges, the chi-squared fit is biased, and the synthetic validation against the same MARCS library measures internal consistency rather than the fidelity of the models to real stars.","fun_headline_variants_meta":{"raw":{"variants":["Asexual GA fits APOGEE-2 cool stars with little bias","tonalli: GA decodes APOGEE-2 cool stars","Genetic search fits 3200-6250 K stars from APOGEE-2","Low-bias stellar parameters via asexual genetic algorithm","tonalli: asexual GA finds best MARCS fit for APOGEE-2"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000552,"raw_usage":{"total_tokens":2628,"prompt_tokens":936,"completion_tokens":1692,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":552,"completion_tokens_details":{"reasoning_tokens":1609}},"tokens_in":552,"tokens_out":1692,"duration_ms":11635,"temperature":1.0,"reasoning_tokens":1609,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:24:42.044317+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply tonalli to APOGEE-2 spectra of benchmark cool stars with independent effective temperatures from interferometry and surface gravities from asteroseismology, spanning 3200 to 6250 K, and check whether the recovered medians agree with the independent values within the claimed bias of less than half a MARCS grid step; a systematic offset that grows as $T_{\\rm eff}$ approaches 6250 K, or an offset on real stars that is absent in the synthetic recovery tests, would refute the working-range claim.","supporting_citations":[{"cited_title":"G., Nordlund , A ., & Plez , B., 2008","cited_arxiv_id":null,"evidence_quote":"Provides the MARCS model atmosphere grid that is the basis of the synthetic comparison library."},{"cited_title":"R., Hutchinson , B., Lingg , R., Stassun , K","cited_arxiv_id":null,"evidence_quote":"Presents the data-driven APOGEE Net approach whose young-star parameter estimates motivate the need for direct synthetic fitting."},{"cited_title":"E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt , S","cited_arxiv_id":null,"evidence_quote":"Supplies the SciPy griddata routine used for the 4D interpolation of each offspring spectrum."},{"cited_title":"A., Shetrone , M., Johnson , J","cited_arxiv_id":null,"evidence_quote":"Defines the APOGEE pixel masks and data quality flags used in cleaning the observed spectra."}],"review_version":1}