{"id":"033460c4-5aed-4adb-ad43-8b4207e0e2d6","arxiv_id":"2504.17600","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Age estimates for red giants from different evolution models diverge by up to about 90 percent when mass is unconstrained, while known-mass stars agree to roughly 10 percent.","lead":"This paper compares four common stellar evolution model grids and finds that inferred ages for red giants can differ by more than 80 percent when a star's mass is not known, while known-mass stars agree to about 10 percent. It offers a practical recipe for adding model disagreement to age uncertainties in galactic archaeology.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"MIST's downsampled EEP tracks could inflate the headline 80–90% age offsets near the RGB bump; the paper never validates that resampling preserves MIST's native age–Teff–log g relation in that region.","rationale":"The central claim is partly quantitative: unknown mass turns ~10% model agreement into 60–90% disagreement, and the paper's proposed uncertainty recipe uses that spread as a theoretical error bar. The paper's own §2.1 contains the seed of the weakest link: the authors added EEPs at the late end for three grids but could not do so for MIST, whose full-resolution tracks are not public. The region that matters most—the RGB bump and the tip—is exactly where the largest offsets are reported in Figures 3 and 4. If MIST's default EEP spacing is too coarse there, kiauhoku's interpolation can distort MIST's track shape, changing its inferred mass (and hence age) at fixed Teff–log g or Teff–L. Because the headline metric is a maximum over grids, one biased grid can set the number. This concern lands on the quantitative values and on the uncertainty prescription, not on the qualitative finding: the APOKASC-3 spectroscopic averages (~80%) and the synthetic mean offsets (51–69%) are large even if MIST's contribution is partly numerical, and the known-mass case consistently shows ~10% agreement. I therefore keep the reader's CONDITIONAL verdict. Credit is due for sharing reproducible notebook code, publishing the modified grids on Zenodo, and explicitly flagging the MIST resolution limitation; those facts make the issue testable rather than fatal. A direct re-run with full-resolution MIST is the single check that would settle whether the headline 80–90% offsets are physical. There is also a minor internal numerical slip in the Conclusion ('reaching a mean offset of 90%') versus the Figure 3 caption means (51–69%) and APOKASC means (~80%), but that does not change the verdict.","tokens_in":22950,"tokens_out":7481,"duration_ms":75788,"concrete_test":"Regenerate the MIST grid at full resolution with the public MESA inlists (or obtain full-resolution history files from the authors), resample it with kiauhoku using the same dense late-stage EEP scheme applied to YREC/DSEP/GARSTEC, and recompute the maximum fractional age offsets in Figure 3 and Figure 4 for the post-bump region (log g ≲ 2.0 and L ≳ 30 L☉). If the full-resolution MIST ages shift enough to change the reported maxima by more than ~10 percentage points at any metallicity, the headline 80–90% offsets are partly numerical rather than purely physical.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing quantitative claim is that unknown mass turns ~10% model spread into 60–90% age offsets (Abstract; §3.3; §3.4; §4.1.2). For that number to be physical, every grid must be compared at equal fidelity. §2.1 states that YREC, DSEP, and GARSTEC were resampled with extra EEPs at the late end, while MIST 'only provides the downsampled EEP tracks' and is therefore used at default resolution. §2.2.2 repeats this. No check is reported that MIST's default EEP spacing resolves the RGB bump and the rapid ascent to the tip—precisely the regions where §3.3 and Figure 3 report the maximum offsets (≈60% before the bump, ≈90% after) and where §3.4/Figure 4 report ≈80% after the bump. If kiauhoku's interpolation across a coarse EEP interval smooths or shifts the bump and tip structure in Teff–log g–luminosity space, MIST's inferred mass and age at those observables acquire a numerical error. Since the headline metric is the maximum fractional offset across only four grids, a single grid's resampling artifact near the bump can set the headline value and then propagate directly into the theoretical-uncertainty recipe of Section 5.2, which takes the multi-grid spread at face value. The qualitative conclusion—unknown mass makes model choice the dominant error—is robust, but the specific 80–90% numbers and the calibrated uncertainty prescription are not securely physical until this is tested.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript compares age estimates for red giants inferred from four common stellar evolution grids (YREC, MIST, DSEP, GARSTEC) using the kiauhoku EEP-interpolation package. For synthetic stars with known mass, metallicity, and surface gravity, the maximum fractional age offset among grids is small, with mean offsets of 9–12%; when mass is not known and age is inferred from spectroscopic observables (effective temperature, surface gravity, metallicity) or Gaia photometric observables (luminosity, effective temperature, metallicity), maximum offsets grow to roughly 60–90% and are largest after the RGB bump. The same pattern appears in an APOKASC-3 sample of about 4,300 giants: asteroseismic input gives mean grid-to-grid offsets of 5–9%, spectroscopic input gives 77–83%, and Gaia input gives 42–57%. The paper argues that model choice, not observational precision, dominates the age error budget for red giants without mass constraints and recommends adding a multi-grid theoretical uncertainty; it releases the resampled grids on Zenodo and a Jupyter notebook for reproduction.","tokens_in":23178,"tokens_out":7495,"duration_ms":69972,"significance":"If the quantitative claims are secure, the paper addresses an important practical problem: red giant ages used in Galactic archaeology are often derived without direct mass measurements, and the paper shows that the choice among standard grids can matter more than measurement error. The qualitative finding is robust across synthetic grids and the APOKASC-3 sample, and the use of open-source interpolation, published grids, and a reproducible notebook is a real strength. I see no circularity concern, since the multi-grid spread is an output of the comparison rather than a fitted target. The quantitative headline values and the proposed uncertainty recipe are not yet secure, however, because the grids are compared at unequal EEP resolution and the central statistic is a maximum across four grids; these issues do not negate the central argument but require validation or reinterpretation before the 80–90% numbers can be taken at face value.","major_comments":[{"comment":"The comparison is not at equal EEP fidelity. The text states that YREC, DSEP, and GARSTEC tracks were resampled with additional EEPs at the late end, while MIST is used only at the default downsampled EEP resolution because full-resolution tracks are not publicly available. No test is reported that MIST's default EEP spacing resolves the RGB bump and the rapid ascent to the tip, which are precisely the regions where Figures 3 and 4 place the maximum age offsets (≈60–90%) and where Section 5.2's multi-grid spread is calibrated. If kiauhoku's interpolation across a coarse EEP interval smooths the bump/tip structure in Teff–logg–L, MIST's inferred masses and ages acquire a numerical error, and because the headline statistic is the maximum across only four grids, a single grid's resampling artifact can set the headline value. I ask for a validation: compare MIST default-EEP and full-resolution or artificially EEP-densified tracks over the RGB, and show that the maximum-offset regions are not dominated by EEP-spacing differences; if full-resolution MIST is truly unavailable, the 80–90% claims and the Section 5.2 prescription should be explicitly downgraded to illustrative.","section":"§2.1, §2.2.2, Figures 3 and 4"},{"comment":"The headline statistic is the maximum fractional offset among four grids. A maximum over N=4 grids is a noisy and upward-biased estimator of model uncertainty, and the reported means are substantially lower than the maxima: mean offsets are 57–74% in Figure 4 and 77–83% in the APOKASC-3 spectroscopic case, while the maxima reach ≈90%. The paper should report the full distribution of offsets, including per-grid pairwise differences and median/percentile spreads, and the Section 5.2 uncertainty recipe should be defined with a robust statistic rather than the maximum, otherwise the proposed uncertainties inherit the fragility of a four-grid maximum.","section":"§3.3–3.4, Figures 3 and 4"},{"comment":"The APOKASC-3 spectroscopic comparison mixes two different comparisons. The grid-versus-grid offsets in Figure 7 are computed with each grid interpolating from the same spectroscopic parameters, which is appropriate, but the grid-versus-catalog offsets in Figure A.2 compare each grid to ages from the APOKASC-3 variant of GARSTEC, so the GARSTEC row is a code-variant comparison and the average values (37%, 100%, 45%, 67%) depend on this variant choice. The text acknowledges this, but the 15 Gyr exclusion also interacts with GARSTEC's hotter temperature scale, and the reported sensitivity (mean offset decreases to 66% when all GARSTEC stars with ages above 15 Gyr are removed) shows that the mean is not stable under a small selection change. Please report the spectroscopic grid-versus-grid offsets after removing the GARSTEC-variant complication, or present a sensitivity analysis over the exclusion threshold.","section":"§4.1.2, Figures 7 and A.2"},{"comment":"The paper describes a method for including theoretical uncertainties but gives no concrete prescription in the text. It states that age inference uncertainties should at a minimum account for the variation across multiple grids and points to a notebook, but it does not specify how to combine the per-grid ages into a point estimate and uncertainty (for example, mean and standard deviation, median and percentile range, or maximum spread), how to handle grids that fail to converge, or when the theoretical term should be added in quadrature to the observational term. Add a section with a precise algorithm so that the claimed deliverable is reproducible from the text alone.","section":"§5.2"}],"minor_comments":[{"comment":"The second bullet says 'reaching a mean offset of 90%', but the reported means are 57–74% in the synthetic runs and 77–83% in the APOKASC-3 spectroscopic case; 90% is a maximum, so the bullet should distinguish 'mean' from 'maximum'.","section":"Conclusions, bullet 2"},{"comment":"The caption's second metallicity interval is printed as '−0.5 < [Fe/H] < −0.'; it should read '−0.5 < [Fe/H] < −0.1'.","section":"Figure 9 caption"},{"comment":"The text says model differences reach ≈90% after the bump, while the Figure 3 caption says the mean age offset increases to 80% beyond the bump; the quantity being quoted (mean versus pointwise maximum) should be made consistent in both places.","section":"§3.3"},{"comment":"The statement that the MIST website 'only provides the downsampled EEP tracks' should include the version and date accessed, since the availability of full-resolution tracks is part of the reproducibility record and the distribution could change.","section":"§2.2.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript fits the journal's scope, and the self-citation pattern (YREC and kiauhoku come from the authors' group) is not problematic for the central comparison, since the result is an emergent property of published grids rather than a fitted target. The main risk is the MIST EEP-resolution issue; if the authors can obtain full-resolution MIST tracks or demonstrate EEP convergence, the paper is publishable after a moderate revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid, useful empirical paper. It does not discover that mass constrains giant ages — Feuillet, Tayar, and others have said that — but it quantifies the effect properly across four grids and three data regimes, and it gives the community a practical way to fold model spread into age error bars. I'd send it to a referee.\n\nThe core calculation is transparent: kiauhoku resamples tracks to EEPs, interpolates age for given observables, and reports the maximum fractional offset among grids. Synthetic experiments show roughly 10% offsets when mass is known and 60–90% when it is not, depending on location on the RGB. The APOKASC-3 application is informative: spectroscopic-only ages differ by ~80% across grids on average, and the model spread beats observational error for most of those stars. The paper also ships a notebook and resampled grids, so the result is reproducible.\n\nThe caveats are real but not fatal. The headline statistic is a maximum over only four grids, so one grid with an odd bump can set the number. The MIST grid is used at its default downsampled EEP resolution because full-resolution tracks are not public, and the regions of largest offsets — near the RGB bump and tip — are exactly where coarse EEP spacing could distort the temperature–gravity relation. The authors do not validate MIST's EEP spacing against full-resolution tracks, so part of the 80–90% could be numerical rather than physical. That said, the qualitative conclusion is robust: even if the MIST artifact were removed, the spread among YREC, DSEP, and GARSTEC would still be tens of percent, far above the ~10% mass-known case. Also, there is no ground-truth age baseline, so the numbers are a diagnostic of model dispersion, not a calibrated measurement.\n\nOne minor cleanup: the conclusion bullet says \"mean offset of 90%\" while the body reports means in the 50–70% range and a maximum near 90%. That phrasing should be fixed.\n\nWho is this for? Galactic archaeology practitioners and anyone inferring ages of field giants. A serious referee can handle it; I'd recommend publication after the MIST-resolution point is addressed, either by obtaining full-resolution tracks or by explicitly testing EEP-spacing sensitivity.","headline":"Solid, reproducible quantification of a known qualitative effect: when mass is unknown, model choice dominates red-giant age errors, and the paper's specific numbers should be treated as a range diagnostic pending a MIST resolution check.","tokens_in":23783,"tokens_out":2029,"would_cite":true,"duration_ms":20203,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"For red giants without a measured mass, the choice of stellar evolution model—not the precision of the data—can change the inferred age by up to 90 percent.","keywords":["red giant branch","stellar ages","stellar model grids","galactic archaeology","asteroseismology","theoretical uncertainties","equivalent evolutionary phases"],"falsifier":"Take the APOKASC-3 giants with asteroseismic masses and compare each grid's mass predicted from spectroscopic temperature, gravity, and metallicity against the true masses: if the real masses are consistently captured by one grid's temperature scale rather than spread across all four, the grid-to-grid age spread overestimates the true model uncertainty. Alternatively, recompute the spectroscopy-only offsets with full-resolution MIST tracks; if the roughly 90 percent offsets beyond the RGB bump shrink or shift by more than about 10 percentage points relative to the downsampled EEP tracks, the headline numbers are partly an artifact of MIST's public resolution rather than pure physics.","tokens_in":22671,"feed_emoji":"⭐","tokens_out":8902,"duration_ms":77774,"temperature":0.7,"pith_summary":"Stellar ages cannot be measured; they are inferred by matching observed stars to theoretical model grids, and different grids encode different choices about convection, opacities, abundances, and boundary conditions. This paper shows that for red giant stars the choice of grid matters enormously: when a mass measurement from asteroseismology is available, four widely used grids agree on ages to about 10 percent, but when mass is unknown and must be inferred from temperature, gravity, and luminosity, the same grids disagree by 60 to 90 percent, with the largest offsets near and above the red-giant-branch bump. The result matters because most giants in large spectroscopic and photometric surveys do not have asteroseismic masses, so their published ages carry a model-dependent uncertainty that routinely exceeds the quoted observational error. The paper also offers a practical remedy: treat the spread across grids, computed with the public kiauhoku interpolation package, as a realistic theoretical uncertainty in age inference.","feed_headline":"Red giant ages can differ by up to 90% between model grids","feed_subtitle":"When a giant's mass is unknown, picking the stellar model—not the data—drives the age error.","key_machinery":"The load-bearing tool is a comparison in a common evolutionary-phase coordinate system: kiauhoku resamples each grid's tracks to equivalent evolutionary phases (EEPs), so that the same state of evolution is compared grid-to-grid. The four grids—YREC, MIST, DSEP, and GARSTEC—differ in atmosphere, mixing-length treatment, overshoot, diffusion, opacities, and solar composition, and these differences shift the predicted effective temperature scale along the red giant branch by tens to over a hundred kelvin. The mechanism that converts those temperature shifts into large age differences is the missing mass constraint: without a measured mass, each grid independently solves for the mass that best matches the observed temperature and gravity or luminosity, and a roughly 120 K temperature offset is equivalent to about a 0.1 solar-mass change, which in turn moves inferred ages by tens of percent. The region of maximal disagreement is the RGB bump, where the luminosity of the bump itself shifts between grids because of different convective-overshoot prescriptions.","core_discovery":"The paper's central claim is that, for red giant stars, the choice of stellar evolution model grid is a first-order, usually unquantified source of age uncertainty, and that its size is set by whether the star's mass is known. Resampling the YREC, MIST, DSEP, and GARSTEC grids into a common evolutionary-phase coordinate system, the authors find that when mass, metallicity, and surface gravity are all supplied, the four grids agree on age to a mean of roughly 10 percent (9 to 12 percent depending on metallicity). When mass is removed and age is instead inferred from spectroscopic or photometric observables, grid-to-grid differences reach about 60 percent before the red-giant-branch bump and roughly 90 percent above it in the synthetic demonstrations, with mean offsets of 51 to 69 percent; applying the same procedure to the APOKASC-3 sample gives mean offsets of 77 to 83 percent using spectroscopic parameters and 42 to 57 percent using Gaia photometry. Comparison with Monte Carlo error propagation shows that model choice exceeds the observational error budget for most giants in the spectroscopy-only case and for giants cooler than about 4900 K in the Gaia case. The paper concludes by proposing that the inter-grid age spread be adopted as a theoretical uncertainty term in age inference, so that catalog ages reflect the model dependence explicitly.","pith_inferences":["The uncertainty floor the paper establishes is set by its four grids; a grid with substantially different physics (rotation, magnetic braking, alternative convective models) could widen the quoted spread, so the 90 percent figure is best read as a floor on model-choice error, not the full range of plausible stellar physics.","The largest offsets concentrate at the RGB bump, the luminosity of which the paper notes shifts between grids because of convective-overshoot treatment; this suggests a testable physics discriminator, since observed bump luminosities across metallicity could favor one overshoot prescription over another.","A natural extension is to run the same multi-grid pipeline on subgiant stars, where age correlates tightly with luminosity and model agreement is expected to be better; mapping where the offsets start rising would identify the precise evolutionary stage at which model choice begins to dominate.","The grid-by-grid mass comparison against asteroseismic masses could be turned into an empirical ranking of the grids' temperature scales; the paper refrains from endorsing any single grid, but the data products it releases make such a ranking straightforward."],"forward_implications":["For the majority of observed giants that lack asteroseismic masses, single-grid ages carry an implicit model-dependent error of order 50 to 90 percent, and galaxy-archaeology catalogs built from those ages inherit that as a systematic bias.","Multi-grid averaging with the grid-to-grid spread as the error, as the paper proposes, roughly doubles or triples the quoted age uncertainty for spectroscopic-only samples, making the uncertainty budget honest rather than optimistic.","Because the offsets grow with decreasing surface gravity and with distance from solar metallicity, studies of the most luminous giants and of metal-poor halo populations are the ones most in need of this correction.","The pattern of offsets as a function of the observable set indicates that luminosity-based coordinates (photometric or astrometric) produce smaller, though still large, model dependence than surface-gravity-based coordinates, so analyses should tailor uncertainty prescriptions to their specific observable set."],"supporting_citations":[{"why":"Supplies the kiauhoku interpolation package that resamples tracks to equivalent evolutionary phases, the tool that makes the grid-to-grid comparison possible.","marker":"Claytor et al. 2020a,b"},{"why":"Defines the equivalent evolutionary phase (EEP) scheme the resampling uses so that the same evolutionary state is compared across grids.","marker":"Dotter 2016"},{"why":"Builds the YREC grid, the fitting methodology with a mean-squared-error loss, and the table of input physics replicated in this paper.","marker":"Tayar et al. 2022"},{"why":"Supplies the MIST grid, one of the four model sets, including its downsampled public EEP tracks used at default resolution.","marker":"Choi et al. 2016"},{"why":"Supplies the DSEP grid, one of the four model sets.","marker":"Dotter et al. 2008"},{"why":"Supplies the GARSTEC grid, one of the four model sets.","marker":"Serenelli et al. 2013"},{"why":"Provides the APOKASC-3 catalog used as the real-data test case and the quoted typical observational uncertainties for mass, metallicity, and surface gravity.","marker":"Pinsonneault et al. 2025"},{"why":"Provides the earlier result that photometric mass estimates lead to roughly 41 percent age uncertainties, the comparison point for the spectroscopy-only case.","marker":"Feuillet et al. 2016"}],"fun_headline_variants":["Stellar model choice swings red giant ages by 80%+","Red giant age errors up to 90% when mass is unknown","Model choice, not data, drives red giant age errors","Red giant ages vary up to 90% with model grid","Stellar models disagree on red giant ages by 80-90%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes that the four grids, resampled onto a common evolutionary-phase grid, truly span the range of physically plausible model predictions at every point of the red giant branch—including the region near the RGB bump where MIST is only available at downsampled resolution, so the claimed offsets could partly be numerical artifacts of the resampling rather than physical model differences.","fun_headline_variants_meta":{"raw":{"variants":["Stellar model choice swings red giant ages by 80%+","Red giant age errors up to 90% when mass is unknown","Model choice, not data, drives red giant age errors","Red giant ages vary up to 90% with model grid","Stellar models disagree on red giant ages by 80-90%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000788,"raw_usage":{"total_tokens":3491,"prompt_tokens":976,"completion_tokens":2515,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":592,"completion_tokens_details":{"reasoning_tokens":2426}},"tokens_in":592,"tokens_out":2515,"duration_ms":16492,"temperature":1.0,"reasoning_tokens":2426,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:35:26.759233+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the APOKASC-3 giants with asteroseismic masses and compare each grid's mass predicted from spectroscopic temperature, gravity, and metallicity against the true masses: if the real masses are consistently captured by one grid's temperature scale rather than spread across all four, the grid-to-grid age spread overestimates the true model uncertainty. Alternatively, recompute the spectroscopy-only offsets with full-resolution MIST tracks; if the roughly 90 percent offsets beyond the RGB bump shrink or shift by more than about 10 percentage points relative to the downsampled EEP tracks, the headline numbers are partly an artifact of MIST's public resolution rather than pure physics.","supporting_citations":[],"review_version":1}