{"id":"fb0bf620-10e7-4948-b312-d4962a585fbc","arxiv_id":"2607.22309","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A new catalog of ~1.5 million stellar ages built from the median vertical action of chemically similar stars, calibrated on subgiant ages and cross-validated against clusters, asteroseismology, and gyrochronology.","lead":"Astronomers assigned ages to 1.5 million stars by averaging the vertical motions of chemically similar groups, rather than using individual star motions. The ages agree with independent methods to about 2–3 billion years, offering a new statistical tool for Milky Way history.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Universal lnJz-age relation calibrated on subgiants is assumed for all populations; the paper's own Figures 5–6 show breakdowns for old and metal-poor stars, so catalog ages in those regimes may be systematically biased.","rationale":"The reader identified the universality of the subgiant-calibrated lnJz–age relation as the weakest assumption; I agree. The paper's own figures demonstrate that the relation does not hold in the old high-α and metal-poor low-α regimes, and Section 3.1 explicitly disclaims lower-main-sequence ages. Since the catalog includes these populations and the abstract claims full-HR applicability, this is a load-bearing gap rather than a minor caveat. The circular subgiant validation is a separate issue, but it is secondary because Figure 8 provides independent comparisons (APOKASC-3, open clusters, gyro, wide binaries) that broadly support ~2 Gyr scatter. Nevertheless, those independent checks are dominated by near-solar metallicity, intermediate-age stars and do not test the problematic regimes. A population-split calibration/validation would directly test universality and could be done with existing data. The paper is otherwise honest and well-structured, and the method is useful if restricted to the calibrated regime, so the appropriate verdict remains CONDITIONAL: the central claims should be revised or the universal relation empirically demonstrated.","tokens_in":20573,"tokens_out":5208,"duration_ms":45262,"concrete_test":"Hold out the high-α subgiant population: calibrate Equation (1) using only low-α subgiants from Xiang & Rix (2022) (age<10 Gyr, as in their fit), then predict ensemble kinematic ages for the high-α subgiants (age>9 Gyr) and compare to their subgiant ages. If the median residual exceeds ~2 Gyr or shows a systematic trend with age or [Fe/H] beyond the low-α internal scatter, the universal relation is rejected. As a complement, repeat the calibration excluding the 7–8 Gyr, [Fe/H]≈−0.75 bin identified as region '2' in Figure 6 and check whether predicted ages for that bin recover the subgiant ages.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that ensemble kinematic ages are accurate to ~2 Gyr (~30%) and extend age estimates across the full HR diagram. The method hinges on Equation (1), a single lnJz–age–[Fe/H] polynomial fitted to subgiant ages from Xiang & Rix (2022), and then applied to 1.5M LAMOST stars spanning all evolutionary stages and disk populations. For this to hold, the subgiant-calibrated relation must be universal across the high-α disk, metal-poor low-α stars, and lower main-sequence stars. The paper itself falsifies universality in places: Figure 5 shows a plateau and growing bias for subgiant ages >12.5 Gyr, and Figure 6 marks the ~7.5 Gyr, [Fe/H]~−0.75 low-α region (dark gray '2') where ages are not reproduced. Section 3.1 also states lower main-sequence stars cannot be reliably aged. These are not edge cases: old high-α stars and metal-poor populations are precisely the targets of Galactic archaeology. If Eq. (1) does not transfer across these populations, the catalog's ages for those stars carry undetermined systematics, and the 'robust population-level tool' claim is unsupported for a large fraction of the sample. The validation against subgiant ages is also partly circular because the same sample is used for calibration and validation, although independent checks (APOKASC-3, clusters) provide some support.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes 'ensemble kinematic ages': for each target star, the authors isolate a local ensemble of ~20–50 stars with similar atmospheric and chemical parameters, compute the median vertical action ln J_z of that ensemble, and convert it to an age using Eq. (1), a fifth-order polynomial in ln J_z with a linear [Fe/H] scaling. The relation is calibrated on the subgiant ages of Xiang & Rix (2022). The method is validated against that same subgiant sample (MED = 0.9 Gyr, MAD = 2.6 Gyr, ~30%), then applied to ~1.5 million LAMOST DR5 stars, with comparisons to APOKASC-3, [C/N] ages, open clusters, wide binaries, gyrochronology, and isochrone ages. The paper also constructs empirical isochrones and discusses old low-α and young high-α populations.","tokens_in":20959,"tokens_out":6718,"duration_ms":65396,"significance":"If the central claim holds, the resulting catalog would be a valuable population-level age indicator that can be applied across much of the HR diagram and used as a cross-calibration tool between different age-dating techniques. The paper's strengths include the public release of the catalog (Table 1), the explicit enumeration of assumptions in Sec. 2.2, and the candid discussion of failure regions in Secs. 3.1 and 5. The key weaknesses are that the headline accuracy is measured on the calibration sample itself, and that the universal ln J_z–age–[Fe/H] relation is applied to populations for which the paper's own figures show systematic breakdowns. The independent checks provide partial support but are not broad enough to establish a uniform ~2 Gyr accuracy across the full HR diagram.","major_comments":[{"comment":"The headline '~30% accuracy' is not an independent validation. Equation (1) is fitted to the running median of ln J_z against subgiant age using the same Xiang & Rix (2022) sample that is then compared in Fig. 5. The quoted MED = 0.9 Gyr and MAD = 2.6 Gyr therefore largely measure the quality of the calibration fit, not predictive accuracy. The paper itself acknowledges this in Sec. 4.1 ('our ensemble kinematic ages show the smallest bias relative to the subgiant ages, since the lnJz–age relation is calibrated on this sample'). Please add a held-out test (e.g., fit on half of the subgiants and validate on the other half) or calibrate on an independent age scale; otherwise the abstract's 'accuracy of ~30%' should be reframed as scatter relative to the calibration scale.","section":"Sec. 2.4, Sec. 3, Fig. 5"},{"comment":"The universality of Eq. (1) is load-bearing but not established. Fig. 2 shows different slopes and offsets for the high-α and low-α disks; Fig. 6 identifies ages >12.5 Gyr and the ~7.5 Gyr, [Fe/H] ~ -0.75 low-α region as places where the relation fails; Sec. 3.1 states that lower main-sequence stars cannot be robustly aged. Nevertheless, Eq. (1) is applied to all 1.5 million LAMOST stars, including those in these breakdown regions, and the abstract claims ages 'across the full HR diagram.' The catalog should either restrict the applicable parameter ranges, provide per-population validity flags, or calibrate and validate separate relations for high-α, metal-poor, and lower-main-sequence populations. The old high-α and metal-poor stars are precisely the populations of greatest interest to Galactic archaeology, so this is not a minor edge case.","section":"Sec. 2.4, Sec. 4, Figs. 2 and 6"},{"comment":"The independent validations are too partial to support a uniform '~2 Gyr' uncertainty. The APOKASC-3 comparison shows a systematic ~2 Gyr overprediction, which the paper attributes to age-scale offsets but does not resolve. Open clusters are concentrated near solar metallicity, as the paper itself notes in Sec. 5, so they cannot test the metal-poor regime. Wide binaries are predominantly main-sequence stars, where stellar parameters provide weak age discrimination. The lack of a strong log g trend in Fig. 12 is suggestive but does not demonstrate that dwarf and giant ages sit on the same absolute scale across metallicity and age. Please report residuals separately for each evolutionary state and metallicity regime, and quantify how age-scale offsets and selection effects affect the claimed accuracy.","section":"Sec. 4.1, Figs. 8 and 12"},{"comment":"Applying the method to stars without full 6-D kinematics assumes that the subset of LAMOST stars with Gaia DR3 radial velocities is representative of each parameter bin. The paper states only qualitatively that selection effects matter. If stars with and without radial velocities differ in distance, brightness, or Galactic location, the median ln J_z of the kinematic subset can be biased relative to the full target population. Please quantify the fraction of LAMOST stars lacking radial velocities, test for kinematic differences within bins, or provide selection weights. This directly affects the catalog ages for all stars without measured radial velocities.","section":"Sec. 4 and Table 1"}],"minor_comments":[{"comment":"Please state the units of J_z (kpc km s^-1) and the valid range of ln J_z, age, and [Fe/H] over which the polynomial was calibrated. The text later excludes stars with median ln J_z < 0.5, but the calibration range is not specified.","section":"Eq. (1)"},{"comment":"The definition of MAD as 'Median(|ensemble kinematic ages − Median(subgiant ages)|)' is not a paired residual statistic. The usual MAD for a comparison would use paired differences (ensemble age − subgiant age). As written, it measures scatter of ensemble ages around the median subgiant age rather than the accuracy of individual age recovery. Please correct the definition or state clearly what statistic is shown.","section":"Sec. 3, Fig. 5"},{"comment":"The caption reads 'Similar to Figure 4 but over-plotting the best-fit model,' which appears to be a self-reference; it should refer to Figure 1 or Figure 3.","section":"Fig. 4 caption"},{"comment":"Notation is inconsistent: the text uses 'ln J_z' and 'J_z' interchangeably, and Figure 1 labels the axis 'J_z' while the text and Eq. (1) use 'ln J_z'. Please standardize.","section":"Throughout"},{"comment":"Minor language issues: 'an age-J_z relations' should be 'an age–J_z relation'; 'the low-α disk' is sometimes written 'low-α disk' without the article; in Sec. 4 the phrase 'If these stars are not the result of systematics... they may indicate either parallel disk formation...' is incomplete ('either' without a second option).","section":"Sec. 1 and Sec. 2.3"}],"recommendation":"major_revision","confidential_remarks":"The core idea is promising and the catalog could be a useful community resource, but the central accuracy claim is currently in-sample and the transfer to all HR-diagram populations is overreaching relative to the paper's own caveats. I do not recommend rejection; the needed changes are substantial but feasible: held-out calibration, explicit applicability flags, and per-population validation. The authors' Secs. 3.1 and 5 are more cautious than the abstract, and aligning the abstract with those sections would improve the paper substantially."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuine extension of ensemble kinematic ages from rotation-selected dwarf groups (Lu et al. 2021) to abundance-selected groups across the HR diagram, and it ships a public catalog. The method is plausible and the paper is honest about its own limitations. But the headline validation against subgiant ages is partly circular: Eq. 1 is fit to the moving median of those same subgiant ages, so the ~30% agreement in Section 3 is not an independent check. The authors acknowledge this in Section 4.1, and they do provide independent checks (APOKASC-3, open clusters, wide binaries, [C/N], gyro) that broadly support the ~2 Gyr scatter. It is also true, as the stress-test says, that a universal lnJz-age-[Fe/H] relation calibrated on subgiants is assumed to hold for all populations. The paper itself shows breakdowns at ages >12.5 Gyr and in the ~7.5 Gyr metal-poor low-alpha region, and it explicitly says lower main-sequence stars cannot be aged reliably. Those are real soft spots, and they are more than edge cases for Galactic archaeology. Still, the independent validations and the fact that the authors map where the relation fails mean the central tool is not empty; it's a solid population-level product with documented domain limits.\n\nWhat's actually new: applying ensemble kinematic ages to a 1.5M-star LAMOST sample using 13 stellar parameters, with a public catalog and reported median vertical action so readers can calibrate their own relation. The empirical isochrone demonstration is a nice bonus, not the main event. The use of detailed abundances rather than just Teff and rotation is a meaningful generalization.\n\nSoft spots, in proportion: (1) circularity in the subgiant validation is real but partial; independent checks carry weight. (2) The universal-relation assumption is load-bearing and only partially tested; the paper identifies failure regions but does not quantify how many catalog stars fall in them. (3) No per-star uncertainties in the catalog; for a statistical method, that's an omission. (4) The abstract says 'extending age estimates across the full HR diagram,' which overstates given the lower main sequence caveat. These are fixable with a revised abstract, an uncertainty column or at least per-star group size/action distribution, and a more explicit statement that calibration and validation samples overlap.\n\nWho this is for: anyone doing Galactic archaeology or stellar population work with LAMOST; it's a useful cross-calibration tool. I'd cite the catalog if I worked in that area.\n\nRecommendation: send it to peer review. A good referee should push on the calibration/validation separation and ask for uncertainty handling, but this deserves serious engagement.","headline":"A useful, well-calibrated statistical age catalog for 1.5M LAMOST stars, with a real circularity problem in the headline validation and a clear need for per-star uncertainties, but more than worth refereeing.","tokens_in":21467,"tokens_out":2275,"would_cite":true,"duration_ms":19854,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that the median vertical motion of stars with similar temperature, surface gravity, and chemistry can be turned into stellar ages with ~30% accuracy, producing a self-consistent age catalog for 1.5 million stars in the Mil","keywords":["stellar ages","kinematic ages","vertical action","Milky Way disk","age-velocity relation","galactic archaeology","ensemble age inference","stellar populations"],"falsifier":"Compare the inferred ensemble kinematic ages against precise asteroseismic ages for a sample of old, metal-poor red giants with [Fe/H] below about –0.5. If the ages disagree by more than the claimed ~2 Gyr scatter in a way that grows with age or with metallicity, the calibrated relation is not universal. The paper itself flags ages above ~12.5 Gyr and the metal-poor low-alpha region around 7.5 Gyr as regions where the relation already breaks down.","tokens_in":20434,"feed_emoji":"⭐","tokens_out":4785,"duration_ms":40445,"temperature":0.7,"pith_summary":"This paper tries to establish that stellar kinematics—averaged over carefully chosen groups of similar stars—are a usable stellar age indicator across nearly the whole Hertzsprung-Russell diagram. Each star is assigned the median vertical action of 20–50 stars with matching temperature, surface gravity, metallicity, and detailed element abundances, and that median action is converted to an age using a metallicity-dependent relation calibrated on precise subgiant ages. If correct, ages become available for millions of stars, including giants and dwarfs where traditional methods fail, and different dating techniques can be cross-checked against one another on a common scale. The paper validates the method against subgiant, asteroseismic, cluster, gyrochronology, and wide-binary ages, finding a typical scatter of about 2 Gyr (roughly 30%).","feed_headline":"Star motions yield ages for 1.5 million stars to ~30%","feed_subtitle":"Averaging vertical motions of similar stars puts dwarfs and giants on one age scale, checked against five independent methods.","key_machinery":"The central object is the vertical action J_z, the phase-space area enclosed by a star's vertical oscillation about the Milky Way's midplane. Because J_z is an adiabatic invariant, it accumulates the vertical heating history of the disk and can be computed per star. The method groups stars that are observationally indistinguishable in stellar-parameter space (effective temperature, absolute magnitude or surface gravity, [Fe/H], [alpha/Fe], plus several elemental abundances), takes the median ln J_z of each group of 20–50 stars, and converts that median to an age through a fifth-order polynomial in ln J_z that is scaled linearly by [Fe/H]. The ensemble construction is what converts noisy indi","core_discovery":"The central claim is that a single, continuous metallicity-dependent relation between vertical action and age holds across both the high- and low-alpha disk populations, and that this relation can be calibrated using isochrone ages of subgiants. Using this relation, ages computed as the median action of parameter-matched ensembles reproduce the calibration ages with a median absolute deviation of 2.6 Gyr (about 30%), comparable to carbon-to-nitrogen-based ages, and agree with independent methods across different evolutionary stages. Applied to roughly 1.5 million stars, the method yields an internally consistent age catalog spanning dwarfs and giants, with the caveats that ages above roughly","pith_inferences":["If the subgiant-calibrated relation holds in the metal-poor regime, this method could map age structure in the oldest disk and inner halo, where current indicators are weakest.","The requirement that parameter-space bins be approximately coeval implies that combining ensemble kinematics with rotation periods could tighten gyrochronology beyond about 4 Gyr, a range where rotation-age relations are poorly calibrated.","Because the method averages over bins, real age spreads within a bin are smeared; the width of the action distribution inside each bin could itself be used as a diagnostic of the remaining age spread.","With the next astrometric data release roughly doubling the number of stars with full 6-D motions, the same calibration should yield larger samples and finer bins, improving the precision floor."],"forward_implications":["If the method works as claimed, ages become available for roughly 1.5 million stars, including many giants and dwarfs lacking asteroseismic or rotation data.","Because the same relation assigns ages across evolutionary stages, it provides a common age scale for comparing dwarf and giant populations, with no strong dependence on surface gravity.","The approach can serve as a cross-calibration tool: offsets between ensemble kinematic ages and asteroseismic or [C/N]-based ages reveal systematic differences between age scales.","Empirical isochrones built from these ages can be compared directly with theoretical tracks, offering a data-driven check of stellar evolution models.","The catalog reveals bimodal age distributions in the [Mg/Fe]–[Fe/H] plane, including an old population at low [Mg/Fe] that, if real, bears on how the two disks formed."],"fun_headline_variants":["Ensemble motions put 1.5M stars on one age scale","1.5M LAMOST stars get ages from ensemble kinematics","Team kinematics give 1.5M star ages accurate to ~30%","Motion-averaging yields ages for 1.5M stars, dwarfs and giants","New age catalog: 1.5M stars via ensemble vertical motions"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that a single age–vertical-action relation, calibrated on subgiants that mostly belong to the intermediate-age disk, applies equally well to every other stellar population and evolutionary stage—especially old metal-poor high-alpha stars and lower main-sequence stars; if the relation is not universal, the derived ages for those groups shift systematically.","fun_headline_variants_meta":{"raw":{"variants":["Ensemble motions put 1.5M stars on one age scale","1.5M LAMOST stars get ages from ensemble kinematics","Team kinematics give 1.5M star ages accurate to ~30%","Motion-averaging yields ages for 1.5M stars, dwarfs and giants","New age catalog: 1.5M stars via ensemble vertical motions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00068,"raw_usage":{"total_tokens":2950,"prompt_tokens":795,"completion_tokens":2155,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":539,"completion_tokens_details":{"reasoning_tokens":2055}},"tokens_in":539,"tokens_out":2155,"duration_ms":12238,"temperature":1.0,"reasoning_tokens":2055,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T05:07:58.618970+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the inferred ensemble kinematic ages against precise asteroseismic ages for a sample of old, metal-poor red giants with [Fe/H] below about –0.5. If the ages disagree by more than the claimed ~2 Gyr scatter in a way that grows with age or with metallicity, the calibrated relation is not universal. The paper itself flags ages above ~12.5 Gyr and the metal-poor low-alpha region around 7.5 Gyr as regions where the relation already breaks down.","supporting_citations":[],"review_version":1}