{"id":"9988a876-5197-4cbc-b85e-0cee2fd1b23b","arxiv_id":"1908.04548","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A new inversion method recovers the underlying log(age) distribution of large stellar ensembles from stacked UniDAM age PDFs, tested on simulations and applied to APOGEE, GALAH, RAVE-on and LAMOST.","lead":"The authors build a method to recover the true age distribution of a large sample of stars from noisy, ambiguous per-star age estimates made by isochrone fitting. The method deconvolves the stacked age probability distribution using simulated mono-age stellar populations, and it recovers known age distributions in tests and plausible age-metallicity trends in real survey data.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Recovered age distributions hinge on an untested MAMC kernel; self-consistent simulations and acknowledged model systematics leave real-data inversions vulnerable to bias.","rationale":"The reader's verdict (CONDITIONAL) matches my assessment. The inversion method is a coherent and useful contribution: the forward problem is clearly formulated, the regularization and cross-validation are sensible, and the simulated tests demonstrate that the method can recover smooth input distributions for sufficiently large N. However, the central claim 'capable of properly recovering the true log(age) distribution' is conditional on the MAMC kernel being a faithful representation of the real survey smearing. The paper's own wording supports this conditionality: Section 5.3 explicitly notes that systematics are zero by construction in the tests, and Section 6.1 lists PARSEC-model offsets as a possible cause of the observed differences. The within-bin IMF weighting is an additional, more specific assumption that is not tested at all. A direct external validation with asteroseismic ages would either confirm the method or quantify the bias. Because the concern is real but addressable, and the paper's mathematical core is not invalidated, the correct verdict remains CONDITIONAL rather than ACCEPT or REJECT. My read does not change the reader's verdict; the test proposed would be the deciding check.","tokens_in":17207,"tokens_out":7074,"duration_ms":64545,"concrete_test":"Apply the inversion to more than 10^3 APOGEE red giants with independent asteroseismic ages (e.g., APOKASC-2). Construct APOGEE-matched MAMCs exactly as in Section 4, run UniDAM on the observed parameters, and invert the stacked PDF. Compare the recovered N(τ) with the asteroseismic age distribution using a two-sample test over the same age range; the comparison tests the full kernel, including the within-bin IMF-weighted sampling. If the recovered distribution lies outside the 68% uncertainty band, the MAMC kernel is biased for real data and the central claim fails; agreement would validate the method externally.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the mono-age mock catalogue (MAMC) kernel P(τ,τ~) accurately describes how a real star of true age τ~ maps into UniDAM log(age) PDFs. This is the load-bearing element of Equation (4), but it is not validated on real data. Section 5 tests are self-consistent: simulated surveys are built from the same PARSEC/UniDAM pipeline as the MAMCs, so any systematic error in the isochrones, the adopted uncertainty distributions, or the mock construction cancels by construction. The paper itself acknowledges this in Section 5.3 ('systematic uncertainties are zero by definition...') and Section 6.1 ('PARSEC models and spectroscopic measurements can have a systematic offset...'), but does not quantify the bias. A further unvalidated choice is made in Section 4: within each (logg, [Fe/H]) bin, models are drawn with IMF weights rather than matching the survey's actual Teff distribution. Magnitude-limited surveys can have a stellar mix within a (logg, [Fe/H]) bin that differs from the IMF-weighted model mix, so even a perfect stellar model would yield a biased kernel. If P is biased, no inversion can recover the true N(τ); the bimodal age distributions and age-metallicity trends in Section 6 could be artefacts. This is therefore the single most load-bearing concern.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a method for recovering the underlying log(age) distribution N(τ) of a large stellar sample from the stacked log(age) probability density functions produced by the UniDAM isochrone-fitting pipeline. The central idea is to model the observed stacked PDF C(τ) as a linear superposition of mono-age mock catalogue (MAMC) PDFs P(τ,τ~), leading to the integral equation C(τ)=∫P(τ,τ~)N(τ~)dτ~, discretized and solved with non-negative least squares plus Tikhonov regularization. The regularization parameter is chosen by a cross-validation procedure. The method is tested on simulated surveys built from the same PARSEC/UniDAM pipeline for six input age distributions and varying sample and mock sizes. It is then applied to APOGEE, GALAH, RAVE-on, and LAMOST, including mono-metallicity slices, and the resulting age-metallicity relations are compared with Feuillet et al. (2018) and with chemical evolution models. The paper concludes that such an inversion is necessary to recover underlying age distributions for samples with N>10^3 stars.","tokens_in":17496,"tokens_out":8611,"duration_ms":90392,"significance":"If validated, the method would be a valuable tool for Galactic archaeology, because it directly addresses the well-known bias and multimodality of per-star isochrone ages and provides a way to deconvolve population-level age distributions from large spectroscopic surveys. The mathematical formulation is clear, the non-negativity and regularization choices are standard, and the simulations with known input distributions demonstrate that the method works in an idealized setting. The paper is also honest about several limitations, including the self-consistency of the mock tests and the presence of possible systematic offsets between models and real data. The independent comparison to Feuillet et al. (2018) is a useful external benchmark, although it tests only the mean age-metallicity relation rather than the full recovered distribution. The main weakness is that the load-bearing MAMC kernel is never validated against real data or against an alternative stellar model, so the real-data inversions could be biased.","major_comments":[{"comment":"The kernel P(τ,τ~) is the load-bearing element of Eq. (4), but its construction is not guaranteed to match the real survey. The paper states that within each (logg,[Fe/H]) bin, models are selected with probabilities proportional to the initial-mass-function weight, rather than matching the actual Teff distribution of the survey in that bin. Main-sequence and giant stars contribute very differently to the log(age) PDF, as shown in Fig. 3, so a mismatch in the internal mix within a bin will directly bias P and, through Eq. (4), the recovered N(τ). The authors should either demonstrate that the IMF-weighted sampling reproduces the survey's Teff distribution in each bin or quantify the bias introduced by this approximation.","section":"Section 4, MAMC construction and Eq. (4)"},{"comment":"The validation in Section 5 is self-consistent: both the simulated 'observed' survey and the MAMCs are produced from the same PARSEC isochrones, the same uncertainty-assignment procedure, and the same UniDAM pipeline, so any systematic error in the stellar models or in the assumed scatter cancels by construction. The paper acknowledges this in Section 5.3 ('systematic uncertainties are zero by definition') and Section 6.1, but it never quantifies how a mismatch between the true kernel and the assumed MAMC kernel propagates into the recovered N(τ). Given that the central claim is that the method recovers the underlying age distribution of real surveys, the authors should add controlled experiments that perturb the isochrone set, the Teff/logg/[Fe/H] zero-points, or the uncertainty scale, and report the induced changes in the recovered distributions.","section":"Section 5, tests with mock data"},{"comment":"The derivation of Eq. (4) assumes that the smearing kernel is the same for all stars of a given true age, independent of their other properties. For the full-survey inversion this requires that age does not depend on logg or [Fe/H] within the sample, an assumption the authors themselves question in Section 6.1. For the mono-metallicity slices in Section 6.2, the same issue persists within each metallicity bin because a magnitude-limited survey selects luminous low-logg stars at larger distances, so the age distribution may vary with logg even at fixed [Fe/H]. The bimodal age distributions and the high-metallicity upturn at [Fe/H]≥+0.2 are left with competing explanations, two of which are artefacts of this assumption. This needs to be modeled or otherwise bounded before the real-data inversions can be interpreted as underlying age distributions.","section":"Sections 6.1 and 6.2, assumption of age independence"},{"comment":"Equations (19) and (20) do not define conditional mean ages or mean metallicities as the text claims. For a two-dimensional density C(τ,[Fe/H]), the mean log(age) in a metallicity bin is ∫ τ C(τ,[Fe/H]) dτ / ∫ C(τ,[Fe/H]) dτ, not ∫ C(τ,[Fe/H]) dτ / (τmax−τmin). The same problem affects Eq. (20). If the blue and red curves in Figs. 13–15 were computed with these formulas, they do not measure what the captions claim, and the comparison with Feuillet et al. (2018) and the chemical evolution models is not quantitative. The formulas should be corrected, or the normalization and definition of C(τ,[Fe/H]) should be clarified.","section":"Section 6.3, Eqs. (19) and (20)"}],"minor_comments":[{"comment":"The Toeplitz matrix in Eq. (9) is difficult to read as typeset; the dots and spacing appear garbled. Please typeset it cleanly, for example using \\ddots and aligned columns.","section":"Eq. (9)"},{"comment":"The sentence 'There are no clear trends visible for F(ω) values with nsim and nmock' contains a subject-verb agreement error; it should read 'There are no clear trends visible in F(ω) as a function of nsim and nmock.'","section":"Section 5.3"},{"comment":"The caption of Figure 2 lists line styles, but in the printed version it is not always obvious which line corresponds to the histogram of mean, median, or mode values. A legend or explicitly labeled subpanels would improve readability.","section":"Figure 2"},{"comment":"The conclusion that the inversion is 'necessary' to recover underlying age distributions is stronger than what the tests demonstrate. The simulations show that the method works for the considered input distributions and survey sizes, but they do not establish necessity in general. Suggest softening this wording.","section":"Abstract and Section 7"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for Astronomy & Astrophysics and represents a useful contribution if the kernel-validation gap and the Eq. (19)/(20) issue are addressed. I do not see a circularity problem with the external Feuillet et al. comparison; the main concern is that the real-data inversions rest on an unvalidated MAMC kernel. The method is promising, but the central claim as currently stated is stronger than the evidence supports."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core idea here is solid and useful: instead of stacking mean or mode ages from isochrone-fitting PDFs, represent the stacked PDF as a weighted sum of mono-age mock catalogues and invert with regularized NNLS. That is a real step forward for Galactic archaeology, where age PDFs are often multi-modal and mean/mode proxies are biased. The math is clearly laid out, the cross-validation scheme for choosing lambda is sensible, and the simulations demonstrate good recovery for N > 10^3. I also credit the comparison to Feuillet et al. (2018) and to chemical evolution models; that is independent external benchmarking, not just self-validation.\n\nThe soft spots are real but mostly proportional. The biggest one is the load-bearing kernel P(τ,τ~): it is built from the same PARSEC/UniDAM pipeline that produces the \"observed\" PDFs, so the simulation tests are self-consistent by construction. The authors explicitly acknowledge this in Sections 5.3 and 6.1, but they do not quantify the bias that a real systematic offset between models and data would introduce. There is also a subtler unvalidated choice in Section 4: within each (logg,[Fe/H]) bin, models are drawn with IMF weights rather than matching the survey's actual Teff distribution. For a magnitude-limited survey, the stellar mix inside a bin can differ from the IMF-weighted mix, which biases the kernel even if the stellar models are perfect. That concern lands on reading the paper; it is not a manufactured flaw.\n\nThe full-survey bimodal age distributions and the upturn at high metallicity are interesting but could be artefacts of these kernel biases. The abstract's \"necessary\" is too strong for the external evidence; the paper shows the inversion is useful, not that it is strictly required. Minor issues: no released code, and the uncertainty analysis shows the reported 3-sigma intervals behave like 2-sigma, which is worth stating more prominently.\n\nWho is this for? Anyone working on ages from large spectroscopic surveys, especially APOGEE/RAVE/LAMOST/GALAH. The method is worth engaging with seriously, and the authors have made a genuine methodological contribution. It needs revision rather than rejection: quantify the kernel sensitivity (e.g., vary the IMF weighting, add a simple systematics model), release the code, and soften the conclusion to match the evidence. A serious referee should get this.\n\nRecommendation: send to peer review; expect heavy revision but the core method will survive.","headline":"A genuinely useful inversion method for recovering ensemble age distributions from stacked isochrone-fit PDFs, with self-consistent validation only; the real-data results are suggestive but not yet robust to unquantified model systematics.","tokens_in":18063,"tokens_out":937,"would_cite":true,"duration_ms":10904,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The smeared age PDF of a large stellar sample can be inverted to recover the true underlying age distribution, provided the smearing kernel is simulated from mono-age mock catalogues.","keywords":["stellar ages","age inversion","Bayesian isochrone fitting","spectroscopic surveys","age-metallicity relation","non-negative least squares","Tikhonov regularization","Galactic archaeology"],"falsifier":"Take a sample of stars from clusters with independently known ages, embed it in a survey-like set, and run the inversion: if the recovered age distribution peaks away from the known cluster ages by more than the quoted confidence intervals, the smearing kernel built from mock catalogues is not the kernel acting on real stars.","tokens_in":16953,"feed_emoji":"🌌","tokens_out":14090,"duration_ms":129970,"temperature":0.7,"pith_summary":"Individual stars often cannot be assigned a single trustworthy age: isochrone fitting produces broad, multi-peaked $\\log(\\mathrm{age})$ probability distributions, so histograms of mean, median, or modal ages are biased. This paper claims that for any large sample (more than about $10^3$ stars) the underlying $\\log(\\mathrm{age})$ distribution can nonetheless be recovered by solving $C(\\tau)=\\int P(\\tau,\\tilde{\\tau})N(\\tilde{\\tau})\\,d\\tilde{\\tau}$, where $C$ is the stacked observed age PDF and $P$ is the age PDF of simulated mono-age mock catalogues matched to the survey's $\\log g$–[Fe/H] distribution. The solution is a weighted mixture of mono-age populations, found with non-negative least squares plus a smoothness penalty, and tests on simulated data show it recovers the true input distribution where the stacked PDF does not. Applied to RAVE-on, LAMOST, APOGEE, and GALAH, the method gives bimodal age distributions and age–metallicity trends consistent with a high-precision local APOGEE sample and with chemical evolution models. If correct, this makes inverting the smearing kernel a necessary step for age-based studies of large spectroscopic surveys.","feed_headline":"Age inversion removes the bias hidden in stacked stellar age PDFs","feed_subtitle":"For large spectroscopic samples, undoing the age-smearing equation recovers the true age mix hidden in stacked PDFs.","key_machinery":"The load-bearing object is the mono-age mock catalogue (MAMC): a simulated set of stars all sharing one $\\log(\\mathrm{age})$ $\\tilde{\\tau}$, built from the PARSEC stellar evolution models, with each model star replaced by a small cluster of perturbed copies carrying uncertainties sampled from the base survey, and binned in $\\log g$–[Fe/H] to match the survey footprint. Each MAMC is run through the same Bayesian isochrone-fitting pipeline as the real data, yielding the conditional PDF $P(\\tau,\\tilde{\\tau})$ — the smearing kernel. The method turns the integral equation $C(\\tau)=\\int P(\\tau,\\tilde{\\tau})N(\\tilde{\\tau})\\,d\\tilde{\\tau}$ into a discretized linear system and solves it with non-negative least squares plus Tikhonov regularization on the first derivative of $N$, the regularization parameter being selected from a piecewise-linear fit to cross-validation residuals. The key identity is that an observed stacked PDF is a mixture of mono-age PDFs, so the mixture weights are precisely the underlying age distribution.","core_discovery":"The central claim is that the smeared, stacked $\\log(\\mathrm{age})$ probability density function of a large stellar ensemble can be inverted to recover the underlying $\\log(\\mathrm{age})$ distribution $N(\\tau)$, as long as the smearing kernel $P(\\tau,\\tilde{\\tau})$ is known. The kernel is built from mono-age mock catalogues: simulated populations of a single age $\\tilde{\\tau}$, constructed with the same stellar models, the same $\\log g$–[Fe/H] distribution, and survey-matched uncertainties as the observed sample, then run through the same Bayesian isochrone-fitting pipeline that produced the observed PDFs. Discretizing $C(\\tau)=\\int P(\\tau,\\tilde{\\tau})N(\\tilde{\\tau})\\,d\\tilde{\\tau}$ gives the linear system $C_j=\\sum_i P_{j,i}N_i$, solved by non-negative least squares with a Tikhonov smoothness penalty whose strength is set by cross-validation across five survey and mock realizations. On simulated data with known input distributions, the inversion recovers the input age distribution for samples above about $10^3$ stars, whereas the stacked PDF and mean/mode histograms do not. On real surveys, the inversion removes artefacts such as low-age spikes and yields age–metallicity relations close to those from a high-precision local sample and from chemical evolution models. The paper concludes that such an inversion is necessary, not optional, for large ($N>10^3$) stellar samples.","pith_inferences":["The same mixture-deconvolution scheme could be applied to any posterior PDF from Bayesian parameter estimation, such as mass, distance, or initial mass estimates, whenever a smearing kernel can be simulated from mono-value mocks.","Because the paper finds fractional uncertainty scales roughly as $n_{\\mathrm{sim}}^{-0.4} n_{\\mathrm{mock}}^{-0.15}$, mock-catalogue sizes beyond a few thousand give diminishing returns; for future million-star surveys, the practical bottleneck is generating enough mocks, so cheaper surrogate kernels for $P(\\tau,\\tilde{\\tau})$ deserve attention.","An independent cross-check would be to repeat the inversions with a second, independent set of stellar models: wherever the recovered $N(\\tau)$ shifts substantially, model systematics rather than survey noise dominate the derived age–metallicity trends."],"forward_implications":["For any spectroscopic sample of more than about a thousand stars, stacked $\\log(\\mathrm{age})$ PDFs and mean-age histograms cannot be trusted as age tracers; the inversion must be applied first.","Age–metallicity relations derived from large, lower-precision surveys can be brought into agreement with high-precision local samples, because the inversion reproduces trends that would otherwise be hidden or distorted by isochrone smearing.","The recovered full-survey age distributions are bimodal, implying that a single mean age for a survey hides two distinct stellar populations and that splits by metallicity, distance, or sky position are needed to interpret them.","Because $\\tau([\\mathrm{Fe/H}])$ and $[\\mathrm{Fe/H}](\\tau)$ differ unless the two-dimensional age–metallicity distribution is tight, comparisons with chemical evolution models must specify which function is measured; after inversion the two nearly coincide.","Spurious low-age spikes in stacked PDFs of old, metal-poor populations disappear after inversion, so samples selected by abundance will show smoother age distributions than the raw PDFs suggest."],"supporting_citations":[{"why":"Supplies the PARSEC stellar models used to construct the mono-age mock catalogues.","marker":"Bressan et al. (2012)"},{"why":"Describes UniDAM, the isochrone-fitting pipeline whose log(age) PDFs provide both the observed C(τ) and the mock P(τ,τ̃).","marker":"Mints & Hekker (2017)"},{"why":"Provides the non-negative least-squares algorithm used to solve the discretized inversion system.","marker":"Lawson & Hanson (1995)"},{"why":"Defines the RAVE-on survey data and uncertainties used as one base survey for mocks and real-data tests.","marker":"Casey et al. (2016)"},{"why":"Defines the APOGEE survey data used as a base survey and as a real-data application for the inversion.","marker":"Majewski et al. (2017)"},{"why":"Defines the LAMOST survey data used for mono-metallicity inversions and age-metallicity trends.","marker":"Luo et al. (2015)"},{"why":"Supplies the high-precision local APOGEE age-metallicity reference that the inverted trends are compared against.","marker":"Feuillet et al. (2018)"},{"why":"Provides the chemical evolution model predictions for age-metallicity and metallicity-age relations.","marker":"Minchev et al. (2013)"}],"fun_headline_variants":["Age inversion uncovers true age mix in large surveys","Invert smeared age PDFs to recover real stellar ages","Stacked age PDFs? Invert them for the true distribution","Age inversion removes bias from stacked stellar ages","Recovering underlying age distributions via inversion"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The mono-age mock catalogues must faithfully reproduce how real survey stars get smeared in log(age) by the fitting procedure; if the stellar models or the assigned uncertainties are systematically wrong, the recovered age distribution inherits that wrongness.","fun_headline_variants_meta":{"raw":{"variants":["Age inversion uncovers true age mix in large surveys","Invert smeared age PDFs to recover real stellar ages","Stacked age PDFs? Invert them for the true distribution","Age inversion removes bias from stacked stellar ages","Recovering underlying age distributions via inversion"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000263,"raw_usage":{"total_tokens":1697,"prompt_tokens":1137,"completion_tokens":560,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":753,"completion_tokens_details":{"reasoning_tokens":483}},"tokens_in":753,"tokens_out":560,"duration_ms":5531,"temperature":1.0,"reasoning_tokens":483,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:39:30.870653+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a sample of stars from clusters with independently known ages, embed it in a survey-like set, and run the inversion: if the recovered age distribution peaks away from the known cluster ages by more than the quoted confidence intervals, the smearing kernel built from mock catalogues is not the kernel acting on real stars.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the non-negative least-squares algorithm used to solve the discretized inversion system."},{"cited_title":"The RAVE-on catalog of stellar atmospheric parameters and chemical abundances for chemo-dynamic studies in the Gaia era","cited_arxiv_id":"1609.02914","evidence_quote":"Defines the RAVE-on survey data and uncertainties used as one base survey for mocks and real-data tests."},{"cited_title":"K., Bovy, J., Holtzman, J., et al","cited_arxiv_id":null,"evidence_quote":"Supplies the high-precision local APOGEE age-metallicity reference that the inverted trends are compared against."}],"review_version":1}