{"id":"8d76ab5f-ae78-4168-b2c8-7bea2e802418","arxiv_id":"2412.04921","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Adding Gaia-based luminosities to grid-based asteroseismic modelling improves mass precision for main-sequence benchmarks but leaves radii systematically underestimated, and 1 per cent precise interferometric radii would yield under 1.5 per cent mass precision.","lead":"The paper tests whether adding Gaia parallax-based luminosities to asteroseismic grid searches improves stellar mass and radius estimates for seven benchmark stars. It finds masses become more precise (scatter down from 1.9 per cent to 0.8 per cent) while radii stay underestimated by about 1.9 per cent, and that a 1 per cent precision interferometric radius would pin masses to under 1.5 per cent.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fixed solar-calibrated mixing length in the MESA grid is the least secure link: a Teff-dependent radius bias could produce both the -1.9% radius offset and the apparent mass-scatter improvement.","rationale":"The reader's weakest_assumption is exactly the grid radius-scale bias, and I agree that it is the most load-bearing concern. The central claim is a compound: (a) mass precision improves when Gaia luminosity is added, and (b) radii are systematically underestimated by about -1.9%. Claim (b) is directly an inference about the model grid: the radii are compared to interferometric radii, so any offset is a statement about how well the grid reproduces true stellar radii. Claim (a) is also vulnerable to a grid bias because the reference masses (Set 3) come from the same grid as the compared masses (Set 2). If the grid has a Teff-dependent radius error, adding the Gaia luminosity could pull models along that biased radius-mas s relation and reduce the apparent ensemble scatter without bringing masses closer to truth. The paper provides no test of the grid's radius scale beyond the Sun (which is the calibration point and thus not a fair test) and the two alpha Cen stars, whose offsets are smaller than the sample average. A concrete test is to vary alpha_MLT (or use an independent grid) and see whether the offsets and scatter persist. The reader already flagged this as the weakest assumption and issued a CONDITIONAL verdict; my read does not change that verdict. I note other issues, such as the internal inconsistency about the number of l=0 modes used for alpha Cen A (10 available but 12 claimed in Figure 8) and the undisclosed Ball-Gizon coefficients, but these are less central to the headline results. The grid-radius concern is the single point on which the central claims rest.","tokens_in":22389,"tokens_out":8444,"duration_ms":88723,"concrete_test":"Re-run the AIMS inference for the same 8 benchmark stars with a second grid in which alpha_MLT is either free (e.g., a coarse grid of alpha_MLT values, with marginalization) or calibrated per star to the interferometric radius, keeping all other inputs identical. If the -1.9% radius offset and the 1.9%->0.8% mass-scatter reduction do not persist, the central claims are grid artifacts. As a cheaper check, test whether the Figure 2 radius residuals correlate with Teff or [Fe/H] in the direction expected from a fixed-alpha_MLT bias.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Both headline results—the radius underestimation (Section 3.1, Figure 2) and the mass-scatter reduction from 1.9% to 0.8% when a Gaia luminosity is added (Section 3.1, Figure 3)—assume the MESA grid's radius/mass scale is accurate for every benchmark star. The grid fixes alpha_MLT=1.71 (solar-calibrated) and exponential overshoot f=0.01 for all models (Section 2.1). For main-sequence stars, radius at fixed Teff and L depends sensitively on alpha_MLT; using a single solar-calibrated value can introduce a Teff- and metallicity-dependent bias. The reported offsets are indeed non-uniform: the Sun (the calibration anchor) matches to 0.2%, while Perky is inferred 4.6% low and Doris 3.4% low—a pattern consistent with a fixed-alpha_MLT bias. Furthermore, the mass-scatter comparison uses Set 3 (the same grid) as the reference (Section 3.1); if the grid biases Set 2 and Set 3 in a correlated Teff-dependent way, adding L could artificially reduce the scatter without improving true mass accuracy. Thus the headline percentages may reflect the grid's radius scale rather than the effect of Gaia-based constraints.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper uses the AIMS grid-search code with a MESA/GYRE model grid to infer masses and radii of seven main-sequence benchmark stars plus the Sun under five combinations of seismic and atmospheric constraints. The constraint sets are Set 1 ([Fe/H], Teff, L), Set 2 (individual frequencies, [Fe/H], Teff), Set 3 (Set 2 plus L), Set 4 ([Fe/H], Teff, interferometric R), and Set 5 (Set 4 plus frequencies). The main claims are that adding a parallax-based luminosity (Gaia for five of the non-solar stars, Hipparcos for alpha Cen A/B) reduces the scatter of inferred masses relative to Set 3 from 1.9% to 0.8%; that radii inferred with seismic constraints are systematically underestimated by -1.9 ± 0.7% with about 1.9% scatter compared with interferometric radii; that an interferometric radius with precision better than about 1% yields masses with precision around 1.5-2.5%; and that l = 0-only frequency sets of more than eight modes, combined with atmospheric constraints, can still give robust masses and radii. The paper also compares Gaia- and Hipparcos-based luminosities, finding a 1.4% scatter and a -0.5 ± 0.6% offset.","tokens_in":22615,"tokens_out":12528,"duration_ms":115872,"significance":"The question is relevant for PLATO preparation and for the asteroseismic community: it quantifies the value of Gaia parallax-based luminosities and of high-precision interferometric radii in grid-based forward modelling. The manuscript's strengths are that it uses a uniform pipeline for all stars, includes genuinely external validation data (interferometric radii, alpha Cen A/B dynamical masses, the Sun), and tabulates the inferred quantities so that the main comparisons are transparent. The claimed scatter values are, however, based on a small and heterogeneous sample and are tied to one particular model grid, so the results are more indicative than conclusive; with appropriate reframing and sensitivity tests they would be a useful contribution.","major_comments":[{"comment":"The interpretation of the headline radius offset and mass scatter is tied to a single grid prescription. Section 2.1 fixes alpha_MLT = 1.71 (solar-calibrated) and exponential overshoot f = 0.01 for all models. At fixed Teff and L, the model radius of a main-sequence star depends on alpha_MLT, and a fixed solar value can introduce a Teff- and metallicity-dependent bias. The pattern in Figure 2 (Sun within 0.2%, alpha Cen within about 1%, Doris about -3.4%, Perky and Saxo2 about -4.6% for Set 3) is qualitatively consistent with such a bias. The quoted systematic scatter from Eq. (11) therefore measures the internal scatter of this specific grid, not a general systematic uncertainty of the method. I ask the authors to either repeat the fits with alpha_MLT (and, if feasible, overshoot) varied within plausible ranges, or to explicitly restrict the conclusions to the adopted grid and remove the implication that the -1.9% offset is a property of grid-search modelling per se.","section":"Section 2.1 and 3.1"},{"comment":"The mass-scatter improvement from 1.9% to 0.8% is computed using the Set 3 masses from the same grid as the reference (Section 3.1: 'we considered the inferred masses from Set 3 as a reference'). For five of the seven stars there is no independent mass, so this comparison tests consistency between constraint sets, not accuracy. The external benchmarks that do exist (alpha Cen A/B dynamical masses and the Sun) should be used to report accuracy separately. In addition, the scatter is computed over N = 7 with no uncertainty on the scatter itself; the sample mixes Kepler-quality seismic data with ground-based data for alpha Cen A/B, and for the latter two the 'Gaia-based' luminosity is actually from Hipparcos (Section 2.3, Table 2). The quoted 1.9% and 0.8% values therefore need a stated N and list of included stars, a bootstrap or similar uncertainty on the scatter, and a version computed without alpha Cen A/B to show the Gaia-only effect.","section":"Section 3.1, Figure 3, and Eq. (11)"},{"comment":"There is a numerical inconsistency in the mass-precision claim from interferometric radii. Section 3.2 states that for a radius uncertainty of about 1 per cent the inferred mass uncertainty is about 2.5 per cent, and that a radius uncertainty of 0.5 per cent gives a mass uncertainty of about 1.5 per cent. The abstract and Section 4 instead state that a radius precision of about 1 per cent yields a mass precision of about 1.5 per cent. Please reconcile these statements; if the 1.5 per cent mass precision only holds at about 0.5 per cent radius precision, that is a materially different recommendation for interferometric campaigns.","section":"Section 3.2 and abstract/conclusions"}],"minor_comments":[{"comment":"The expression for chi2_[Fe/H] appears to use Teff in both the numerator and denominator rather than the observed and model [Fe/H]; this is presumably a typographical error but should be fixed.","section":"Section 2.4, Eq. (10)"},{"comment":"The phrase 'high quality sesimic data' should read 'high quality seismic data'.","section":"Section 2.2"},{"comment":"The bottom-panel y-axis label 'Fractional difference on Mass' should specify that the difference is (M_inferred - M_dynamical)/M_dynamical.","section":"Figure 6"},{"comment":"The reference to 'Kamulali et al. (in preperation)' should be 'in preparation'.","section":"Section 4"},{"comment":"The reference 'Pijpers, F. P. 2003' would be easier to read as a standard author-year citation in the text.","section":"Section 2.3"},{"comment":"Because the quoted scatter values in Section 3.1 depend on the exact posterior medians, the printed two-decimal values should be supplemented (for example in an online table or appendix) so that Eq. (11) can be reproduced.","section":"Tables 4 and 5"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the scope of the journal and the data products are useful for future PLATO-related work. The main risk is that the mass-scatter claim is presented as an accuracy improvement when it is largely a consistency measure within a single grid; the revision should address this and the alpha_MLT dependence. I do not see grounds for rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one thing to know: this paper gives PLATO work packages a concrete set of design numbers. Adding a Gaia-based luminosity to seismic plus spectroscopic constraints tightens the inferred mass scatter from 1.9% to 0.8%, the inferred radii come out about 1.9% low against interferometric radii, and an interferometric radius at or below 1% uncertainty buys mass precision around 1.5%. Those are exactly the kinds of statements mission planners want.\n\nWhat is genuinely new is that they do this uniformly across a benchmark sample with Gaia DR3 parallaxes in a grid-based forward model, and they validate against independent interferometric radii, dynamical masses for alpha Cen A and B, and the Sun. The l=0-only mode requirement (more than 8 precise radial modes plus atmospheric constraints) is a concrete, useful threshold. The paper is clearly written and honest about the model dependence of the results.\n\nThe soft spots are real but not fatal. The numbers come from seven stars with mixed data quality - four with Kepler data, two with ground-based data, plus the Sun. The scatter in equation 11 has no uncertainty attached, so the 1.9% versus 0.8% comparison is not as sharp as it looks. The mass accuracy check uses Set 3 as the reference, which is the same grid, so part of the apparent improvement from adding L could be correlated bias rather than true accuracy gain.\n\nThe bigger concern is the fixed solar-calibrated mixing length (alpha_MLT = 1.71) and overshoot f = 0.01 across the whole grid. The radius offset is not uniform: the Sun matches to 0.2%, while Perky comes out 4.6% low and Doris 3.4% low. That pattern is consistent with a Teff- and metallicity-dependent radius bias in the grid, meaning the -1.9% radius underestimate and even part of the mass-scatter reduction could be grid artifacts rather than properties of the Gaia constraint. The Ball-Gizon surface correction coefficients are also not reported, which makes the seismic fits harder to reproduce.\n\nThere is also a small internal inconsistency: the text says only 10 l=0 modes are available for alpha Cen A, yet Figure 8 shows a 12-mode run for that star. That needs clarification.\n\nOn balance, the central claims are plausible and align with earlier simulation work by Creevey et al. The exact percentages should be treated as grid-specific, not universal. This paper deserves a serious referee. With the surface correction coefficients disclosed, the l=0 mode count fixed, and the radius underestimate framed as \"for this grid,\" it would be a solid MNRAS contribution. I would send it to peer review.","headline":"A useful empirical benchmark for grid-based stellar characterization with Gaia luminosities, but the headline percentages rest on a fixed-mixing-length grid and only seven heterogeneous stars.","tokens_in":23199,"tokens_out":2188,"would_cite":true,"duration_ms":23074,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Including a Gaia-based luminosity in asteroseismic grid searches tightens inferred stellar masses to 0.8 per cent scatter while radii come out about 1.9 per cent low.","keywords":["asteroseismology","Gaia parallax","grid-based modelling","stellar mass","stellar radius","interferometric radii","main-sequence stars","benchmark stars"],"falsifier":"Re-run the same grid search on the same stars with an independent stellar grid (for instance, varying the mixing-length parameter over 1.5–2.0 or using a different evolution code) and check whether the inferred radii still fall about 1.9 per cent below the interferometric values. If the offset disappears or changes sign, the reported radius underestimate is a grid artefact; if it persists across grids, it indicates a genuine systematic in the modelling or in the interferometric comparison.","tokens_in":113,"feed_emoji":"🔭","tokens_out":10426,"duration_ms":154047,"temperature":0.7,"pith_summary":"This paper asks whether adding a luminosity derived from Gaia parallaxes to the standard set of asteroseismic and spectroscopic constraints improves the masses and radii that grid-based forward modelling returns for main-sequence stars. Working uniformly with eight benchmark stars that have interferometric radii (and, for Alpha Centauri A and B and the Sun, independent masses), the authors find that the added luminosity reduces the systematic scatter on inferred masses from 1.9 per cent to 0.8 per cent. The same tests show that grid-inferred radii come out smaller than the interferometric radii, with an offset of -1.9 ± 0.7 per cent and a scatter of about 1.9 per cent. The paper also establishes that an interferometric radius with precision of about 1 per cent or better yields masses with precision near 1.5 per cent, and that reliable masses and radii can still be obtained from l=0 oscillation frequencies alone if more than eight precise frequencies are combined with atmospheric constraints.","feed_headline":"Gaia luminosities push stellar mass scatter down to 0.8 per cent","feed_subtitle":"Better masses for solar-type stars, plus a clear warning that model radii run ~1.9 per cent low.","key_machinery":"The argument runs on a pre-computed grid of main-sequence stellar tracks generated with the MESA code, the GYRE code for adiabatic oscillation frequencies, and the AIMS code for the Markov-chain Monte Carlo grid search that interpolates between models. The optimisation minimises a chi-squared that combines seismic frequencies, corrected for surface effects with the Ball-Gizon two-term formula, with atmospheric constraints of effective temperature, metallicity, and either a Gaia-parallax luminosity or an interferometric radius. Gaia luminosities are built from DR3 parallaxes via the standard distance-luminosity relation with bolometric corrections and reddening, while interferometric radii combine angular diameters with parallaxes. Accuracy is judged by the offset and scatter between the inferred values and these independent measurements.","core_discovery":"The central claim is that a parallax-based luminosity from Gaia acts as a genuinely informative extra constraint in asteroseismic grid searches: adding it to individual oscillation frequencies, effective temperature, and metallicity lowers the scatter on inferred stellar masses from 1.9 per cent (seismic plus spectroscopic constraints) to 0.8 per cent. When the inferred radii are checked against model-independent interferometric radii, they are systematically lower by -1.9 ± 0.7 per cent with a scatter near 1.9 per cent, indicating a persistent small bias in the model grid or the comparison itself. Injecting an interferometric radius with uncertainty ≲1 per cent into the optimisation yields masses with uncertainty ≲1.5 per cent, and this benefit saturates once the radius uncertainty exceeds about 1.5 per cent, after which the seismic data dominate. Finally, using only radial l=0 oscillation frequencies, robust masses and radii are still attainable provided more than eight precise l=0 frequencies are paired with atmospheric constraints, including the Gaia-based luminosity where available.","pith_inferences":["If the -1.9 per cent radius offset survives tests with other stellar grids, asteroseismically calibrated radii—and exoplanet and Galactic properties built on them—may carry a small systematic low bias for solar-type stars.","The fixed solar-calibrated mixing length and exponential overshoot in the grid are the most plausible grid-side causes of the offset; a grid with varied mixing length would separate a grid artefact from a genuine data-model discrepancy.","The l=0-only result implies that for space missions with sparse mode visibility, such as short-cadence TESS targets, atmospheric constraints and Gaia luminosities become the limiting factor for parameter precision.","A natural extension is to apply the same uniform luminosity injection to a larger sample with eclipsing-binary masses, where mass and radius accuracy can be checked simultaneously."],"forward_implications":["Adding a Gaia-parallax luminosity to seismic and spectroscopic constraints reduces the systematic scatter on inferred stellar masses from 1.9 per cent to 0.8 per cent.","Grid-inferred radii for these benchmark stars sit about 1.9 per cent below interferometric radii, implying a small but consistent bias in the modelling or comparison.","An interferometric radius with precision ≲1 per cent, when included in the optimisation, yields stellar masses with precision ≲1.5 per cent.","When only l=0 modes are available, more than eight precise l=0 frequencies combined with atmospheric constraints are needed to obtain reliable masses and radii.","The results argue for pushing interferometric radius measurements of solar-type stars toward 1 per cent precision, a step relevant to PLATO's stellar-characterisation work packages."],"supporting_citations":[{"why":"Supplies the MESA stellar model grid used in the forward modelling.","marker":"Paxton et al. 2011, 2013, 2015, 2018, 2019"},{"why":"Computes the adiabatic oscillation frequencies for the grid models.","marker":"Townsend & Teitler 2013"},{"why":"The AIMS code performs the MCMC grid-search optimisation that infers parameters.","marker":"Rendle et al. 2019"},{"why":"The two-term surface correction that matches model frequencies to observed ones.","marker":"Ball & Gizon (2014)"},{"why":"Provides the precise parallaxes from which the added luminosity constraint is derived.","marker":"Gaia Collaboration et al. 2018, 2023"},{"why":"Interferometric angular diameters and linear radii for 16 Cyg A and B, the accuracy benchmark.","marker":"White et al. (2013)"},{"why":"Interferometric radii for Doris, Perky, and Saxo2 used as accuracy benchmarks.","marker":"Huber et al. (2012b)"},{"why":"Interferometric radii and dynamical masses for Alpha Centauri A and B.","marker":"Kervella et al. (2017)"},{"why":"The Kepler-quality oscillation frequencies adopted for the sample stars and the Sun.","marker":"Lund et al. (2017)"},{"why":"Prior simulations predicting that a model-independent radius strongly constrains mass, which the authors confirm.","marker":"Creevey et al. (2007)"}],"fun_headline_variants":["Gaia luminosity tightens stellar mass scatter to 0.8%","Gaia data sharpen stellar masses, but radii run low","Interferometric radii boost mass precision to 1.5%","Robust stellar masses from l=0 modes with Gaia","Gaia clues: better masses, but radii underestimated"],"cache_read_input_tokens":25216,"weakest_assumption_plain":"The grid's radius scale—fixed by the solar-calibrated mixing length of 1.71, the exponential overshoot parameter of 0.01, and the Ball-Gizon surface correction—is assumed to represent the true radius scale of every benchmark star to better than the reported ~1.9 per cent offset; if that scale is biased, the radius underestimate and the mass scatter comparison are artefacts of the grid rather than properties of the data.","fun_headline_variants_meta":{"raw":{"variants":["Gaia luminosity tightens stellar mass scatter to 0.8%","Gaia data sharpen stellar masses, but radii run low","Interferometric radii boost mass precision to 1.5%","Robust stellar masses from l=0 modes with Gaia","Gaia clues: better masses, but radii underestimated"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000762,"raw_usage":{"total_tokens":3432,"prompt_tokens":1048,"completion_tokens":2384,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":664,"completion_tokens_details":{"reasoning_tokens":2299}},"tokens_in":664,"tokens_out":2384,"duration_ms":16080,"temperature":1.0,"reasoning_tokens":2299,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T21:08:26.274679+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same grid search on the same stars with an independent stellar grid (for instance, varying the mixing-length parameter over 1.5–2.0 or using a different evolution code) and check whether the inferred radii still fall about 1.9 per cent below the interferometric values. If the offset disappears or changes sign, the reported radius underestimate is a grid artefact; if it persists across grids, it indicates a genuine systematic in the modelling or in the interferometric comparison.","supporting_citations":[{"cited_title":"L., Monteiro M","cited_arxiv_id":null,"evidence_quote":"Prior simulations predicting that a model-independent radius strongly constrains mass, which the authors confirm."}],"review_version":1}