{"id":"4a763294-e7c9-4e5b-a3c3-a221ea851689","arxiv_id":"2608.11440","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"New interferometric measurements give roughly 1% radius and 1.5% temperature uncertainties for 27 nearby solar-type stars, anchoring stellar models and exoplanet host characterization.","lead":"Astronomers used the CHARA interferometer to measure the angular sizes of 27 nearby Sun-like stars, then combined these with distances and brightness data to derive precise radii and temperatures. The results provide independent benchmarks for stellar evolution models and for the properties of exoplanet host stars.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Calibrator diameter zero-point is the load-bearing risk: JMMC θ_est errors are not propagated into the quoted 1% radius uncertainties, and at PAVO baselines a 3% calibrator θ error can shift science diameters by ~1%.","rationale":"The central assertion is a precision claim: 27 stars with ~1% radii and ~1.5% Teff, suitable as empirical anchors. The weakest point in the chain is the absolute visibility calibration, because a coherent error in the adopted calibrator diameters would shift every measured diameter and hence every radius and temperature in the same direction. The paper's own description of the Monte Carlo procedure confirms that calibrator θ_est values are not sampled, and the bootstrap only captures random bracket-to-bracket scatter. At the long baselines and visible wavelengths used by PAVO, the calibrators are not unresolved: their finite diameters matter, so the missing propagation is numerically relevant rather than academic. I do not regard this as a rejection: the data reduction is standard, the code is public, and the quoted random errors may well be correct. The appropriate response is to require a concrete sensitivity test, either by perturbing Table B1 diameters or by cross-checking against an independent diameter scale, before the ~1% precision claim is accepted as a total uncertainty. The reader's verdict of ACCEPT with moderate confidence is reasonable, but the load-bearing assumption should be upgraded from an implicit assumption to a demonstrated result.","tokens_in":20702,"tokens_out":9768,"duration_ms":96143,"concrete_test":"Re-run the RADPy reduction for HD 22879 (best-sampled, lowest θUD uncertainty) and HD 1461 with all calibrator θ_est values in Table B1 shifted by +1σ and −1σ, and in a second variant sample θ_est from its quoted normal prior inside the existing bootstrap. Record the resulting θLD and R for each variant. If the mean shift exceeds ~0.5% (half the quoted 1% radius error) in either target, the quoted uncertainty budget omits a calibrator zero-point term and the precision claim should be reported as conditional on the JMMC diameter scale; if the shift is <0.3%, the concern is resolved and ACCEPT stands.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The headline 1% radius / 1.5% Teff precision in Table 2 rests on the visibility calibration scale, and that scale inherits every systematic error in the calibrator angular diameters adopted in §2 (Table B1, from Bourges et al. 2014/JMMC). These calibrators are not point sources at the PAVO long baselines: for B=278 m, λ=800 nm, a calibrator with θ_est=0.20 mas has x=πθB/λ≈1.06 and V²≈0.75, not 1.0. A 3% error in θ_est therefore changes the calibrator visibility by ~1–2%, which maps to a ~1% shift in the fitted science θLD—comparable to the claimed radius uncertainty. The RADPy Monte Carlo described in §2 resamples visibilities within brackets and draws the limb-darkening coefficient from a σ=0.02 prior, but it never samples θ_est from the Table B1 errors and cannot capture a JMMC zero-point bias. The paper's statement that the bootstrap 'empirically include[s] ... bracket-to-bracket calibration variability' addresses random scatter, not a coherent scale error. Because all 27 targets are calibrated through the same catalog scale, the central claim of ~1% radii is only as strong as the JMMC diameter zero-point.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents interferometric angular diameter measurements for 27 nearby solar-type stars obtained with the PAVO beam combiner at the CHARA Array. The authors determine uniform-disk and limb-darkened angular diameters, then combine these with Gaia parallaxes and bolometric fluxes from PHOENIX SED fits to derive stellar radii, effective temperatures, and luminosities, claiming typical uncertainties of ~1% in radius and ~1.5% in effective temperature. They compare the resulting empirical H-R diagram positions with four stellar evolutionary model grids (YREC, MIST, Dartmouth, Garstec) via the Kiauhoku interpolation tool to infer masses and ages and to assess grid-to-grid systematics, including discussions of metal-poor stars, young stars, and subgiants. The central derivation, Eq. (1), is a correct rearrangement of the Stefan-Boltzmann law, and the data analysis pipeline is described in considerable detail.","tokens_in":20983,"tokens_out":5966,"duration_ms":52043,"significance":"If the quoted precision holds, the paper provides a valuable homogeneous set of empirical anchors for solar-type stars, useful for testing stellar evolutionary models and for improving exoplanet host star characterization. The manuscript has notable strengths: a well-documented Monte Carlo uncertainty treatment (bracket bootstrapping, limb-darkening prior, and a 2% systematic on bolometric flux), openly available code (RADPy on GitHub), a sample spanning a range of metallicity and evolutionary states, and a transparent discussion of model grid limitations (e.g., the lack of alpha-enhancement in the grids and edge effects for low-mass stars). The main risk is the unpropagated calibrator angular diameter uncertainty, which is load-bearing for the headline precision claim. The model-comparison section is internally consistent but currently understates the role of measurement uncertainties in the mass and age estimates.","major_comments":[{"comment":"The RADPy Monte Carlo described in §2 resamples visibilities within calibration brackets and draws the limb-darkening coefficient from a σ=0.02 prior, but it never samples the calibrator angular diameters θ_est listed in Table B1. This matters because at the PAVO long baselines (B≈278 m, λ≈800 nm) a calibrator with θ_est≈0.20 mas has x≈1.06 and V²≈0.75, so the calibrator is not a point source; a 3% error in θ_est changes the calibrator visibility by roughly 1–2% and maps to an ~1% shift in the fitted science θ_LD, comparable to the claimed radius uncertainties in Table 2. The text's statement that the bootstrap empirically includes bracket-to-bracket calibration variability does not capture a coherent scale error, and since all targets are calibrated through the same catalog scale, the quoted ~1% radius and ~1.5% Teff uncertainties are missing a systematic term. Please propagate the θ_est uncertainties through the calibration, or demonstrate explicitly why the adopted calibration is insensitive to them, and revise the uncertainty claims accordingly.","section":"§2, Table B1, §3"},{"comment":"The model-derived masses and ages in Table 3 are quoted to three decimal places, but the σ_Mass and σ_Age columns are purely grid-to-grid scatter; measurement uncertainties in Teff, L, and [Fe/H] are deliberately not propagated. The text states this, but the abstract and Section 4 nevertheless present 'mass and age estimates' and the 'agreement to within ~6%' without a total error budget. Because the input uncertainties are not negligible (e.g., 1.5% in Teff), the quoted fractional offsets should be explicitly labeled as model-systematic only, and the conclusion about mass agreement should be tempered or accompanied by a propagated measurement-error example. This does not invalidate the grid-comparison methodology, but the presentation is currently misleading.","section":"§4, Table 3"}],"minor_comments":[{"comment":"The capitalization of 'Section' is inconsistent (e.g., 'In section 4' appears after 'In Section 2' and 'Section 3').","section":"Introduction"},{"comment":"The captions for Figures C1–C5 are identical; please indicate which stars appear in each figure or state that they are grouped by observing season.","section":"Appendix C"},{"comment":"The paper should briefly justify why linear limb-darkening coefficients in the R band (Claret and Bloemen 2011) are used for PAVO data dispersed over 630–950 nm, and how sensitive the diameters are to this choice.","section":"§2"},{"comment":"The 'Bad Calibrators' table lists stars without stating the rejection criterion; a short note on why each star was excluded (e.g., binarity, rapid rotation, or poor visibility fit) would aid reproducibility.","section":"Table B2"},{"comment":"For stars flagged as unreliable or non-convergent (such as HD 22879 and HD 42807), add an explicit symbol in the table itself rather than only in the table notes.","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"The calibrator zero-point issue is the main technical risk and should be resolved before acceptance. The authors have provided a strong, reproducible analysis, but the unpropagated calibrator diameters directly affect the paper's central precision claim, so I recommend a revision that addresses this point rather than accepting the manuscript as is."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Khaled, quick read on the new Boyajian et al. diameters paper. If you work in stellar interferometry or exoplanet host characterization, it's useful. It's the seventh in the series, so the method is established: CHARA/PAVO visibilities, limb-darkened diameters, Gaia parallaxes, PHOENIX SED fluxes, then Stefan-Boltzmann radii and Teff. The new content is the 27-star sample itself — several metal-poor stars, several exoplanet hosts — and a four-grid comparison of inferred masses and ages. The measurements look carefully reduced: Monte Carlo bootstrap with bracket resampling, a prior on limb-darkening, and a 2% systematic added to bolometric fluxes. They also give honest discussion of where models fail: the metal-poor stars get ages older than the Universe, and they say so.\n\nThe soft spot is the calibrator diameter zero-point. Table B1 uses JMMC estimated diameters for calibrators, and at the long PAVO baselines a 0.2 mas calibrator has V^2 ~ 0.75, not 1. A 3% error in θ_est changes the calibrator visibility by ~1-2%, which propagates almost one-to-one into the fitted science diameter. The bootstrap you describe resamples the measured visibilities within brackets, so it captures random noise and night-to-night transfer function scatter, but it never draws the calibrator diameter from its error distribution. If the JMMC catalog has a coherent zero-point bias, every one of the 27 radii shifts in the same direction. That means the quoted ~1% radii are optimistic; the true absolute accuracy is more like 1.5-2% unless the calibrator errors average out. This is a fixable caveat, not a reason to reject.\n\nOther notes: the masses and ages from Kiauhoku are presented without propagating the input Teff/L uncertainties; the authors explicitly flag it, so it's fine for their grid-to-grid purpose. The 'model-independent' label is fair for R and Teff, since those do not invoke evolutionary models. The plots and table are clear. I'd accept this paper for peer review, but I would ask the authors to quantify the calibrator diameter systematic, even just a sensitivity test or an added quadrature term. It deserves a referee slot, and it will be a standard reference for these stars.\n\nThe calibration issue is a good excuse to bring it to reading group, if you want to see how the field handles transfer-function calibration.","headline":"Solid, incremental interferometric survey of 27 solar-type stars with careful measurements, but the quoted ~1% radii omit calibrator diameter systematics that could push the true accuracy to ~2%.","tokens_in":21549,"tokens_out":4272,"would_cite":true,"duration_ms":38262,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper reports interferometric angular diameters for 27 nearby solar-type stars and combines them with parallaxes and bolometric fluxes to obtain model-independent radii, effective temperatures, and luminosities, claiming typical…","keywords":["stellar angular diameters","optical interferometry","solar-type stars","effective temperature","stellar radii","stellar evolution models","exoplanet host stars","limb darkening"],"falsifier":"Observe a subset of the same 27 stars with an independent calibration route that does not depend on the catalog calibrators, such as lunar occultation diameters or phase-referenced interferometry, and compare the resulting radii; a systematic offset between the two methods would reveal the calibrator zero-point error.","tokens_in":20531,"feed_emoji":"🔭","tokens_out":3943,"duration_ms":48680,"temperature":0.7,"pith_summary":"This paper measures the angular sizes of 27 nearby solar-type stars with long-baseline optical interferometry and turns those diameters, together with parallaxes and bolometric fluxes, into radii, temperatures, and luminosities that do not depend on stellar models. The central claim is that these values reach about 1% accuracy in radius and about 1.5% in effective temperature, making the stars high-precision empirical benchmarks. The authors then compare the measured radii and temperatures with four independent stellar evolutionary grids and find that inferred masses agree to roughly 6% for most stars, while inferred ages scatter much more widely, with four metal-poor stars giving ages older than the Universe in every grid. If right, these results tighten the empirical anchors used to test stellar evolution theory and to characterize stars that host planets.","feed_headline":"27 nearby stars get radii accurate to ~1 percent","feed_subtitle":"Interferometric diameters plus parallaxes give model-free benchmarks for stellar evolution and exoplanet hosts.","key_machinery":"The load-bearing mechanism is long-baseline interferometric measurement of stellar angular diameters: visibility data are fit to uniform-disk and limb-darkened disk models, with flux-conserving R-band limb-darkening coefficients, and uncertainties are propagated by Monte Carlo resampling at the calibration-bracket level. Temperatures then follow from the Stefan-Boltzmann relation expressed as $T_{\\rm eff}(\\mathrm{K}) = 2341 (F_{\\rm bol}/\\theta_{\\rm LD}^2)^{0.25}$, where $F_{\\rm bol}$ is the bolometric flux and $\\theta_{\\rm LD}$ the limb-darkened angular diameter, so the final radii and temperatures are anchored directly to observables rather than to stellar models.","core_discovery":"The paper's central discovery is a homogeneous set of model-independent fundamental properties for 27 solar-type dwarfs and mildly evolved subgiants spanning 4830 to 6390 K. For each star, limb-darkened angular diameters from interferometric visibility fits are combined with bolometric fluxes from spectral-energy-distribution fitting and zero-point-corrected Gaia parallaxes, giving radii with typical 1% uncertainties and effective temperatures with typical 1.5% uncertainties. Cross-comparison with four stellar evolutionary model grids shows that mass estimates agree within about 6% for most stars, while age estimates differ far more between grids, especially for metal-poor, alpha-enhanced stars whose model ages can exceed the age of the Universe, and for young and low-mass stars where grid tracks are closely spaced or insensitive to age.","pith_inferences":["If the calibrator zero-point is later revised, all 27 radii shift coherently; the relative ordering of stars is largely preserved, so population-level comparisons are safer than any single absolute radius.","Applying the same pipeline to a volume-limited sample of solar-type stars would turn these benchmarks into a statistical test of the mass-radius relation in the solar neighborhood, which the current target list cannot do by itself.","The authors' planned re-analysis with models that include alpha-enhanced compositions self-consistently is a natural test: if those models bring the four metal-poor stars below 14 Gyr, it would directly confirm that missing alpha-enhancement, not the data, drives the unphysical ages.","A direct comparison of the measured temperatures against asteroseismic temperatures for any stars in the sample that overlap with pulsation surveys would provide an independent check on the 1.5% temperature claim."],"forward_implications":["The 27 stars become model-independent calibration points on the H-R diagram, giving stellar evolution codes fixed empirical targets near the main sequence and subgiant branch.","Exoplanet host stars in the sample get radii and temperatures that can shrink the fractional uncertainties in planet radius and density for their known planets.","The large grid-to-grid age scatter, including unphysical ages for four metal-poor stars, identifies where current evolutionary models need better treatment of abundances, alpha-enhancement, and boundary conditions.","Inferred masses agreeing within about 6% across the four grids supports moderate confidence in model-based mass estimates for solar-type stars, while ages should be treated as plausible ranges rather than precise values.","The two evolved subgiants show that model discrimination is strong on the subgiant branch but weak near the base of the giant branch, where tracks overlap and inferred masses and ages become unreliable."],"supporting_citations":[{"why":"Supplies the calibrator angular diameters used to set the visibility scale for every science target.","marker":"Bourges et al. (2014)"},{"why":"Provides the parallaxes that turn angular diameters into linear radii and bolometric fluxes into luminosities.","marker":"Gaia Collaboration (2020)"},{"why":"Gives the zero-point correction applied to the Gaia parallaxes.","marker":"Lindegren et al. (2021)"},{"why":"Defines the uniform-disk and limb-darkened visibility functions used in the diameter fits.","marker":"Hanbury Brown et al. (1974)"},{"why":"Supplies the flux-conserving linear limb-darkening coefficients in R band.","marker":"Claret and Bloemen (2011)"},{"why":"Provides the weighted mean metallicities input to the stellar evolution grids.","marker":"Soubiran et al. (2022)"},{"why":"Provides the Kiauhoku interpolation toolkit that carries out mass and age estimation from the model grids.","marker":"Claytor et al. (2020)"},{"why":"Implements the Kiauhoku-based fitting procedure used here for the four-grid comparison.","marker":"Tayar et al. (2022)"},{"why":"The RADPy pipeline performs the visibility fitting, SED fitting, and bootstrap uncertainty estimation.","marker":"Elliott (2025)"}],"fun_headline_variants":["27 sunlike stars get 1% radii from interferometry","Model-free benchmarks: 27 stars with precise sizes","Solar-type stars measured to 1% radius accuracy","27 nearby stars: precise diameters, temps, and radii","Interferometry yields 1% radii for 27 sunlike stars"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The adopted angular diameters of the calibrator stars are assumed to be accurate and unbiased, so any systematic error in those diameters shifts every measured stellar diameter, radius, and temperature together.","fun_headline_variants_meta":{"raw":{"variants":["27 sunlike stars get 1% radii from interferometry","Model-free benchmarks: 27 stars with precise sizes","Solar-type stars measured to 1% radius accuracy","27 nearby stars: precise diameters, temps, and radii","Interferometry yields 1% radii for 27 sunlike stars"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000227,"raw_usage":{"total_tokens":1439,"prompt_tokens":879,"completion_tokens":560,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":495,"completion_tokens_details":{"reasoning_tokens":477}},"tokens_in":495,"tokens_out":560,"duration_ms":26384,"temperature":1.0,"reasoning_tokens":477,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:12:12.216298+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Observe a subset of the same 27 stars with an independent calibration route that does not depend on the catalog calibrators, such as lunar occultation diameters or phase-referenced interferometry, and compare the resulting radii; a systematic offset between the two methods would reveal the calibrator zero-point error.","supporting_citations":[],"review_version":1}