{"id":"28f8833c-839d-4094-a590-4b1a5e1ba54a","arxiv_id":"2608.07331","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"After age-mass matching and bias correction, host and non-host stars are chemically similar overall, but Earth-like and sub-Neptune hosts show opposite O, Mg, and Si trends, and sub-Neptune hosts are less active and born closer to the Galactic center.","lead":"This paper compares 28,383 Kepler stars with and without detected planets, matching them by age and mass, and finds that most apparent chemical differences vanish once matching and bias correction are applied. After splitting planets by size, Earth-like and sub-Neptune host stars show opposite abundance trends and different birth radii, hinting that the two planet types form under different conditions.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline abundance trends are below the paper's own 3-sigma detection thresholds and are selected post hoc, so the radius-dependent chemistry claim is not currently supported by the data.","rationale":"The reader's weakest-assumption choice, the fake-host baseline correction, is legitimate and may be the most fundamental methodological risk: if the baseline subtraction is inappropriate, all corrected offsets become suspect. However, that concern would require a new fake-host experiment to confirm. The concern I identify is already established from the paper's own text: Section 3.6 gives the minimum intrinsic abundance offset detectable at 3-sigma with 68% power, and the observed high-abundance-end offsets are roughly 0.02-0.05 dex, well below those thresholds. The paper is honest about its limitations and does not claim a formal detection, but the abstract's phrasing that radius separation 'reveals distinct trends' overstates what the data can support, especially because the trends are chosen post hoc from many bins with no multiple-testing correction. The birth-radius result is partly a reparameterization of [Fe/H] and age, so it is not independent confirmation. The activity offset is also small and could be influenced by Kepler detection completeness. For these reasons, the appropriate disposition is UNVERDICTED: the central claim cannot be evaluated as a positive finding until a permutation/FDR analysis is supplied and shown to survive. If it does survive, the paper would merit a conditional acceptance; if it does not, the claim should be downgraded to an upper limit.","tokens_in":18107,"tokens_out":13093,"duration_ms":125161,"concrete_test":"Take the exact high-abundance-end bins highlighted in Figures 7-9 and run a permutation test in which planet-radius labels (Earth-like vs sub-Neptune) are randomly reassigned among the planet hosts, recomputing the bias-corrected host-non-host offsets under each permutation. Apply false-discovery-rate control over the full grid of elements, abundance bins, and radius classes, and report adjusted q-values. Separately, compare the maximum observed absolute corrected offset in any claimed bin with the Section 3.6 thresholds of 0.070, 0.112, 0.112, and 0.079 dex for C, O, Mg, and Si. If no high-end trend survives at q < 0.1, or if all observed amplitudes remain below those thresholds, the radius-dependent chemistry result should be reported as an upper limit rather than as evidence for distinct formation environments.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The abundance pillar of the central claim fails the paper's own sensitivity criterion. Section 3.6 states that intrinsic differences of approximately 0.070 dex in [C/Fe], 0.112 dex in [O/Fe], 0.112 dex in [Mg/Fe], and 0.079 dex in [Si/Fe] are required to achieve a 68% chance of a 3-sigma detection with the current sample sizes. The headline trends in Figures 7-9 are binned offsets of order +/-0.02 to 0.05 dex in the high-abundance bins, with the [O/Fe] offset for Earth-like hosts reaching only about +0.02 dex. These offsets are therefore smaller than the minimum intrinsic difference this sample can detect. In addition, the trends are selected from a grid of 4 elements, about 8 abundance bins, and 3 radius classes, with no multiple-testing control presented. The paper's caveats do not resolve the tension: the title says 'hints,' but the abstract states that separating planets by radius 'reveals distinct trends.' The concrete magnitude of the offsets, combined with the post hoc selection, means the radius-dependent chemistry result is not statistically anchored.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper compares 629 Kepler planet-host stars with age- and mass-matched non-host stars drawn from a homogeneous Kepler-LAMOST-Gaia sample of 28,383 dwarfs and subgiants, applying a fake-host Monte Carlo baseline correction to remove distribution-induced matching biases. The authors report that most host-non-host differences vanish for the full planet sample, but that after splitting planets by radius, Earth-like hosts appear more O-rich yet Mg- and Si-poor at the high-abundance end, while sub-Neptune hosts show opposite tendencies, lower chromospheric activity, and smaller birth radii. The paper also reports tentative contrasts between hot-Jupiter and longer-period Jupiter hosts. The central claim is that these radius-dependent differences indicate distinct formation environments for Earth-like planets and sub-Neptunes, though the authors frame the results as hints in the title and discuss statistical sensitivity limitations.","tokens_in":18374,"tokens_out":4036,"duration_ms":35321,"significance":"If the radius-dependent trends are real, the paper would provide a valuable population-level clue that small-planet host stars differ systematically from non-host stars only when separated by planet size, potentially linking formation pathways to stellar chemistry, activity, and Galactic birth environment. The methodology is a strength: the sample is homogeneous and well-characterized, the age-mass matching is clearly described, the fake-host baseline correction is a thoughtful treatment of a known bias, and the Monte Carlo sensitivity analysis is an honest assessment of the detectability limits. However, the headline abundance trends are of smaller amplitude than the paper's own stated detection thresholds, and the post hoc selection of the high-abundance end weakens the statistical foundation. The activity and birth-radius results are also modest in amplitude and would benefit from more explicit significance testing.","major_comments":[{"comment":"The paper's own Monte Carlo sensitivity analysis (Section 3.6, Figure 13) states that intrinsic abundance differences of approximately 0.070 dex in [C/Fe], 0.112 dex in [O/Fe], 0.112 dex in [Mg/Fe], and 0.079 dex in [Si/Fe] are required to achieve a 68% probability of a 3-sigma detection with the current sample sizes. The headline abundance offsets shown in Figures 7-9 are of order 0.02-0.05 dex, with the [O/Fe] offset for Earth-like hosts reaching only about +0.02 dex at the high-abundance end. Therefore, the abstract's claim that 'separating planets by radius reveals distinct trends' in O, Mg, and Si is not supported by the paper's own sensitivity criterion; these trends are below the detection threshold and should be presented as non-detections or as upper limits, not as revealed trends.","section":"Section 3.6, Figures 7-9"},{"comment":"The radius-dependent abundance trends are identified by scanning a grid of four elements, roughly eight abundance bins, and three planet-radius classes, and then focusing on the high-abundance end where the patterns appear. No multiple-testing correction (e.g., false discovery rate) or permutation-based control is presented. Under the null hypothesis of no host-non-host differences, such a scan would be expected to produce some 'opposite tendencies' in isolated bins by chance. The central claim depends crucially on these selected bins, so the analysis should report how many independent comparisons were made and whether the trends survive a multiple-testing correction.","section":"Section 3.2"},{"comment":"The fake-host baseline correction subtracts the mean offset obtained when random non-host stars are matched to other non-host stars, implicitly assuming that the matching-induced bias for true host stars is identical to that for non-host stars of the same X. If host stars are drawn from a different parent distribution in age-mass or abundance space, or if the bias depends on variables not captured by the eight-bin interpolation, the correction could over-subtract genuine host-non-host signals or leave residual artifacts. Because the headline trends are only 0.02-0.05 dex, the analysis should include robustness tests varying the number of nearest neighbors k (e.g., 20 and 100), the number of bias-correction bins (e.g., 6 and 12), and the interpolation scheme, to demonstrate that the qualitative conclusions are insensitive to these choices.","section":"Section 2.6, Eq. (5)"},{"comment":"The activity offset for sub-Neptune hosts (about -0.04 dex at log R'HK > -5) and the birth-radius offsets (about -0.2 to -0.4 kpc in the R_b ~4-7 kpc range) are presented as coherent trends, but their statistical significance is not quantified in a way that accounts for the same post hoc selection of bins and sub-populations. Please provide significance levels (e.g., p-values or confidence intervals that include a multiple-testing correction) for these specific claims, or explicitly label them as tentative.","section":"Section 3.3 and 3.4"}],"minor_comments":[{"comment":"The phrase 'more Mg- and Si-poor' should be 'more Mg-poor and Si-poor' for parallel construction and clarity.","section":"Abstract and Section 3.2"},{"comment":"The caption and text refer to 'hot-Jupiter and longer-period Jupiters' with inconsistent singular/plural usage; use 'hot-Jupiter hosts and longer-period Jupiter hosts' throughout.","section":"Figure 12"},{"comment":"The reference list contains Johnson et al. (2010) twice with identical bibliographic information (PASP, 122, 905); the duplicate should be removed.","section":"References"},{"comment":"The text says 'using the NearestNeighbors algorithm implemented in scikit-learn'; more precisely, the NearestNeighbors class from scikit-learn is used to perform the matching. The phrasing is acceptable but could be clarified for readers unfamiliar with the library.","section":"Section 2.6"},{"comment":"Open symbols marking edge bins are difficult to distinguish from filled symbols in grayscale; consider using different marker shapes or adding a legend entry for edge bins.","section":"Figures 5-11"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for ApJ and the methodology is largely sound, but the central abundance claim is not supported by the paper's own sensitivity analysis. The authors would need to either substantially temper the abstract and summary and present the results as sub-threshold hints, or add a more sensitive analysis (e.g., continuous regression) and proper multiple-testing control. The current gap between the abstract's 'reveals distinct trends' and the body's 'should be interpreted with caution' is a consistency issue that the editor may want the authors to address. I do not see grounds for rejection, as the methodological framework is reusable and the activity/birth-radius results, if confirmed with better statistics, could still be valuable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The main new thing here is the fake-host baseline correction applied to a homogeneous Kepler–LAMOST–Gaia sample, and the radius-stratified host–nonhost comparisons for O, Mg, Si, activity, and birth radius. The method is thoughtful: age–mass matching with k=50, bootstrap uncertainties, a Monte Carlo sensitivity analysis, and an explicit statement that non-hosts may contain undetected planets. The full-sample null results after correction are a useful caution to the field. I also respect that the authors report their own detection limits rather than hiding them.\n\nThat said, the central claim is weaker than the abstract suggests. The headline abundance offsets in the high-abundance bins are roughly 0.02–0.05 dex, while the paper's own sensitivity analysis (Figure 13) says intrinsic differences of 0.07–0.11 dex are needed for a 3-sigma detection with these sample sizes. So the abundance trends are below the threshold the authors themselves define. On top of that, the high-abundance end is selected after looking at the data, and there is no multiple-testing control across four elements, ~8 bins, and three radius classes. The stress-test note holds up on reading.\n\nThe birth-radius result is also less independent than it appears. Equation (1) makes R_birth a deterministic function of [Fe/H] and age, so after matching on age and mass, the R_birth comparison is largely a rescaling of the [Fe/H] comparison. The activity offset for sub-Neptune hosts is small and confined to a selected regime. None of this is fatal, but it means the paper's strongest advertised conclusions are not statistically anchored.\n\nWhat the paper does well is set up the framework. The fake-host correction is a genuine contribution that should be applied in future host–nonhost studies, and the honesty about sensitivity is commendable. But the authors should either soften the abstract to match the \"hints\" language used later, or add the missing multiple-testing control and present the abundance trends as exploratory. No code or data are provided, which makes it hard to check the subtle bias correction.\n\nWho is this for? Researchers working on host-star comparisons and planet–formation environment studies. It deserves a serious referee: the method is worth engaging with, and the null results are valuable, even if the headline trends do not survive close inspection. I would send it to review, with the expectation of major revision to temper claims and add proper statistical controls.","headline":"A careful, honest study with a useful bias-correction method, but the headline radius-dependent chemistry trends are below the sample's own detection thresholds and look post hoc.","tokens_in":18886,"tokens_out":1711,"would_cite":true,"duration_ms":16228,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Host stars differ from similar non-host stars only after planets are split by size, with Earth-like and sub-Neptune hosts showing opposite chemical trends.","keywords":["exoplanet host stars","Kepler","LAMOST","stellar abundances","stellar activity","birth radius","sub-Neptunes","planet formation"],"falsifier":"Recompute the corrected offsets using only stars with near-complete planet detection, where non-detection can rule out planets down to Earth size (e.g., using injection-recovery completeness maps for each Kepler star), and check whether the high-abundance [O/Fe], [Mg/Fe], and [Si/Fe] offsets for Earth-like and sub-Neptune hosts persist; if they vanish or reverse sign, the central claim fails.","tokens_in":17912,"feed_emoji":"🪐","tokens_out":9089,"duration_ms":72783,"temperature":0.7,"pith_summary":"Using 28,383 Kepler–LAMOST–Gaia dwarf and subgiant stars, including 629 planet hosts, this paper asks whether planet-hosting stars differ from otherwise similar stars without detected planets. After matching each host to 50 non-host stars of similar age and mass and correcting for matching biases, the full planet sample shows almost no host–non-host differences. The differences appear when planets are separated by radius: at the high-abundance end of [O/Fe], [Mg/Fe], and [Si/Fe], hosts of Earth-like planets (radii below $2\\,R_\\oplus$) tend to be oxygen-rich but magnesium-poor and silicon-poor relative to matched non-hosts, while sub-Neptune hosts (radii $2$–$4\\,R_\\oplus$) tend toward the opposite. Sub-Neptune hosts also tend to be less chromospherically active and to have been born closer to the Galactic center. The result matters because it suggests that Earth-like planets and sub-Neptunes form in distinct chemical and Galactic environments rather than through one universal formation process.","feed_headline":"Planet size splits host stars into two chemical families","feed_subtitle":"Age-mass matched non-hosts reveal opposite abundance trends for Earth-like vs sub-Neptune systems","key_machinery":"The analysis rests on age–mass matched control samples plus a Monte Carlo fake-host baseline correction. Each host star is compared with the mean of its 50 nearest non-host stars in normalized age–mass space; then, for each parameter $X$, a fake-host experiment using only non-host stars measures the average offset that the matching procedure alone produces as a function of $X$, and this baseline is subtracted via $\\Delta X_{\\rm corr} = \\Delta X_{\\rm raw} - \\Delta X_{\\rm bias}(X)$. Planet radii are recomputed from Kepler DR25 transit depths and updated stellar radii, and the radius valley near $2\\,R_\\oplus$ sets the division between Earth-like planets and sub-Neptunes.","core_discovery":"The central discovery claim is that host–non-host differences in stellar chemistry, activity, and birth radius are largely absent when all detected planets are pooled together, but become visible once planets are divided by size. After age–mass matching and a fake-host baseline correction, Earth-like hosts ($R_p<2\\,R_\\oplus$) show positive corrected offsets in [O/Fe] and negative corrected offsets in [Mg/Fe] and [Si/Fe] at the high-abundance end, whereas sub-Neptune hosts ($2 \\le R_p < 4\\,R_\\oplus$) show the opposite pattern. Sub-Neptune hosts additionally show lower chromospheric activity in the relatively active regime and systematically smaller birth radii over $R_b\\sim4$–$7$ kpc, while Earth-like hosts show no such birth-radius offset. The paper reads these patterns as evidence that Earth-like planets and sub-Neptunes may form in different chemical environments, inherit different rocky building blocks, and follow different evolutionary paths, with the tentative result that hot-Jupiter hosts are more metal-rich, more active, and born at smaller Galactic radii than hosts of longer-period Jupiters.","pith_inferences":["If the trends survive, one testable consequence is that sub-Neptune atmospheric compositions (e.g., C/O ratios from transmission spectra) should differ systematically from rocky-planet systems, reflecting the different Mg/Si and O abundances of their host stars.","The fake-host correction's validity could be checked by injecting synthetic host–non-host differences of known size into mock catalogs and verifying that the corrected offsets recover the input; this would isolate any residual bias in small high-abundance bins.","A larger sample, ideally with measured detection completeness per star, could turn the tentative hot-Jupiter versus longer-period Jupiter signal into a robust test of whether migration pathways depend on stellar metallicity and birth environment.","If sub-Neptune hosts are truly born at smaller Galactic radii, then planet occurrence models that include Galactic chemical evolution should predict a radial dependence in sub-Neptune frequency; this could be checked with future transit surveys."],"forward_implications":["Future host–non-host comparisons that do not split planets by radius may miss real differences, so demographic studies should analyze planet-size classes separately.","The opposite Mg/Si trends imply that Earth-like and sub-Neptune systems may inherit systematically different rocky building blocks, which could be tested through interior-composition modeling of individual systems.","Lower chromospheric activity among sub-Neptune hosts is consistent with weaker high-energy irradiation that helps preserve volatile envelopes, linking stellar activity to atmospheric retention.","Sub-Neptunes' smaller birth radii suggest inner-disk formation environments, connecting planet demographics to the chemical and dynamical evolution of the Milky Way.","The tentative hot-Jupiter versus longer-period Jupiter contrasts indicate that close-in giant planet formation may be tied to metal-rich, active, inner-disk stellar populations."],"supporting_citations":[{"why":"Supplies the quality-controlled Kepler–LAMOST–Gaia parent sample of dwarfs and subgiants from which hosts and matched non-hosts are drawn.","marker":"X. Chen et al. 2025"},{"why":"Provides the LAMOST DR9 DD-Payne atmospheric parameters and elemental abundances used to characterize every star in the sample.","marker":"M. Zhang et al. 2025"},{"why":"Original Kepler stellar properties catalog that is the starting point for the sample construction.","marker":"T. A. Berger et al. 2020"},{"why":"Kepler DR25 planet catalog whose transit depths are combined with updated stellar radii to compute planet radii.","marker":"S. E. Thompson et al. 2018"},{"why":"Calibration tables relating ISM metallicity gradient to age, used to estimate stellar birth radii.","marker":"Y. L. Lu et al. 2024"},{"why":"Calibration that converts S-index measurements to the R+HK chromospheric activity index.","marker":"M. Mittag et al. 2013"},{"why":"Provides the methodology for deriving stellar luminosities, radii, masses, and ages used throughout the analysis.","marker":"X. Chen et al. 2026"},{"why":"Framework used to derive precise ages for large dwarf and subgiant samples from LAMOST and Gaia constraints.","marker":"M. Xiang & H.-W. Rix 2022"}],"fun_headline_variants":["Size separates host star chemistry, not planet presence","Earth-like and sub-Neptune hosts show opposite abundance trends","Host-non-host differences emerge only by planet radius","Sub-Neptune hosts have smaller birth radii and lower activity"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the fake-host correction removes exactly the bias created by the matching process, and does not remove any real difference between planet hosts and non-hosts.","fun_headline_variants_meta":{"raw":{"variants":["Size separates host star chemistry, not planet presence","Earth-like and sub-Neptune hosts show opposite abundance trends","Host-non-host differences emerge only by planet radius","Sub-Neptune hosts have smaller birth radii and lower activity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001226,"raw_usage":{"total_tokens":5091,"prompt_tokens":1046,"completion_tokens":4045,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":662,"completion_tokens_details":{"reasoning_tokens":3979}},"tokens_in":662,"tokens_out":4045,"duration_ms":25191,"temperature":1.0,"reasoning_tokens":3979,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T10:10:42.159973+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the corrected offsets using only stars with near-complete planet detection, where non-detection can rule out planets down to Earth size (e.g., using injection-recovery completeness maps for each Kepler star), and check whether the high-abundance [O/Fe], [Mg/Fe], and [Si/Fe] offsets for Earth-like and sub-Neptune hosts persist; if they vanish or reverse sign, the central claim fails.","supporting_citations":[{"cited_title":"2025, ApJ, 995, 33, doi: 10.3847/1538-4357/ae12a2","cited_arxiv_id":null,"evidence_quote":"Supplies the quality-controlled Kepler–LAMOST–Gaia parent sample of dwarfs and subgiants from which hosts and matched non-hosts are drawn."},{"cited_title":"2025, ApJS, 279, 5, doi: 10.3847/1538-4365/add016","cited_arxiv_id":null,"evidence_quote":"Provides the LAMOST DR9 DD-Payne atmospheric parameters and elemental abundances used to characterize every star in the sample."}],"review_version":1}