{"id":"524c4e0b-fee4-4111-9e0c-5ebdb238b8f9","arxiv_id":"2504.21763","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Missing planets, especially middle ones, inflate a system's gap complexity, so the observed even spacing in Kepler multi-planet systems is likely astrophysical.","lead":"By deleting planets from observed and simulated Kepler systems, this paper shows that missed planets make planetary spacing look less regular, while mass similarity and coplanarity are barely affected. The results strengthen the case that the 'peas-in-a-pod' pattern of evenly spaced, similarly sized planets is real astrophysics rather than a Kepler detection artifact.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The astrophysical-spacing conclusion depends on the SysSim model's simulated missing-planet population, which the paper itself shows overpredicts irregularity; an independent detection model could reverse the direction.","rationale":"The reader's weakest assumption correctly identifies SysSim fidelity as the load-bearing condition. The L24 jackknife establishes the directional effect for detected planets, but the extrapolation to undetected planets requires the SysSim simulation. The paper is transparent about the model's imperfection and hedges with 'likely astrophysical,' so the central claim is modest enough to accept. The multiplicity-cut mismatch in Section 2.3 (L24 4+ vs SysSim 3+) is a secondary flaw that does not affect the within-sample directional result. The proposed independent-completeness test would substantially strengthen confidence, but its absence does not invalidate the paper.","tokens_in":11805,"tokens_out":16659,"duration_ms":172284,"concrete_test":"Recompute the SysSim underlying-vs-detected gap complexity change using an independent Kepler detection completeness map (e.g., the DR25 pipeline completeness from Burke et al. 2015) applied to the same underlying SysSim catalogs. If the median Δ50% system for gap complexity is no longer positive (or is <0.05), the paper's directional conclusion is model-dependent.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that observed even spacing is astrophysical requires that the planets missed by Kepler are ones whose removal increases gap complexity. The L24 jackknife cannot establish this, because it removes already-detected planets; only the SysSim experiment simulates a realistic missing-planet population (Section 2.2). Yet Section 2.3 admits the Synthetic-Detected systems overpredict gap complexity relative to L24 (median 0.1 higher). This overprediction may mean SysSim's detection bias removes too many middle planets, inflating the +0.16 effect attributed to missing middle planets. If the real Kepler pipeline misses fewer middle planets (e.g., because small planets at intermediate periods are more detectable than modeled), the bias on gap complexity could shrink below significance, and the observed low gap complexity could still be consistent with detection bias. The paper's conclusion is therefore only as strong as the SysSim model's fidelity in exactly the regime (which planets are missed) that the model is known to represent imperfectly.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This Letter asks how the failure to detect one or more planets in high-multiplicity Kepler systems biases three architectural metrics: gap complexity (spacing regularity), mass partitioning, and impact-parameter dispersion. The authors first perform a jackknife experiment on the observed Lissauer et al. (2024) catalog, removing one planet at a time from systems with four or more transiting planets. They find that removing a middle planet increases gap complexity (median change +0.15), while removing an edge planet has a negligible effect; mass partitioning and impact-parameter dispersion are essentially unchanged. They then repeat the experiment using 100 pregenerated SysSim maximum-AMD synthetic catalogs, comparing underlying systems to detected systems after applying a Kepler-like detection pipeline. The synthetic experiment confirms the direction of the bias: missing at least one middle planet raises gap complexity by a median of +0.16, while missing only edge planets changes it by less than 0.01. Finally, the authors compare the observed L24 systems to the SysSim detected systems and find that L24 systems have lower gap complexity (median 0.08 vs. 0.18), despite the fact that detection bias tends to increase gap complexity. They conclude that the observed even spacing in high-multiplicity systems is likely astrophysical rather than a product of Kepler's detection biases, and that an underlying model producing more regular spacings is needed.","tokens_in":11972,"tokens_out":6578,"duration_ms":71234,"significance":"If the result holds, this is a timely and useful contribution to the debate over the origin of the 'peas-in-a-pod' pattern. The experimental design is clean and falsifiable: it uses well-defined metrics, a homogeneous observed catalog, and a synthetic population with known underlying architectures. The repeated mass-radius sampling (100 times) and the use of 100 SysSim catalogs demonstrate stability of the qualitative conclusions. The comparison between observed and synthetic detected systems is a genuine model-data test that was not used to fit the model, which strengthens the claim that the discrepancy in gap complexity is a real feature of the Kepler population. The distinction between edge and middle planet removals is physically intuitive and well supported by both the observed and synthetic experiments. The paper is appropriately cautious about the imperfections of the SysSim model, although the central inference does lean on the assumption that the direction of the bias (increasing gap complexity) is robust to details of the detection pipeline.","major_comments":[],"minor_comments":[{"comment":"The KS and AD tests treat the 'One Planet Removed' sample as 357 independent systems, but these jackknife samples are clustered within 80 parent systems and are therefore not independent. A paired test on the per-system differences or a bootstrap resampled at the system level would give more defensible p-values; the reported very small p-values likely overstate the significance, although the median effect sizes are large enough that the qualitative conclusion is probably unaffected.","section":"Section 2.1, Figure 2 and Table 1"},{"comment":"The sentence 'This result supports that SysSim underestimates the number of systems with highly uniform spacings' conflates two possible explanations: the underlying SysSim spacing distribution may be too irregular, or the detection pipeline may remove too many middle planets. Since the paper's main conclusion depends on the direction rather than the magnitude of the bias, I suggest adding a sentence clarifying that the robust result is the positive sign of the gap-complexity bias, and that the overprediction of the detected SysSim systems does not change the direction of the inference.","section":"Section 2.3, Figure 4"},{"comment":"The statement 'the observed systems have more evenly spaced planets than the observation-bias-applied synthetic systems' is presented as a key difference. Because the SysSim model is known to be imperfect, I recommend explicitly noting in the abstract or conclusions that this difference could in principle reflect a deficiency in the synthetic model, and that the astrophysical conclusion rests on the direction of the bias in both experiments rather than on the absolute agreement between L24 and SysSim.","section":"Abstract and Section 2.3"},{"comment":"The argument against false positives as a driver of the pattern is brief but adequate; however, the sentence 'This false alarm rate is so low that we expect our Astrophysical—Observed catalog to have <1 false positive' should specify that this expectation applies to the high-multiplicity subset under consideration, since the Lissauer et al. (2014) estimate was made for a different sample.","section":"Section 2.4"}],"recommendation":"minor_revision","confidential_remarks":"The manuscript is well suited for ApJL. The central claim is sound and the experiments are clearly described. The main technical caveat—non-independence of the jackknife samples—does not appear to overturn the qualitative result, but the authors should address it to make the statistical claims rigorous. The SysSim model is co-authored by one of the present authors, but the paper uses it as a pre-existing, independently published model and performs a falsifiable comparison, so I do not see a circularity problem. The minor revisions requested are local and do not require new observations or major reanalysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a clean, convincing demonstration that missing a planet—especially a middle one—makes a multi-planet system look less regularly spaced, not more. That directly undercuts the idea that Kepler's detection biases produce the peas-in-a-pod pattern. The paper is modest in scope but well-executed.\n\nWhat's new: although the directional expectation was floating around in prior work, this is the first explicit, systematic test. The jackknife on L24 systems and the comparison of SysSim underlying vs detected catalogs isolate which architecture metrics respond to missed planets. The two experiments agree: gap complexity increases, mass partitioning and coplanarity are basically unaffected.\n\nWhat it does well: the metrics are clearly defined, the synthetic analysis handles the large-sample p-value problem by computing per-catalog p-values, and the mass-radius sampling is repeated 100 times to show stability. The authors also check three other catalogs and get consistent results. They are appropriately candid about SysSim's limitations.\n\nSoft spots: the SysSim overprediction of gap complexity in the detected sample is a real worry, and the stress-test note is right to flag it. If SysSim's detection pipeline misses too many middle planets, the +0.16 effect size could be an overestimate. But the direction is not in serious doubt: the L24 jackknife already shows, free of any model, that removing a middle planet from an observed system increases gap complexity. For detection bias to explain the observed even spacing, one would need missing planets to preferentially smooth spacing, which is the opposite of what you get for any realistic transit-detection bias. So the paper's central conclusion—that detection bias alone is insufficient—holds, even if the magnitude is uncertain. The non-independence of jackknife samples is a minor statistical quibble; the coplanarity proxy is crude, but the result there is a null effect. No code release is a minor annoyance, not a flaw.\n\nThis paper is for anyone working on Kepler architecture statistics or planet formation models that try to reproduce peas-in-a-pod. It deserves a serious referee; the analysis is logical, the claims are properly scoped, and the limitations are stated rather than hidden. I'd be happy to see it published after minor revision.","headline":"A clean, modest paper showing that missing a middle planet increases gap complexity, which supports the astrophysical origin of peas-in-a-pod and warrants a serious referee.","tokens_in":12514,"tokens_out":3175,"would_cite":true,"duration_ms":31669,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The failure to detect a planet, especially a middle one, biases planetary systems toward irregular spacing, so the even spacing seen in Kepler multi-planet systems is a genuine astrophysical signal rather than a detection artifact.","keywords":["exoplanet systems","Kepler","gap complexity","detection bias","peas-in-a-pod","planetary architecture","transit photometry"],"falsifier":"Re-run the experiment on systems where the missing planets are already known: take the L24 four-plus planet systems, add every independently confirmed transiting or non-transiting planet back into those systems, and recompute the median gap complexity. The paper's claim predicts a decrease comparable to the +0.15 middle-removal offset; if the median instead rises or stays flat, the asserted bias direction would be refuted.","tokens_in":11583,"feed_emoji":"🪐","tokens_out":10209,"duration_ms":100961,"temperature":0.7,"pith_summary":"The paper asks whether failing to detect a planet can distort the architectural story we read from a multi-planet system. Using the 80 observed Kepler systems with four or more transiting planets and thousands of synthetic systems, the authors remove planets and recompute three metrics: gap complexity (the irregularity of orbital spacings), mass partitioning (how unequal the planet masses are), and impact-parameter dispersion (how coplanar the orbits are). They find that losing a planet, especially one in the middle of the system, pushes gap complexity upward, whereas mass uniformity and coplanarity are barely affected. Because the Kepler-detected systems are more evenly spaced than the synthetic catalogs after detection bias is applied, the paper concludes that the even spacing of high-multiplicity systems is likely astrophysical rather than an artifact of missing planets.","feed_headline":"Missing middle planets fake gaps in Kepler systems","feed_subtitle":"Removing a planet raises gap complexity by +0.15, so the even spacing seen by Kepler is astrophysical.","key_machinery":"The load-bearing tool is the jackknife experiment applied to two samples, monitored through three metric definitions. Gap complexity, the central metric, is a single number between 0 (perfectly evenly spaced planets) and 1 (maximally irregular spacing in log-period), and the paper computes how this number changes when a planet is removed. The L24 observed sample provides the real-world test of removing exactly one planet at a time, while the SysSim synthetic system catalogs provide the test of removing however many planets a Kepler-like detection pipeline would miss, with full knowledge of the underlying true system. Mass partitioning is evaluated with masses drawn from a probabilistic mass-radius relation, and coplanarity is estimated from impact-parameter dispersion. The comparison between the synthetic-detected and observed gap-complexity distributions is what carries the astrophysical conclusion.","core_discovery":"The paper's central discovery is a directional bias: detection incompleteness does not manufacture the 'peas-in-a-pod' regularity seen in high-multiplicity Kepler systems; it works against it. In the observed L24 catalog, removing a middle planet from a four-plus planet system raises that system's gap complexity by a median of +0.15, while removing an edge planet changes it by less than 0.01. In synthetic SysSim systems, missing at least one middle planet raises the median system gap complexity by +0.16, versus +0.01 when only edge planets are missed. Mass partitioning shifts by less than 0.01 in both experiments, and the impact-parameter dispersion of transiting planets is similarly unaffected. Since the bias from missed planets is to make spacing look more irregular, the fact that observed systems are still more evenly spaced than the bias-applied synthetic population implies the regularity is intrinsic to the systems, not produced by Kepler's detection biases.","pith_inferences":["A direct aggregate test of the paper's logic would add every independently confirmed hidden planet back into the L24 systems and check whether median gap complexity falls by roughly the +0.15 amount; the paper identifies hidden planets in specific systems but does not run this population-level test.","If the upward gap-complexity bias is universal, dynamical-excitation and stability estimates that treat observed transiting planets as complete may systematically overstate how dynamically hot these systems are.","Because mass partitioning is insensitive to removed planets, studies of intra-system mass uniformity can proceed with incomplete catalogs, but this particular result inherits the assumptions of the probabilistic mass-radius relation used to assign masses.","Extending the same jackknife to period ratios or to recently discovered higher-multiplicity systems would test whether the +0.15 bias grows with the number of planets and whether it is uniform across orbital architectures."],"forward_implications":["A system with high gap complexity is a plausible hiding place for one or more undetected middle planets; a large gap at the edge does not carry the same signal.","Detection bias works against the peas-in-a-pod pattern, so the observed regular spacing is not an artifact of missed planets.","Mass homogeneity and orbital coplanarity measured from transiting samples are robust to moderate detection incompleteness.","Any synthetic model of planetary architectures that aspires to match Kepler must reproduce the observed low gap-complexity distribution; the current synthetic model overproduces irregular spacing.","The quantitative bias, roughly +0.15 gap-complexity units per missed middle planet, gives a concrete benchmark for judging how surprising a large intra-system gap is."],"supporting_citations":[{"why":"Supplies the homogeneous observed Kepler planet catalog used for the jackknife experiment and for comparison with the synthetic detected systems.","marker":"J. J. Lissauer et al. (2024, L24)"},{"why":"Introduces the SysSim modeling framework and the approximate Bayesian computation used to fit the underlying planet distribution to Kepler observations.","marker":"M. Y. He et al. (2019)"},{"why":"Provides the pregenerated maximum-AMD synthetic catalogs whose underlying and detected systems are compared here.","marker":"M. Y. He et al. (2020)"},{"why":"Defines the gap-complexity and mass-partitioning metrics that serve as the paper's architectural diagnostics.","marker":"G. J. Gilbert & D. C. Fabrycky (2020)"},{"why":"Supplies the probabilistic mass-radius relation used to assign planet masses for the mass-partitioning calculation.","marker":"J. Chen & D. Kipping (2017)"},{"why":"Motivates the interpretation that high gap complexity can flag the presence of undetected planets in a system.","marker":"M. Y. He & L. M. Weiss (2023)"},{"why":"Provides the false-positive statistics used to argue that the observed multi-planet catalog is essentially free of false positives.","marker":"J. J. Lissauer et al. (2014)"}],"fun_headline_variants":["Missed planets skew spacing, not sizes or coplanarity","Missing a mid planet hikes gap complexity by 0.15","Even Kepler spacing is real, not a detection artifact","Detection bias can't fake peas-in-a-pod patterns"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"That the observed even spacing is real and astrophysical presupposes that the SysSim synthetic population, including its detection model, is a faithful stand-in for the true Kepler planet population; an unrealistic period-spacing prior or detection pipeline in the model would make the comparison between synthetic and observed systems misleading.","fun_headline_variants_meta":{"raw":{"variants":["Missed planets skew spacing, not sizes or coplanarity","Missing a mid planet hikes gap complexity by 0.15","Even Kepler spacing is real, not a detection artifact","Detection bias can't fake peas-in-a-pod patterns"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000423,"raw_usage":{"total_tokens":2185,"prompt_tokens":972,"completion_tokens":1213,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":588,"completion_tokens_details":{"reasoning_tokens":1145}},"tokens_in":588,"tokens_out":1213,"duration_ms":10186,"temperature":1.0,"reasoning_tokens":1145,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:54:22.256954+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the experiment on systems where the missing planets are already known: take the L24 four-plus planet systems, add every independently confirmed transiting or non-transiting planet back into those systems, and recompute the median gap complexity. The paper's claim predicts a decrease comparable to the +0.15 middle-removal offset; if the median instead rises or stays flat, the asserted bias direction would be refuted.","supporting_citations":[],"review_version":1}