{"id":"a18b4d82-0cd3-4b53-bf73-fef82c2dfd1c","arxiv_id":"2509.07297","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"An HWO-like survey needs about 30 Earth-sized planets to recover a strong albedo step at the habitable zone, but 80-90 for weaker trends.","lead":"By simulating direct-imaging surveys with the Bioverse code, the authors estimate how many Earth-sized planets a mission like the Habitable Worlds Observatory must detect to recover trends in planetary albedo across the habitable zone. The results help mission planners decide whether a 25-exoEarth sample is enough to test theories of habitability or whether stronger telescope designs are needed.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"KDE plug-in Bayes factor makes the 25–30 EEC threshold for strong trends an artifact of the chosen bandwidth; Appendix B shows substantial sensitivity.","rationale":"The reader identified the KDE/in-sample Bayes factor as the weakest assumption; I agree and sharpen it to a specific, testable calibration issue. The paper itself, in Appendix B, acknowledges sensitivity to the kernel bandwidth, stating that reported values 'should be taken with a grain of salt.' This is an explicit limitation, but the main text still presents the 25–30 EEC number as a headline result with no uncertainty. A concrete numerical test can determine whether the KDE approximation is the culprit. Given that the reader’s verdict was already CONDITIONAL and my concern reinforces that conditionality without showing the argument is fundamentally broken, the appropriate verdict is unchanged (still CONDITIONAL). The concern does not warrant rejection because the underlying simulation and sensitivity analyses are transparent and the central qualitative conclusion—that weaker trends require many more EECs—is likely robust even if the exact thresholds shift.","tokens_in":20675,"tokens_out":6051,"duration_ms":81075,"concrete_test":"Recompute the power curves of Figures 4–5 using the exact generative distributions (available from the Bioverse forward model) instead of KDEs: e.g., estimate the joint density of (log S_eff, beta) with a very large Monte Carlo sample or a fine histogram, then calculate Bayes factors using these exact densities for both Trend and Null. If the 95% required N_EEC for ΔA=0.4 shifts by more than ~20% (outside roughly 24–42), the headline numbers are KDE-dependent and must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central sample-size claims (e.g., 30–35 EECs for 95% power at K>10 for ΔA=0.4) are computed from a Bayes factor that uses kernel density estimates (KDEs) of the forward-modeled survey distributions as likelihoods (Sec. 2.6, Eq. 6). The KDEs are trained on the same simulation output used to generate the test samples, making this an in-sample plug-in density-ratio test rather than a marginal likelihood. The paper’s statement that 'there is no parameter to marginalize over' overlooks both the KDE bandwidth (a smoothing parameter) and the fact that the trend model’s parameters (A0, ΔA) are fixed rather than marginalized. Appendix B shows that decreasing the bandwidth by 50% noticeably lowers the true positive rate, and the curve may asymptote below 100% power. Because the Bayes factor is a product of many density estimates, small calibration errors propagate. The reported required sample sizes therefore reflect the specific KDE choice, not an intrinsic property of the survey. This is the load-bearing step: if the KDE is miscalibrated, the headline conclusion that ~25–30 EECs suffice for the strongest trend loses its quantitative support.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper adapts the Bioverse statistical framework to simulate direct-imaging surveys for HWO-like architectures, injecting a toy step-function albedo trend (lower albedo inside the conservative Kopparapu HZ, A0=0.7 outside, ΔA up to 0.4). It forward-models the survey distribution of normalized contrast β versus instellation S_eff, then draws samples with a specified number of exoEarth candidates (EECs) and computes Bayes factors from KDE likelihoods (Eq. 6) to estimate true and false positive rates. The main findings are that a strong ΔA=0.4 trend requires roughly 30-35 EECs for 95% power at K>10, weaker trends (ΔA=0.3) require 80-90 EECs, and the Decadal Survey's 25-EEC target suffices only for the strongest trend. Appendices examine sensitivity to telescope design, KDE bandwidth, and simulation count.","tokens_in":21010,"tokens_out":7462,"duration_ms":92931,"significance":"If the quantitative estimates were robust, the paper would provide a useful, falsifiable framework for HWO trade studies and a caution that the 25-EEC target is marginal for population-level albedo trends. Strengths include the explicit and publicly available code fork, the Monte Carlo power-calculation approach, the convergence analysis in Appendix C, and the unusually honest sensitivity tests in Appendices A and B that expose where the numbers are fragile. The qualitative conclusion that weaker trends require much larger samples is likely robust; however, the headline sample-size numbers are not yet established because the central statistic depends on KDE bandwidth choices that are not fully quantified.","major_comments":[{"comment":"The Bayes factor is computed as a product of KDE densities, with no marginalization over the KDE bandwidth or over the trend parameters (A0, ΔA, HZ boundaries). The statement 'there is no parameter to marginalize over' is therefore misleading. Appendix B and Figure 10 show that a 50% reduction in bandwidth noticeably lowers the true positive rate and may asymptote below 100% power, directly affecting the required N_EEC values in §3.2–§3.3. Please quantify this sensitivity: report required N_EEC for each bandwidth in Figure 10, validate the KDE with held-out samples or cross-validation, and give confidence intervals on the headline sample sizes.","section":"§2.6, Eq. (6); Appendix B"},{"comment":"The false positive rate is reported to increase with trend strength and sample size in parts of the parameter space, and the authors state that 'the exact reason for the increase in the variance of the Bayes factor distribution remains unclear.' A hypothesis test whose false positive rate grows with sample size is not well calibrated; since the FPR is used to argue that the survey is power-limited rather than false-positive-limited, this behavior must be explained. Please diagnose the cause (e.g., heavy tails of the log-Bayes-factor distribution, KDE edge effects) and report numerical FPR values with uncertainties, or restrict the conclusions to the calibrated regime.","section":"§3.3, Figure 6"},{"comment":"The simulations inject a step-function albedo dip at the known Kopparapu HZ boundaries with fixed A0 and ΔA. The analysis measures the power to detect that known-shape dip, not the ability to infer the HZ boundary location or the trend shape. Claims of 'empirically constrain the Habitable Zone' therefore overstate what is computed. Please clarify this distinction in the abstract and discussion, or add a version of the test in which the HZ boundaries (and ideally ΔA) are free parameters and are marginalized over.","section":"Title/Abstract; §2.2; §3.3"},{"comment":"The reported power is computed under the assumption that the Bioverse forward model is the true data-generating process. Any mismatch between the assumed planet distribution function, exozodi level, noise model, and reality will degrade the real-world detection power. The paper acknowledges this qualitatively in §4 but does not quantify it. At least for the headline number (30–35 EECs for ΔA=0.4), please state explicitly that this is an upper bound under the adopted model and provide a sensitivity test on the most uncertain inputs (e.g., exozodi level, η⊕, PDF shape).","section":"§2.5; §4"}],"minor_comments":[{"comment":"Typo: 'T rends' should be 'Trends'.","section":"Title"},{"comment":"Duplicate word: 'would would' should be 'would'.","section":"§4.3"},{"comment":"The text uses 10,000 simulations for the N_EEC grid and 1,000 for the ΔA–N_EEC grid. Please state this clearly in both sections, since the FPR contours in Figure 6 may be noisy at 1,000 simulations.","section":"§3.1 vs §3.3"},{"comment":"The description of the bottom-left corner is confusing; it is unclear whether high FPR occurs at low ΔA/low N or low ΔA/high N. Annotate the figure or revise the text to make the behavior unambiguous.","section":"Figure 6"},{"comment":"The 'insensitivity' claim is tested only for telescope diameter and minimum contrast. Exozodi level, which is a known major determinant of yield and detectability, is not varied in the survey-distribution comparison; please note this limitation explicitly.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for astro-ph.EP and will be of interest to the mission-design community. The central issue is statistical: the main quantitative claims depend on the KDE-based Bayes factor, and the authors' own Appendix B shows non-negligible sensitivity to kernel bandwidth. I would be willing to review a revision that calibrates the density estimator and quantifies the effect on required sample sizes. The unexplained FPR behavior in Figure 6 also needs attention. The title overstates the boundary-constraining claim and should be softened or supported by an analysis that marginalizes over HZ boundaries."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this for the Bioverse-HWO application, not for the precise thresholds. The paper extends Bioverse with HPIC target prioritization, an exposure time calculator, and phase-angle sampling, then runs a Monte Carlo power analysis for recovering an albedo step at the HZ. The main result—25–30 exoEarths recover a strong step (ΔA=0.4), 80–90 for a weaker one (ΔA=0.3)—is plausible and the qualitative conclusion (25 EECs only buys the strongest trend) is likely robust. That's a useful input to mission trade studies.\n\nThe novel parts are the new Bioverse modules and the specific application; they look properly implemented. Appendix A's finding that the survey distribution barely changes with telescope design is a nice result, and Appendix C's convergence check on the number of simulations is the kind of thing most papers skip.\n\nThe soft spot is the statistical core. The Bayes factor is computed by plugging KDEs of the forward-modeled distributions into Eq. 6, a product of density estimates. The paper says 'there is no parameter to marginalize over,' but the KDE bandwidth is a parameter, and the trend parameters are fixed not marginalized. More importantly, small calibration errors in the density estimates multiply across the product. Appendix B shows halving the bandwidth lowers the true positive rate substantially and the power curve may asymptote below 100%. The authors acknowledge this and say to take the numbers with a grain of salt, but the abstract and conclusions state the thresholds as if they were solid. A mission planner could easily over-read them.\n\nThe in-sample nature of the test is a related but milder concern: the KDEs are trained on the same simulator that generates the data, so the power calculation is conditional on the simulator being right. That's standard for a forward-model power study, but it means the absolute numbers are upper bounds if the real planet population differs. The paper flags the sensitivity to occurrence rates and exozodi in the text.\n\nOverall, this is a solid, honest simulation study with a clear result and good sensitivity checks. The qualitative message survives the KDE concern; the exact sample sizes do not. That combination merits a serious referee: the paper should be sent out, and the referee should ask for error bars or a robustness range on the headline numbers and a more careful statement about the KDE's role.","headline":"A careful simulation with a useful qualitative answer, but the headline numbers rest on a KDE that the paper itself admits is sensitive—worth refereeing with requests for robustness bounds.","tokens_in":21477,"tokens_out":3273,"would_cite":true,"duration_ms":40506,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A space-based direct imaging survey with 25-30 habitable-zone exoEarths can recover only the strongest albedo step, while weaker trends need 80-90.","keywords":["direct imaging","habitable zone","exoEarth yield","albedo trends","Bayes factor","statistical power","comparative planetology","Habitable Worlds Observatory"],"falsifier":"Repeat the sample-size analysis with a different KDE bandwidth or a different occurrence-rate model; if the number of exoEarths needed for 95% power shifts by more than a factor of two, the headline numbers are not robust. Alternatively, apply the same Bayes-factor test to the real normalized-contrast data from the first ~30-exoEarth survey; if the trend is not recovered at K>10 in at least ~90% of repeated realizations, the claim fails.","tokens_in":20607,"feed_emoji":"🔭","tokens_out":7336,"duration_ms":72291,"temperature":0.7,"pith_summary":"This paper asks whether a future space-based direct imaging mission such as the Habitable Worlds Observatory can test the classical prediction that rocky planets inside the habitable zone have lower albedos than planets outside it. Simulating an 8-meter coronagraphic survey with injected step-function trends in albedo, the authors find that the strongest plausible trend (albedo dropping from 0.7 outside to 0.3 inside the habitable zone) is recoverable with high confidence from about 25-30 exoEarth candidates (Earth-sized planets in the habitable zone). Weaker trends, with a drop of 0.3, require roughly 80-90 exoEarths, far more than the Decadal Survey's 25-exoEarth target. The result matters because it gives mission designers a quantitative science metric - the number of exoEarths needed to empirically constrain the habitable zone - and suggests that population-level albedo science sits on the edge of feasibility for near-term flagship designs.","feed_headline":"25 exoEarths buy only the strongest habitable-zone albedo trend","feed_subtitle":"Weaker albedo signatures would need 80-90 planets, beyond the Decadal Survey's 25-planet target.","key_machinery":"The central object is beta, a normalized planet-star contrast that is the closest directly observable proxy for albedo when radius, phase, and albedo are degenerate: beta = A_g (R_p/1AU)^2 Phi(alpha). The paper injects a step-function albedo model A_g = A0 - Delta A inside the habitable zone and A0 outside (for rocky planets R_p < 1.4 R_Earth), forward-models detectability with an exposure-time calculator, builds smooth survey distributions with kernel density estimates, and then uses the Bayes factor from Equation 6, the product of per-point KDE likelihoods, to measure how often random N_EEC samples separate trend from null at K=3 and K=10 thresholds.","core_discovery":"The paper's central claim is that statistical power to detect an albedo-instellation trend grows sharply with the strength of the trend, not just with sample size. Using forward-modeled distributions of the directly observable quantity beta = C(L/Lsun)^{-1}/S_eff (a normalized contrast proportional to geometric albedo times squared radius times phase function), the authors draw random samples of N_EEC exoEarths and compute Bayes factors between a step-function trend model and a no-trend null model via kernel density estimates. For a strong trend (Delta A = 0.4), a 95% true-positive rate at Bayes factor K>10 needs about 30-35 exoEarths, close to the Decadal Survey target. For Delta A = 0.3, t","pith_inferences":["The reported sample sizes are likely optimistic: they treat the synthetic forward model as the true data-generating process, and Appendix B shows the true-positive rate is sensitive to KDE bandwidth; a real survey with different occurrence rates, exozodi, or phase-function behavior would probably need more exoEarths.","The same Bayes-factor machinery could be applied to other HZ tests the authors list (water-vapor fraction, CO2 dependence on instellation), turning each into a sample-size requirement before the mission is built.","If high-resolution spectroscopy becomes cheap enough, direct albedo estimates from retrievals would break the beta degeneracy; the paper leaves open whether the reduced sample size of well-characterized planets would outweigh the gain, which a follow-up simulation could settle.","The false-positive rate increases with sample size and trend strength in a way the authors do not fully explain; if that variance growth is real, future significance thresholds may need to be calibrated per N_EEC rather than fixed at K=10."],"forward_implications":["If correct, the Decadal Survey's 25-exoEarth requirement is sufficient only for the strongest plausible albedo step; weaker but arguably more realistic trends would demand a larger telescope or a longer survey.","Mission trade studies can use 'exoEarths needed to recover a population trend' as a design metric alongside raw EEC yield.","Because the survey distribution is insensitive to 6-10 m telescope diameter and coronagraph depth, the key decision for this science case is yield, not the exact shape of the detection bias.","Reducing measurement noise (contrast and semi-major axis uncertainty) by 10x improves statistical power only slightly, so chasing precision per target is a poor substitute for sample size.","A survey that only characterizes habitable-zone planets would lose the comparison sample of rocky planets outside the HZ, which is essential for trend detection."],"supporting_citations":[{"why":"Sets the 25-exoEarth goal for the Astro2020-recommended flagship that this paper tests.","marker":"National Academies of Sciences & Medicine 2021"},{"why":"Defines the classical habitable-zone climate theory whose albedo prediction motivates the step-function trend.","marker":"Kasting et al. 1993"},{"why":"Provides the habitable-zone boundaries used to place the step in the injected albedo model.","marker":"Kopparapu et al. 2013"},{"why":"Supplies the statistical comparative-planetology simulator this paper adapts for direct imaging.","marker":"Bixel & Apai 2021"},{"why":"Provides the exposure-time and coronagraph detection model used for target prioritization and survey simulation.","marker":"Stark et al. 2014"},{"why":"Supplies the median exozodiacal light level (3 zodi) adopted in the exposure-time calculation.","marker":"Ertel et al. 2020"},{"why":"Defines the Bayes-factor thresholds (K=3 and K=10) used to call trend detections.","marker":"Jeffreys 1998"}],"fun_headline_variants":["Strong albedo trends need 25 exoEarths; weak ones need 80+","Detecting habitable-zone albedo trends demands strong signals","HWO's 25-planet target only catches bold albedo trends","Albedo trend detection: strength matters more than sample size","25 exoEarths barely enough for clearest albedo signal"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The entire statistical-power calculation assumes that the simulated survey's forward-modeled distribution, including the injected step-function albedo, the assumed occurrence rates, exozodi level, noise, and KDE likelihood, is the true distribution a real survey would produce; if the real planetary population or measurement behavior differs, the required sample sizes will be larger.","fun_headline_variants_meta":{"raw":{"variants":["Strong albedo trends need 25 exoEarths; weak ones need 80+","Detecting habitable-zone albedo trends demands strong signals","HWO's 25-planet target only catches bold albedo trends","Albedo trend detection: strength matters more than sample size","25 exoEarths barely enough for clearest albedo signal"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000366,"raw_usage":{"total_tokens":1851,"prompt_tokens":836,"completion_tokens":1015,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":580,"completion_tokens_details":{"reasoning_tokens":920}},"tokens_in":580,"tokens_out":1015,"duration_ms":9176,"temperature":1.0,"reasoning_tokens":920,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T22:27:15.415722+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the sample-size analysis with a different KDE bandwidth or a different occurrence-rate model; if the number of exoEarths needed for 95% power shifts by more than a factor of two, the headline numbers are not robust. Alternatively, apply the same Bayes-factor test to the real normalized-contrast data from the first ~30-exoEarth survey; if the trend is not recovered at K>10 in at least ~90% of repeated realizations, the claim fails.","supporting_citations":[{"cited_title":"C., Roberge, A., Mandell, A., & Robinson, T","cited_arxiv_id":null,"evidence_quote":"Provides the exposure-time and coronagraph detection model used for target prioritization and survey simulation."},{"cited_title":"2020, The Astronomical Journal, 159, 177","cited_arxiv_id":null,"evidence_quote":"Supplies the median exozodiacal light level (3 zodi) adopted in the exposure-time calculation."},{"cited_title":"1998, The theory of probability (OuP Oxford)","cited_arxiv_id":null,"evidence_quote":"Defines the Bayes-factor thresholds (K=3 and K=10) used to call trend detections."}],"review_version":1}