{"id":"fdd3ad42-9aa6-4e71-be71-08e9f3e7cafd","arxiv_id":"1908.02217","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":8,"one_line_summary":"Injection-recovery simulations show that quasi-periodic stellar activity biases retrieved small-planet radial-velocity amplitudes upward by up to 100% for active stars, and that periodogram searches recover the planet in under 5% of cases.","lead":"This paper simulated 80,000 radial-velocity datasets containing a small planet plus quasi-periodic stellar activity, then tested how well standard analysis tools recover the planet. It finds the planet's wobble amplitude can be overestimated by up to a factor of two for active stars, and periodogram searches alone recover the planet in fewer than 5% of cases.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Matched-kernel protocol: activity is both generated and fit with the same quasi-periodic GP kernel (Eq. 1), so the quoted bias magnitudes may not survive kernel misspecification.","rationale":"The reader's verdict is CONDITIONAL, with the weakest assumption being the matched-kernel fidelity of the quasi-periodic GP. I agree that this is the single most load-bearing concern. The paper is careful and internally consistent: the simulations are large, the sampling is realistic, and the recovery framework is standard. The bias estimates are likely reliable for the stated ideal model, and the paper explicitly acknowledges the working hypothesis. However, because the central purpose is to guide real RV follow-up and mass measurements, the quantitative transfer of the bias map depends on the generative model being representative of real activity. The proposed test directly targets this by changing only the activity generator while keeping the analysis pipeline fixed. The concern does not require changing the verdict, since the reader already conditioned acceptance on exactly this issue and on code availability; no additional adjustment is needed. I also note that the paper's own internal checks, such as varying the tau_AR prior and comparing against jitter-only fits, support the robustness of the Kb-bias results within the quasi-periodic family, which strengthens confidence that the issue is not an internal inconsistency but a scope limitation.","tokens_in":25923,"tokens_out":3889,"duration_ms":46910,"concrete_test":"Re-run the Case III and Case IV injection-recovery simulations (active star, short and long tau_AR, one and two seasons) with 5000 mock datasets per scenario, using the same MultiNest/GEORGE fitting pipeline and priors, but generating the activity from a different physical model, such as SOAP 2.0 starspot simulations or a Matern-periodic GP with the same h, Prot, tau_AR, and w. Compare the medians of Kb,ratio 50% and Kb,50%/sigma_Kb to Table 4. If the medians shift by more than ~0.3 in Kb ratio, or if the significance crosses the 2-sigma threshold, the headline bias magnitudes are model-dependent; if they match within uncertainties, the matched-kernel objection is mitigated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central bias map rests on a matched-kernel protocol: the stellar activity component is generated by drawing from the quasi-periodic GP kernel of Eq. (1) with GEORGE (Sect. 2.5), and the same kernel is used to model the activity in the retrieval (Sect. 3.1, where the authors state: 'Our working hypothesis is that the choice of a quasi-periodic kernel is justified'). Under this protocol, the model is never misspecified, so the reported bias magnitudes, e.g., Kb,ratio 50% ~ 2.1 and ~1.9 for the active-star, Prot~Porb, short-tau_AR case with one and two seasons (Table 4), are estimates of bias under an ideal assumption. Real stellar activity includes granulation, convection, differential rotation, and complex spot evolution that are not captured by the quasi-periodic GP; these components could either absorb more of the planetary signal or be more separable, changing the bias numbers. The paper's own robustness checks (Sect. 3.2.2) vary priors within the same kernel family and show that improving the tau_AR fit does not improve Kb recovery, but this does not test kernel misspecification. Thus the load-bearing assumption is that the quasi-periodic kernel is a faithful enough description of activity that the bias map transfers to real datasets.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents an extensive injection-recovery study of radial velocity (RV) time series containing a small planetary signal (K_b = 1 m/s) embedded in quasi-periodic stellar activity. The authors generate 80,000 mock datasets across 16 scenarios that vary activity amplitude (h = 3 and 15 m/s), rotation period, active-region lifetime (short vs. long tau_AR), the proximity of P_orb to P_rot, and one versus two observing seasons. They fit each dataset with a Keplerian plus a quasi-periodic Gaussian-process (GP) activity term using MultiNest and GEORGE, and quantify the recovered K_b, the significance K_b/sigma, the GP hyperparameters, and the fitted jitter. They also analyze a subset of datasets with GLS, BGLS, and FREDEC periodograms to measure completeness and reliability in recovering P_rot and P_orb. The main claims are that for K_b = 1 m/s the detection significance is generally below 2-sigma, that K_b is severely overestimated (by factors up to ~2) for active stars with P_orb ~ P_rot and short tau_AR, that the 68.3rd percentile of the K_b posterior is often closer to the injected value than the median, and that simple periodogram searches have very low completeness for the planet signal.","tokens_in":26242,"tokens_out":6529,"duration_ms":69680,"significance":"If the results hold, they provide a useful quantitative caution for the RV follow-up of small transiting planets, particularly in the context of TESS and PLATO targets. The study's strengths are its large simulation volume (5,000 datasets per scenario), the systematic coverage of activity regimes, the use of a realistic observing calendar, the explicit robustness checks in Appendix B and Sect. 3.2.2, and the comparison of three widely used periodogram tools. The paper also makes the simulated datasets available, which is a practical contribution. The main limitation is that the activity signal is both generated and fitted with the same quasi-periodic GP kernel, so the quoted bias magnitudes are matched-kernel estimates; the paper acknowledges this in Sect. 2.2 and Sect. 3.1, but the practical conclusions in Sect. 5 are stated more broadly. Subject to that caveat, the internal statistics are solid and the reported trends are clearly tabulated.","major_comments":[{"comment":"The headline statement in Case III, 'The semi-amplitude Kb is overestimated by 100% or more when tau_AR is close to P_rot, even with two seasons of data,' is not fully supported by Table 4. In the active-star, P_orb~P_rot, short-tau_AR row, K_b,ratio,50% is 2.14+0.5-0.4 for one season, but 1.9+0.5-0.4 for two seasons, so the median two-season overestimate is about 90%, not 100% or more. Please rephrase the claim to state the median overestimate (or identify the percentile that exceeds 100%).","section":"Sect. 3.2, Case III and Table 4"},{"comment":"The activity term is generated by drawing from the same quasi-periodic GP kernel (Eq. 1) used in the retrieval, so Table 4 measures bias under a matched-kernel, model-true protocol. The manuscript acknowledges in Sect. 2.2 that the quasi-periodic representation is 'not necessarily complete' and in Sect. 3.1 that the kernel choice is a working hypothesis, but the abstract and Sect. 5 draw practical conclusions about real surveys. Real activity includes granulation, convection, differential rotation, and complex spot evolution, all of which could change the bias magnitudes and the conclusions about GP effectiveness. I recommend either adding robustness simulations with a different activity generator (e.g., a spot-occultation model or a GP with a different kernel) or explicitly restricting the central claims to the exactly quasi-periodic case.","section":"Sect. 2.5 and Sect. 3.1"},{"comment":"Describing the 68.3rd percentile of the K_b posterior as a 'more accurate estimate' is problematic in the low-signal cases treated here, where the posterior is one-sided and this percentile is an upper limit rather than a point estimate. The comparison in Table 4 does not establish accuracy; a coverage statistic, such as the fraction of datasets for which the 68.3% credible interval contains the injected K_b, would be a more meaningful calibration. This issue affects one of the summary bullets in Sect. 5.","section":"Sect. 3.2, 'Upper limits as defined by the 68th percentile'"},{"comment":"The quantity K_b,50%/sigma_Kb^- is repeatedly called a 'detection significance,' but for one-sided posteriors that pile up near zero, the lower uncertainty sigma_Kb^- can be very small, making this ratio a poor proxy for significance relative to a null amplitude. The statement that detection significance 'stays below 2 sigma' should be justified or the quantity should be renamed, for example 'median-over-lower-uncertainty ratio.'","section":"Table 3 and Sect. 3.2"}],"minor_comments":[{"comment":"The text '1 0000 RV mock datasets' should read '10 000 RV mock datasets.'","section":"Sect. 4"},{"comment":"The line 'median sigma_jit, no GP/sigma_jit, with GP' is duplicated within each observing-season block; remove the repeated rows.","section":"Table 5"},{"comment":"In the low-activity, short-tau_AR, two-season GLS row, the entry '30.9.4%' should be '30.9%.'","section":"Table 8"},{"comment":"There is a typo in the introduction: 'Zeng et al. 017b' should be 'Zeng et al. 2017b.'","section":"Section 1"},{"comment":"The notation alternates between K_p (Sect. 2.1) and K_b (Sect. 2.5 onward) for the planetary semi-amplitude; please use a single symbol throughout.","section":"Sections 2.1 and 2.5"}],"recommendation":"major_revision","confidential_remarks":"The simulation campaign is well executed and the paper is within scope for MNRAS, but the matched-kernel design is the main barrier to the broader practical conclusions. I would like to see either a misspecification test or a clearly scoped set of claims, in addition to the correction of the Case III quantitative statement. The internal statistics are otherwise robust."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here is my read on Damasso et al. (arXiv:1908.02217). It is a careful, large-N simulation study that gives the community a quantitative map of a known qualitative worry: for a 1 m/s planet, quasi-periodic stellar activity can bias the recovered semi-amplitude, and in the worst case (active star, P_orb ~ P_rot, tau_AR ~ P_rot) the median Kb is overestimated by roughly a factor of two even with two seasons. The 68.3rd percentile of the Kb posterior is often closer to the injected value than the median; that is a practical, useful recommendation.\n\nWhat is actually new: the 16-scenario grid, 80,000 mock datasets, the explicit Kb bias as a function of activity level, period separation and tau_AR, and the periodogram completeness/reliability tables for GLS, BGLS and FREDEC. Execution is competent: 5,000 datasets per scenario, realistic epoch sampling from a real TNG calendar, clear definitions in Table 3, and the prior-robustness checks in Sect. 3.2.2 show that better tau_AR recovery does not improve Kb recovery. That is an honest, non-obvious result. The references are appropriate and build on the earlier injection-recovery work.\n\nThe main soft spot is not hidden but should shape how the numbers are used. Activity is generated and recovered with the same quasi-periodic GP kernel (Eq. 1), so this is a matched-filter experiment. The authors acknowledge this in Sect. 2.2 ('not necessarily complete representation') and Sect. 3.1, but they do not test a misspecified activity model. Real activity includes granulation, differential rotation and evolving spot geometries, so the bias magnitudes could shift. That is a scope condition, not a fatal flaw. Minor limitations: datasets are only 'available upon request' (no public repository/hash), the 0.5-day P_orb prior is transiting-planet specific, and the main grid fixes Kb=1 m/s with only a few exploratory 2-3 m/s cases.\n\nWho gets value: anyone planning RV follow-up of TESS/PLATO small planets, and anyone fitting RVs with GPs who needs to know what to trust. The paper deserves a serious referee. My asks would be a public data release and, if cheap, one injection test with a different activity model (e.g. a physical spot model or a different GP kernel) to see how the bias map moves. I would conditional-accept on data release, not because the central result is wrong, but because the paper's usefulness depends on replicability.","headline":"A competent, useful simulation study quantifying GP-activity bias on small-planet RVs; the matched-kernel caveat is real but scoped, not fatal.","tokens_in":26800,"tokens_out":4541,"would_cite":true,"duration_ms":46302,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"When a low-mass planet's orbital period nears the star's rotation period, Gaussian-process activity correction can overestimate the planet's Doppler amplitude by 100% or more, with detection significance below 2 sigma even after two…","keywords":["radial velocity","stellar activity","Gaussian process regression","exoplanet masses","injection-recovery simulations","quasi-periodic variability","periodogram analysis","low-mass planets"],"falsifier":"Re-run the Case III scenario (active star, $P_{\\rm orb}\\approx P_{\\rm rot}$, $\\tau_{\\rm AR}\\approx P_{\\rm rot}$, two seasons) with the injected activity drawn from a physical rotating-spot or granulation model instead of sampled from the same GP kernel. If the median $K_b$ ratio then falls well below 2, the claimed 100% overestimate is partly an artifact of the matched kernel; if it stays near 2, the bias persists outside the paper's model-fidelity assumption.","tokens_in":25725,"feed_emoji":"🪐","tokens_out":7092,"duration_ms":73227,"temperature":0.7,"pith_summary":"This paper asks how accurately astronomers can weigh small planets when stellar activity contaminates radial-velocity measurements. Through 80,000 mock datasets spanning 16 scenarios, it shows that the standard quasi-periodic Gaussian-process correction systematically biases the retrieved planetary semi-amplitude: for active stars with the planetary period close to the rotation period, $K_b$ is overestimated by 100% or more even with two observing seasons, and the detection significance stays below 2 sigma in nearly all tested cases. The paper also finds that the activity timescale is often poorly recovered from one season of data, and that fitting the activity better does not automatically improve the planet mass. The result matters because these are exactly the regimes targeted by follow-up mass measurements of small transiting planets found by current and planned transit surveys.","feed_headline":"Stellar activity can double a small planet's measured mass","feed_subtitle":"Mass estimates for small transiting planets are biased in the regimes that RV follow-up surveys actually target.","key_machinery":"The load-bearing object is the quasi-periodic Gaussian process covariance kernel $$K(t,t') = $h^{2}$ \\exp\\left[-\\frac{(t-t')^2}{2\\tau_{\\rm AR}^2} - \\frac{\\$sin^{2}$\\left(\\frac{\\pi(t-t')}{P_{\\rm rot}}\\right)}{$2w^{2}$}\\right] + \\sigma_{\\rm RV}^2(t)\\delta_{t,t'},$$ where $h$ is the activity amplitude, $P_{\\rm rot}$ the stellar rotation period, $w$ the periodic length scale, and $\\tau_{\\rm AR}$ the active-region evolutionary timescale. This kernel both generates the mock stellar activity, via random draws from the GP prior, and serves as the model used to fit it, so the experiment measures how well the recovery pipeline inverts the same stochastic process that created the data. Around that kernel the paper builds an injection-recovery Monte Carlo: 80,000 mock datasets in 16 scenarios, varying activity level, rotation period, activity timescale, period coincidence between planet and star, and one versus two observing seasons, analyzed with nested-sampling GP fits and three periodogram algorithms.","core_discovery":"The paper's central claim is that when stellar activity is modeled as a quasi-periodic Gaussian process, the standard GP regression procedure used to remove it does not reliably retrieve a 1 m/s planetary signal. For the most difficult case, an active star whose rotation period nearly equals the planet's orbital period, the posterior median of the semi-amplitude runs about 1.9--2.1 times the injected value when the activity evolves on the rotation timescale, even with two seasons; with long-lived active regions the overestimate drops to about 20% but remains present. The detection significance is about 1.5--1.6 sigma across active-star cases, formally a non-detection. For quiet stars the median is generally close to the true value, and the 68.3rd percentile of the $K_b$ posterior often gives a more accurate estimate than the median. The paper frames these numbers as expected biases for transit-follow-up mass measurements and as a warning that better constraints on the activity model do not translate into better planetary amplitudes.","pith_inferences":["Beyond the paper: because the activity model is given its best possible chance here, with the same kernel used for generation and fitting, real activity containing granulation, convection, or evolving spot geometries would likely make the planet-activity separation harder rather than easier, so the Case III overestimate may be a floor rather than a ceiling.","Beyond the paper: the same simulation machinery could be used to test whether adding photometry or spectroscopic activity indicators as extra GP inputs reduces the $K_b$ inflation; the paper sets up exactly such a controlled comparison without running it.","Beyond the paper: the strong $P_{\\rm orb}\\approx P_{\\rm rot}$ bias implies that population-level mass-radius studies should flag or down-weight planets whose orbital period is close to the stellar rotation period, a survey-level consequence the paper does not draw."],"forward_implications":["For active stars with $P_{\\rm orb}\\approx P_{\\rm rot}$ and $\\tau_{\\rm AR}\\approx P_{\\rm rot}$, fitted $K_b$ values will be roughly twice the true amplitude even with two seasons of data, so mass estimates in these systems are systematically high.","Across nearly all 16 scenarios, a 1 m/s planet is retrieved with significance below 2 sigma, meaning typical RV follow-up of small transiting planets will yield upper limits rather than detections.","In quiet-star cases with $P_{\\rm orb}$ well separated from $P_{\\rm rot}$ and two seasons of data, the 68.3rd percentile of the $K_b$ posterior is closer to the injected value than the median is, so posterior upper limits are more trustworthy than best-fit medians.","Adding a second season improves recovery of the activity timescale $\\tau_{\\rm AR}$ but does not cure the planet-amplitude bias; better activity modeling does not automatically produce better planet masses.","Blind periodogram searches recover the planetary period in fewer than about 5 percent of datasets, so periodogram peaks alone cannot establish the presence or amplitude of such planets."],"supporting_citations":[{"why":"Supplies the precedent that GP models can recover small-amplitude planetary signals from synthetic RVs and the K/N detectability threshold used to frame the challenge.","marker":"Dumusque et al. 2017"},{"why":"Provides the Gaussian-process library used both to draw random activity realizations and to compute the covariance fits, so removing it breaks the simulation.","marker":"Ambikasaran et al. 2014"},{"why":"Supplies the nested-sampling method used for the Monte Carlo parameter estimation.","marker":"Feroz et al. 2013"},{"why":"Provides the Python interface through which the nested-sampling estimation is run, forming part of the fitting pipeline.","marker":"Buchner et al. 2014"},{"why":"Defines the generalized Lomb-Scargle periodogram whose completeness and reliability are characterized in the blind-search analysis.","marker":"Zechmeister & Kürster 2009"},{"why":"Defines the Bayesian generalized Lomb-Scargle periodogram compared against the other search algorithms.","marker":"Mortier et al. 2015"},{"why":"Defines the multi-frequency decomposer whose false-positive behavior is compared with the other algorithms.","marker":"Baluev 2013"},{"why":"Earlier comparative periodogram study that this work extends to the quasi-periodic activity scenarios.","marker":"Pinamonti et al. 2017"},{"why":"Mass-radius relation used to map the injected $K_b=1$ m/s amplitude to planet masses and radii.","marker":"Weiss & Marcy 2014"},{"why":"Supports the assumption of circular orbits for small, close-in planets used in the simulated planetary signal.","marker":"Van Eylen et al. 2019"}],"fun_headline_variants":["Stellar activity can inflate small-planet masses by 2x","Active stars can double the mass of small exoplanets","Quasi-periodic activity can inflate small planet masses","Stellar noise can double small exoplanet mass estimates","Activity can mimic and double small planet masses"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole exercise assumes that a quasi-periodic Gaussian process kernel faithfully represents stellar activity, because the same kernel generates the mock signals and is then used to remove them; if real activity contains granulation, convection, or evolving spot geometries, the bias sizes could differ.","fun_headline_variants_meta":{"raw":{"variants":["Stellar activity can inflate small-planet masses by 2x","Active stars can double the mass of small exoplanets","Quasi-periodic activity can inflate small planet masses","Stellar noise can double small exoplanet mass estimates","Activity can mimic and double small planet masses"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000595,"raw_usage":{"total_tokens":2829,"prompt_tokens":1034,"completion_tokens":1795,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":650,"completion_tokens_details":{"reasoning_tokens":1714}},"tokens_in":650,"tokens_out":1795,"duration_ms":13355,"temperature":1.0,"reasoning_tokens":1714,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:51:23.410732+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the Case III scenario (active star, $P_{\\rm orb}\\approx P_{\\rm rot}$, $\\tau_{\\rm AR}\\approx P_{\\rm rot}$, two seasons) with the injected activity drawn from a physical rotating-spot or granulation model instead of sampled from the same GP kernel. If the median $K_b$ ratio then falls well below 2, the claimed 100% overestimate is partly an artifact of the matched kernel; if it stays near 2, the bias persists outside the paper's model-fidelity assumption.","supporting_citations":[{"cited_title":"W., O'Neil M., 2014","cited_arxiv_id":null,"evidence_quote":"Provides the Gaussian-process library used both to draw random activity realizations and to compute the covariance fits, so removing it breaks the simulation."},{"cited_title":"V., 2013, @doi [Astronomy and Computing] 10.1016/j.ascom.2013.11.003 , http://adsabs.harvard.edu/abs/2013A","cited_arxiv_id":null,"evidence_quote":"Defines the multi-frequency decomposer whose false-positive behavior is compared with the other algorithms."}],"review_version":1}