{"id":"49564854-e9a8-43a0-a705-bb89cff8f7bb","arxiv_id":"2411.12723","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A homogeneous reanalysis of HARPS RVs for 87 small exoplanets provides RV amplitudes under 12 models and shows that eccentricity priors and GP activity modeling affect derived masses.","lead":"This paper reanalyzes public HARPS radial velocity data for 87 small exoplanets with 12 standardized models to produce a homogeneous set of planet mass measurements. The results give the community a consistent catalog and show that choices about orbital eccentricity and activity modeling can change derived masses.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No injection-recovery or independent validation of the GP activity model: the adopted K values rest on an unvalidated separation of activity from planet signals, and the best-model selection in §6.4 is an indirect heuristic.","rationale":"The reader identified the same weakest assumption: the GP activity model must correctly separate stellar from planetary signals. My stress test agrees and sharpens it: the adopted K values are selected by a heuristic (Section 6.4) that first uses AIC on RV-only models and then substitutes a 3D GP beta model without direct comparison, and the paper contains no injection-recovery or independent cross-validation of the final models. The three failed systems demonstrate that the activity model is not universally safe. This does not destroy the paper's comparative findings about model choices, and the release of all 12 model results is valuable, but the specific claim that the catalog is ready for demographics requires that the adopted K values be unbiased. That bias can only be tested by injecting known signals into the real data and recovering them. I therefore agree with the reader's conditional verdict rather than moving to accept or reject: the scientific approach is sound, but the central demographic claim needs the additional validation before the catalog is used for population studies.","tokens_in":31769,"tokens_out":6977,"duration_ms":72596,"concrete_test":"Run an injection-recovery test on a subset of about 10 stars spanning the activity and stellar-type range of the sample. Inject synthetic two-planet signals with known K values (e.g., 1, 3, and 10 m/s) and periods drawn from the sample period distribution into the real HARPS RVs and activity indicators, then run all 12 Pyaneti models exactly as in the paper. Check whether the posterior median K of the adopted best model (n or e) recovers the injected K within the quoted 1σ uncertainty for at least 90% of injections. If the recovery is biased for active stars or for K values near the GP amplitude, the catalog's demographic-ready claim is not supported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that the paper provides a homogeneous catalog of RV amplitudes suitable for demographics studies. For this to hold, the K value of each planet must be an unbiased estimate of the true amplitude. The weakest point is the stellar-activity separation: the quasi-periodic GP kernels in Pyaneti, with P_GP priors set from the Mamajek & Hillenbrand (2008) rotation relation, are assumed to correctly isolate planetary signals for all 44 stars, but no injection-recovery test or independent validation is presented. The paper itself shows the fragility: TOI-269, TOI-4399, and HD 3167 could not be modelled well (Section 6), so the GP framework is not universally reliable. The adopted catalog values are even less protected because of the best-model selection rule in Section 6.4: the lowest-AIC model is chosen from the RV-only models (a–f); if it is the 1D GP, the authors substitute the 3D GP beta model (n) without performing a model comparison on that substitution, and if it is a no-GP model they substitute the no-GP beta model (e). Models n and e are therefore adopted for almost all planets in Table B.1 without a direct test that they recover known K. The comparison to NASA Archive values in Fig. 5 is not validation: those values are heterogeneous, use different data sets, and are themselves not ground truth. As a result, the headline finding that eccentricity prior and GP choice change K is robust, but the absolute K values intended for demographics are not demonstrated to be unbiased.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper re-analyzes public HARPS radial velocities for 87 small exoplanets (R < 4 R⊕; the abstract also says 85) using the Pyaneti toolkit. For each target it runs 12 models that vary the Gaussian Process dimension (none, 1D, 2D, 3D), the eccentricity treatment (circular, uniform prior, beta prior), and the inclusion of long-term trends. It compares models with AIC/BIC, defines a 'best' model via a hybrid rule, and provides tables of K, eccentricity, period, and m sin i for the adopted models, with the full model grid available online. The authors report that the eccentricity prior can change K by factors up to ~3 in some cases, that the 1D GP gives systematically different K from the 2D/3D GPs, and that long-term trends have a small average effect. They present this as a homogeneous K catalog for demographics studies and give best-practice recommendations.","tokens_in":32119,"tokens_out":15453,"duration_ms":135285,"significance":"If the catalog K values are unbiased, this would be a valuable resource: the first large homogeneous small-planet RV-amplitude catalog from HARPS, directly useful for mass-radius demographics. The paper's strengths are its transparent MCMC setup, explicit prior tables, a 12-model grid applied uniformly, public data products, and honest caveats about archival data. The qualitative conclusions about model-choice sensitivity (eccentricity prior, GP dimension, trend inclusion) appear robust because they are based on differences between models applied to the same data. However, the absolute K values intended for demographics use rest on an unvalidated GP activity-separation assumption, and the sample definition has internal inconsistencies that must be resolved before the catalog can be used as stated.","major_comments":[{"comment":"The final adopted K values are not the AIC-selected models. The rule in §6.4 selects the lowest-AIC model among the RV-only models (a–f), then substitutes model n (3D GP + beta eccentricity) whenever the 1D GP wins and model e (no GP + beta eccentricity) whenever a no-GP model wins. Models e and n are never directly compared with the alternatives on the same likelihood, and model n uses FWHM/BIS as additional data dimensions so AIC is not comparable. No injection-recovery or cross-validation is presented to show that this substitution recovers unbiased K. The comparison to NASA Archive values in Fig. 5 is not a validation because those values are heterogeneous, use different data, and are not ground truth. Since Table B.1 is the primary catalog deliverable, this selection rule needs an explicit validation test, or the AIC-selected model should be listed as primary and the substituted model clearly labeled as a secondary choice.","section":"§6.4, Table B.1"},{"comment":"The GP activity separation is assumed rather than tested. The quasi-periodic kernel in Eq. (1) is applied to all 44 stars with P_GP priors derived from Mamajek & Hillenbrand (2008), but the manuscript contains no injection-recovery tests, no comparison with photometric rotation periods, and no leave-one-out checks to show that the GP does not absorb part of the planet signal or leave activity unmodelled. Section 6 itself shows the framework is fragile: TOI-269, TOI-4399, and HD 3167 could not be modelled well. For a catalog intended for demographic studies, at least a representative subsample of stars should be tested by injecting planetary signals of known K and comparing recovered K; otherwise the absolute K values in Table B.1 are not demonstrably unbiased.","section":"§5, §6"},{"comment":"The model-comparison counts are internally inconsistent. The text states that the lowest-AIC RV-only model is f for 78 planets and g for 11, which sums to 89, exceeding the stated sample of 87 small planets. The 3D GP counts are 80 (k) + 18 (n) + 15 (m, mislabelled as n in the text) = 113, equal to the total number of planets orbiting the target stars including the 26 non-small planets that §2 says are excluded from the model comparison. The counts must be reconciled with the actual sample and the exclusion statement clarified, because these counts describe the basis for the best-model selection.","section":"§6.4"},{"comment":"The sample definition is not reproducible from the tables. Section 2 says the final sample is 87 small planets orbiting 44 stars, but Table C.1 lists 49 stars and omits TOI-4399 (discussed in §6) while including stars with no entry in Table B.1 (for example, HD 18599, HD 15337, HIP 94235). Furthermore, §6 states that removing three targets leaves 83 small planets, which is only consistent if those three stars contribute four small planets total; this is not stated. The tables need to be made consistent and the exact per-star planet counts presented.","section":"§2, §6, Table C.1, Table B.1"}],"minor_comments":[{"comment":"The phrase '15 the uniform eccentric model (n)' should refer to model m, not n; model n is the 3D GP with beta-distributed eccentricity.","section":"§6.4"},{"comment":"The worked example alternates between 'TIC 9870809' and 'TIC 98720809'; the latter matches Table C.1 and should be used consistently.","section":"§6.5"},{"comment":"The λ_e prior is listed as U[1,160] for all stars, but the text says the maximum λ_e is twice the stellar-type-dependent P_GP maximum (20–60 days depending on temperature). Clarify whether the prior is global or per-star, and reconcile the numbers.","section":"Table 2, §5"},{"comment":"The statement that the 3D GP model 'will always have a lower value of AIC compared to the 2D GP case because it has more data points' is incorrect and contradicts the same paragraph's caveat that AIC cannot be compared across different data sets; AIC depends on the likelihood, not simply on the number of points.","section":"§6.4"},{"comment":"The number of small planets is given as 85 in one version of the abstract and 87 in the body; please harmonize all sample-size statements.","section":"Abstract, §2"},{"comment":"The manual RV cuts for TIC 173103335, TIC 220479565, TIC 260004324, and TIC 56815340 are described only in the appendix; please state in the main text how many points were removed per target and confirm that the exact cut thresholds do not affect the final K values.","section":"Appendix A.1, §3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript has the potential to become a standard reference for homogeneous RV masses. I am not recommending rejection because the qualitative sensitivity findings are well supported and the full model grid is transparent. The main reservations are the unvalidated GP activity separation and the sample/count inconsistencies; both are addressable with additional tests and careful table revisions. The editor may also wish to ask whether an independent comparison with another fitting code (e.g., PyORBIT) could rule out Pyaneti-specific effects, given that a coauthor is the primary developer of Pyaneti."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: this is a useful, carefully executed homogeneous reanalysis of archival HARPS RVs for 87 small planets, and the 12-model comparison is the real contribution. It deserves refereeing, but the adopted \"best model\" catalog values rest on a selection rule that is more heuristic than the rest of the paper, and the GP activity separation is not independently validated.\n\nWhat is new: previous homogeneous analyses covered far fewer systems. Running the same 12 models (GP dimension, eccentricity treatment, long-term trend) on every system and publishing K for all models is genuinely useful. The eccentricity-prior result is the most robust finding: a uniform eccentricity prior inflates K and produces implausibly high eccentricities, while a beta prior brings K back in line with circular fits. The GP-dimension comparison, with the 1D GP as the outlier, is also worth knowing. The paper is transparent about manual cuts and caveats, and the recommendations at the end are sensible.\n\nSoft spots: the stress-test note is on target. The adopted catalog uses model n (3D GP + beta eccentricity) or model e (no GP + beta) for almost all planets, chosen by a two-step rule: lowest AIC among RV-only models, then substitution without direct model comparison on the substituted model. So the headline \"best\" K values are not the ones that won the formal comparison. This is a defensible choice, but it deserves a clearer justification and a robustness check. More importantly, there is no injection-recovery test or independent validation that the quasi-periodic GP, with P_GP priors from the Mamajek & Hillenbrand relation, actually separates activity from planets in these 44 systems. The paper's own three failures (TOI-269, TOI-4399, HD 3167) show the GP is not universally reliable. The comparison to NASA Archive values is consistency checking, not validation, since those values are heterogeneous. None of this sinks the paper, but it does mean the catalog should be used with the model spread in mind, not just the adopted column.\n\nMinor issues: abstract says 85 planets while the text says 87 and then 83 after exclusions; Fig. 5 caption/labels appear mismatched; and the machine-readable Table B.2 needs to be released with the code for the catalog to be fully usable.\n\nThis is for exoplanet demographics modelers and RV observers planning surveys, not for people seeking a new method. A serious referee can fix the consistency issues and push for validation. Worth sending out.","headline":"A useful homogeneous HARPS reanalysis with a solid 12-model comparison, but the adopted best-model K values rest on an unvalidated GP activity separation and a heuristic selection rule.","tokens_in":32618,"tokens_out":2446,"would_cite":true,"duration_ms":24281,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A homogeneous reanalysis of 6428 archival HARPS radial velocities for 87 small exoplanets shows that the choice of eccentricity prior and of Gaussian-process activity model can substantially change the derived radial-velocity amplitude…","keywords":["exoplanet masses","radial velocities","HARPS","stellar activity","Gaussian processes","orbital eccentricity","mass-radius relation","homogeneous analysis"],"falsifier":"Compare the homogeneous RV amplitudes for a subset of the 87 planets against independent mass constraints from transit-timing variations or joint RV-plus-photometry fits: if the differences trace the presence of the GP or the eccentricity prior, rather than random scatter, the separation assumption is falsified. A sharper test is an injection-recovery experiment on the same HARPS time series with synthetic planets of known amplitude, checking whether each of the twelve models recovers the injected $K$ without bias.","tokens_in":31583,"feed_emoji":"🪐","tokens_out":7126,"duration_ms":61654,"temperature":0.7,"pith_summary":"This paper argues that the masses of small exoplanets are not as settled as the literature suggests, because nearly every published mass comes from a different modelling recipe. To make the comparison fair, the authors re-fit all publicly available HARPS radial velocities for 87 small planets around 44 stars—6428 measurements in total—with one pipeline and 12 explicitly varied models. They find that modelling the orbit as eccentric with a wide uniform prior can push the RV amplitude up, in some cases by a factor of about three, while a physically motivated beta-distribution prior on eccentricity keeps results close to the circular-orbit case. They also find that adding a Gaussian process to absorb stellar activity changes the amplitude, and that the most consistent results come from multi-dimensional GPs that also fit activity indicators rather than a GP on the RVs alone. If this is right, the community gets the first large homogeneous small-planet RV amplitude catalogue, and demographics conclusions drawn from heterogeneous mass lists will need to be re-examined.","feed_headline":"Model choice shifts masses of 87 small exoplanets","feed_subtitle":"A uniform reanalysis shows eccentricity priors and activity models can shift RV amplitudes by up to a factor of three.","key_machinery":"The argument is carried by a deliberately rigid comparison grid: every system is fitted with the same twelve models, so model choice is the only free variable. The central components are the quasi-periodic Gaussian Process kernel, which models stellar activity as a periodic, evolving signal with a period prior set from the stellar rotation period; the $\\sqrt{e}\\sin\\omega_*$ and $\\sqrt{e}\\cos\\omega_*$ parameterisation of eccentricity, which avoids truncation at zero; and the $\\beta$-distribution prior on eccentricity derived from transit populations, which prevents the unphysically large eccentricities that a uniform prior allows. The grid spans no-GP, 1D, 2D and 3D GPs, circular/eccentric/$\\beta$-distribution orbits, and no/linear/quadratic long-term trends, yielding the $K$ comparisons that ground every conclusion.","core_discovery":"The central claim is that, for the same archival data, the extracted RV amplitude $K$—and hence the planet mass $m\\sin i$—depends systematically on how the fit is configured. Running twelve models that vary the treatment of eccentricity (fixed circular, uniform prior, $\\beta$-distribution prior), the addition of long-term trends, and the dimensionality of a quasi-periodic Gaussian process used to mitigate stellar activity, the authors show that a uniform prior on eccentricity produces inflated, spuriously high eccentricities and higher $K$ values, whereas a $\\beta$-distribution prior yields amplitudes consistent with circular fits. Adding a 1D GP to the RVs alone gives the least consistent amplitudes, while 2D and 3D GPs that jointly fit the FWHM and bisector-span activity indicators behave more consistently. Long-term linear or quadratic trends have little effect. The authors release the $K$ amplitude for every planet under all twelve models and define a 'best' model per target, and they recommend against unconstrained eccentric orbits and in favour of activity-anchored multi-dimensional GPs for large surveys.","pith_inferences":["If the eccentricity-prior effect is general, then published mass-radius relations assembled from heterogeneous eccentric fits may be biased toward higher masses for low-mass planets; a test is to re-fit a literature sample with a beta-distribution prior and compare the mass-radius scatter.","The GP period priors derived from a single activity-rotation relation are a likely weak point: for stars where measured rotation periods are available, replacing that prior with direct rotation-period measurements should reduce activity-signal leakage, and this can be tested on the same data.","The same twelve-model grid could be applied to a control sample of planets with masses independently measured by transit-timing variations; agreement would validate the homogeneous catalogue, while disagreement would localise which model component is biased.","Because the paper uses only HARPS data, the catalogue cannot test instrument systematics; combining with data from other precision spectrographs for a subset would show whether the homogeneous results are instrument-independent."],"forward_implications":["The public catalogue of RV amplitudes under all twelve models lets other groups test how their own modelling choices move planet masses, without redoing the data reduction.","Demographics studies that mix masses from different pipelines carry a model-dependent scatter; a homogeneous sample should tighten the observed mass-radius relation if the systematics are real.","The recommendation to use a beta-distribution prior on eccentricity, or to fix circular orbits, implies that many published eccentric single-planet fits for small planets may be fitting noise.","Multi-dimensional GPs fitted to RVs plus activity indicators should become the default for active stars, since the 1D GP is the least consistent with the other models.","The three systems where no model produced a good fit (TOI-269, TOI-4399, HD 3167) show that some published masses rest on data or modelling beyond what a homogeneous pipeline can reproduce."],"supporting_citations":[{"why":"Supplies the fitting pipeline and MCMC implementation used for all twelve models.","marker":"Barragán et al. 2019b"},{"why":"Provides the multi-dimensional Gaussian process framework with the quasi-periodic kernel used to model stellar activity.","marker":"Barragán et al. 2022"},{"why":"Introduces the multi-dimensional GP approach linking RVs and activity indicators that the 2D and 3D models rely on.","marker":"Rajpaul et al. 2015"},{"why":"Provides the empirical beta distribution on eccentricity used as the informative eccentricity prior.","marker":"Van Eylen et al. 2019"},{"why":"Supplies the rotation-period relation that sets the GP period priors for all stars.","marker":"Mamajek & Hillenbrand 2008"},{"why":"Used to explain the spurious high eccentricities produced by uniform eccentricity priors fitted to outliers.","marker":"Hara et al. 2019"},{"why":"The curated HARPS radial-velocity catalogue used to match target names and assemble the 6428 observations.","marker":"Barbieri 2023"},{"why":"The ensemble sampler underlying the MCMC posterior sampling.","marker":"Foreman-Mackey et al. 2017"}],"fun_headline_variants":["Eccentricity priors shift exoplanet masses by 3x","How model choices skew small planet masses","Uniform reanalysis of 85 exoplanet masses","Activity models alter RV amplitudes up to 3x","HARPS reanalysis: priors matter for planet masses"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The catalogue's amplitudes assume that the quasi-periodic Gaussian process, with rotation-period priors fixed by a single activity-rotation relation, cleanly separates stellar activity from the planetary signal for all 44 stars; the model failures on TOI-269, TOI-4399 and HD 3167 show this separation is not guaranteed.","fun_headline_variants_meta":{"raw":{"variants":["Eccentricity priors shift exoplanet masses by 3x","How model choices skew small planet masses","Uniform reanalysis of 85 exoplanet masses","Activity models alter RV amplitudes up to 3x","HARPS reanalysis: priors matter for planet masses"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000288,"raw_usage":{"total_tokens":1703,"prompt_tokens":972,"completion_tokens":731,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":588,"completion_tokens_details":{"reasoning_tokens":651}},"tokens_in":588,"tokens_out":731,"duration_ms":6681,"temperature":1.0,"reasoning_tokens":651,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:12:48.593492+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the homogeneous RV amplitudes for a subset of the 87 planets against independent mass constraints from transit-timing variations or joint RV-plus-photometry fits: if the differences trace the presence of the GP or the eccentricity prior, rather than random scatter, the separation assumption is falsified. A sharper test is an injection-recovery experiment on the same HARPS time series with synthetic planets of known amplitude, checking whether each of the twelve models recovers the injected $K$ without bias.","supporting_citations":[{"cited_title":"C., Boué, G., Laskar, J., Delisle, J","cited_arxiv_id":null,"evidence_quote":"Used to explain the spurious high eccentricities produced by uniform eccentricity priors fitted to outliers."},{"cited_title":"2023, ESO/HARPS Radial Velocities Catalog Barragán, O., Aigrain, S., Kubyshkina, D., et al","cited_arxiv_id":null,"evidence_quote":"The curated HARPS radial-velocity catalogue used to match target names and assemble the 6428 observations."},{"cited_title":"2017, The Astro- nomical Journal, 154, 220","cited_arxiv_id":null,"evidence_quote":"The ensemble sampler underlying the MCMC posterior sampling."}],"review_version":1}