Pith. sign in

REVIEW 4 major objections 4 minor 8 references

During a Rossiter–McLaughlin transit, machine learning can partially recover the star's true radial-velocity trend from the distorted line profiles that create the anomaly.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Machine-learning models trained on Rossiter-McLaughlin transit observations can partially reconstruct the underlying radial-velocity trend, but performance is uneven and activity correction remains unproven.

T0 review reviewed 2026-08-03 challenge →

load-bearing objection New framing with a fair test design, but the evidence is still qualitative: no metrics, no ablation, and an internal linear trend as target. the 4 major comments →

arxiv 2607.29232 v1 pith:WQJ4GIUP submitted 2026-07-31 physics.gen-ph

Predicting Radial Velocities from Rossiter-McLaughlin Time Series Observations

classification physics.gen-ph
keywords Rossiter–McLaughlin effectradial velocitiesmachine learningcross-correlation functionsline-profile diagnosticsstellar activityESPRESSOexoplanets
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether a machine-learning model can recover the radial-velocity curve a star would have shown during a planet transit if the Rossiter–McLaughlin distortion had not occurred. The authors train regressors on 1,171 ESPRESSO observations of 13 stars across 21 transit nights, using the distorted RVs together with line-profile diagnostics and activity indicators. They show that part of the apparent RV shift caused by flux-blocking can be reconstructed from those diagnostics, so the model's predictions are closer to the uncontaminated trend than the raw observed RVs. Performance is uneven: it works best when the RM-induced residuals correlate strongly with line-shape diagnostics, and it degrades for weak signals or stars unlike the training set. Applied to Sun-as-a-star and Proxima Centauri, the model retains known periodicities but does not fully remove activity-induced variability, so the authors position the result as a proof of concept rather than a finished correction tool.

Core claim

The central claim is that the apparent RV anomaly produced by a transiting planet is not just a nuisance; it is partially predictable from the same line-profile distortions that cause it. Using a voting ensemble of tree-based regressors trained on mean-subtracted RVs, fractional variations of CCF diagnostics (FWHM, BIS, Vspan, Wspan, contrast), and the mean and standard deviation of the Ca activity index, the model predicts a reference RV trend defined by a linear fit to out-of-transit data. In leave-one-star-out tests, the predicted trends match the reference for stars whose RM residuals correlate with the diagnostics, demonstrating that flux-induced line-profile deformation carries recover

What carries the argument

The load-bearing object is the cross-correlation function (CCF) of each ESPRESSO spectrum, from which the apparent RV and a set of line-shape diagnostics are extracted. The regression target is the mean-subtracted linear trend fitted to out-of-transit RVs, which stands in for the star's true orbital velocity during the transit. The model that carries the argument is a voting ensemble of tree-based regressors (Random Forest, Extremely Randomised Trees, XGBoost, LightGBM, CatBoost) whose hyperparameters are tuned with Optuna and whose generalization is tested with a leave-one-star-out scheme. Feature selection, Mahalanobis-distance checks, and transfer-learning experiments with synthetic SOAP

Load-bearing premise

The load-bearing premise is that the straight line fitted to the out-of-transit RVs is a faithful stand-in for the star's true orbital velocity during the transit; if the true velocity curve bends or is contaminated by stellar activity inside the transit, the model is trained to predict the wrong reference.

What would settle it

Find an RM sequence for a system with an eccentric orbit or a transit long enough that the true Keplerian RV curve deviates from the local linear fit by more than the measurement uncertainty; if the trained model reproduces the linear reference instead of the true Keplerian curve, the reference-target assumption is refuted.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If the reconstruction works in general, RV time series of transiting planets can be corrected for the RM distortion, yielding cleaner mass and orbit measurements from transit-night data.
  • The method establishes RM sequences as a calibration ground for flux-induced RV signals, which could help build empirical models of the spot-induced component of stellar activity.
  • The finding that performance depends on similarity between target and training stars implies that future applications should first screen targets by Mahalanobis distance, turning the method into a diagnostic tool as much as a correction tool.
  • The partial success on Sun-as-a-star and Proxima Centauri suggests the approach can preserve planetary periodicities while leaving rotational modulation largely intact — a caution that flux-blocking corrections alone are not enough for activity mitigation.
  • Scaling up the training sample to more stars and more diverse activity levels is the direct next step, and the paper points to spot-dominated stars as the most promising regime.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The paper's reliance on a locally linear reference trend could be relaxed: for eccentric orbits or long transits, a Keplerian reference would test whether the model learns the true physical velocity or merely the linear approximation — an experiment the authors do not run but their data would support.
  • Because the RM effect physically mimics only the flux-imbalance component of starspots, combining this ML mapping with diagnostics sensitive to convective blueshift (such as the asymmetry of the CCF) could close part of the gap toward full activity correction; that combination is a natural, testable extension.
  • The strong dependence on training-set similarity suggests a practical pipeline: compute the Mahalanobis distance of a new star to the training distribution, and only trust the predicted correction when that distance is small; the paper stops short of proposing such a flag but its results justify it.
  • The failure of SOAP-based transfer learning hints that simulated line-profile distortions do not yet capture the full diversity of real RM and spot signals; a systematic comparison of simulated and observed diagnostics would pinpoint which missing physics matters.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The manuscript presents a proof-of-concept machine-learning study aimed at reconstructing the underlying radial-velocity (RV) trend during Rossiter-McLaughlin (RM) events from observed RVs, line-profile diagnostics, and activity indicators. The authors assembled 1171 ESPRESSO observations from 21 RM observing nights of 13 stars, defined a reference RV trend by a linear fit to out-of-transit points, and trained several ML regressors evaluated with leave-one-star-out cross-validation. They report that a voting ensemble performed best, that performance varied among stars, and that applications to Sun-as-a-star and Proxima Centauri data recovered known periodicities but did not convincingly remove activity signals. The central claim is that part of the apparent RV variations arising from flux-induced line-profile distortions can be reconstructed using CCF line-profile diagnostics.

Significance. If the claim were quantitatively established, the paper would offer a novel test bed for studying flux-induced RV distortions and a potential step toward mitigating activity signals in exoplanet searches, since RM events provide a known reference. The study is also useful as a demonstration of leave-one-star-out evaluation in this context. However, as presented, the evidence is qualitative: no numerical metrics, no ablation, and the target is an internal linear trend. The significance therefore remains conditional on additional analyses.

major comments (4)
  1. [Section 3, 'Model performance was assessed...'] The text states that the 'best overall performance, quantified by the smallest mean RMSE' was obtained with a voting ensemble, but no RMSE values, error bars, or per-star metrics are reported anywhere. Without these numbers, the reader cannot judge the magnitude of the reconstruction error or compare it to the RM anomaly amplitude. Please report the RMSE (or equivalent) for each leave-one-star-out fold, the mean and scatter across folds, and an uncertainty estimate from repeated training runs or bootstrap resampling.
  2. [Section 3, feature set; Section 2, target definition] The input features include the mean-subtracted observed RVs together with line-profile diagnostics and activity indicators, while the target is a mean-subtracted linear trend fitted to out-of-transit RVs. Because the observed RV is the sum of this trend and the RM anomaly, a model using only the observed RVs could already approximate the target by interpolation or smoothing, especially for weak RM signals. No ablation or baseline is reported (e.g., models trained on observed RVs alone, on diagnostics alone, or on a trivial smoother). The central claim that the line-profile diagnostics are responsible for the reconstruction is therefore not established. Add explicit ablation experiments and report their numerical performance.
  3. [Section 2, 'To define the regression target...'] The regression target is a linear trend fitted to out-of-transit RVs, and the manuscript acknowledges that a Keplerian model would be ideal but argues the difference is small for nearly circular systems. This is plausible, but the target is still an internal construct derived from the same kind of data used as input features. If the out-of-transit RVs are contaminated by stellar activity (a central motivation of the paper), the fitted trend is biased, and the model is trained to predict that biased trend rather than the true orbital RV. The leave-one-star-out scheme prevents direct overfitting but does not address this label circularity. Please provide a validation using synthetic RM signals with known injected trends, or compare predictions against a Keplerian trend when available, to show that the method recovers the true underlying RV rather than an artifact of the fitting procedure.
  4. [Section 2, sample selection; Section 3, applications to Sun and Proxima] The sample was reduced by visual inspection from 55 observing nights (41 stars) to 21 nights (13 stars), with no quantitative selection criteria. This subjective filtering could bias the sample toward clean, well-sampled RM events and inflate the apparent predictive performance. Please state the explicit criteria used for retention and, if possible, repeat the analysis on the full sample or with objective quality metrics. In addition, the Sun-as-a-star and Proxima tests are explicitly described as lying far outside the training domain (large Mahalanobis distances) and the recovered periodicities are contaminated by rotational modulation; these tests do not currently support the method's utility and should be presented only as null results or omitted.
minor comments (4)
  1. [Abstract and Introduction] The phrase 'controlled laboratory' in the abstract oversells the setup, since the target is a fitted linear trend rather than a known ground truth. Consider rephrasing to 'a test bed with an independently estimable reference trend.'
  2. [Figure 1] The caption states green points indicate out-of-transit measurements used to estimate the reference trend, but the right panel legend (blue, orange, green) is not described in the caption. Please clarify the meaning of each color in both panels.
  3. [Section 3, last paragraph on transfer learning] The claim that 'no consistent improvement was obtained' is not quantified. If this negative result is retained, provide at least a representative metric; otherwise it is unfalsifiable.
  4. [General] No data availability statement or repository link for the code and trained models is provided. For reproducibility, please include access to the sample, feature values, and implementation details (hyperparameters, feature preprocessing).

Circularity Check

0 steps flagged

No significant circularity: the ML target is an estimated reference trend, but the prediction is not equivalent to the input features by construction.

full rationale

The paper trains ML models to predict a reference linear RV trend fitted to out-of-transit measurements, using observed RVs, line-profile diagnostics, and activity indicators as features. The target is not defined in terms of the model outputs, and the model predictions are not used to define the target. Leave-one-star-out evaluation prevents the model from directly memorizing the fitted trends of the evaluated stars. Although the reference trend is an estimate rather than an independent external truth (Section 2: 'the fitted trend provides an estimate of the underlying RV evolution during the observing sequence in the absence of the RM anomaly'), this is a standard supervised-learning setup with an estimated label; it does not make the prediction statistically forced or equivalent to the input by construction. The paper's main weaknesses are evidentiary rather than circular: no numerical RMSE values or ablation analyses are reported, so the specific contribution of the CCF line-profile diagnostics beyond the observed RVs alone is not quantitatively established, and the paper itself flags the Sun-as-a-star and Proxima applications as unreliable because of large Mahalanobis distances ('The reliability of the inferred corrections is therefore uncertain'). These are concerns about support and evaluation, not circularity. The only self-citation with potential load-bearing role, Cristo et al. 2025 for SOAP transfer learning, is explicitly reported as yielding no improvement ('No consistent improvement was obtained'), so it is not load-bearing. No step reduces the derivation to its inputs by definition, and no imported uniqueness theorem or ansatz-by-citation is used.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 0 invented entities

The paper introduces no new physical entities. It relies on domain assumptions about the linear trend proxy, the transferability of RM to activity signals, and the informativeness of CCF diagnostics, plus a sample-selection assumption that is ad hoc to this study.

free parameters (4)
  • Reference RV trend slope and intercept (per observing sequence) = not reported
    Linear fit to out-of-transit RVs defines the regression target; fitted separately for each of the 21 observing nights.
  • ML hyperparameters (Ridge, random forest, XGBoost, LightGBM, CatBoost, MLP, ensembles) = not reported
    Optimized with Optuna; no final values, search ranges, or selected configurations are given.
  • Feature set = mean-subtracted RV, fractional line-profile diagnostics, Ca mean/std
    Selected based on best leave-one-star-out RMSE; the choice is data-driven and not fixed a priori.
  • Per-sequence normalization = per-sequence mean
    Features are normalized independently for each observing sequence, using the data itself to set the scale.
axioms (5)
  • domain assumption The linear trend fitted to out-of-transit RVs accurately approximates the true Keplerian orbital RV during transit.
    Section 2: if the true RV curve deviates from a straight line during transit, the regression target is biased and the model learns to predict an artifact.
  • domain assumption RM flux-blocking is a valid proxy for spot-induced RV signals, so learned corrections transfer to activity-dominated targets.
    Section 4: this transfer is the stated motivation, but the paper admits RM does not reproduce convective-blueshift suppression and the Sun/Proxima tests were unreliable.
  • domain assumption CCF line-profile diagnostics and activity indicators contain enough information to separate the RM anomaly from the underlying trend.
    Section 3: the entire ML input representation rests on this; poor predictions are attributed to weak correlation with the diagnostics.
  • ad hoc to paper The retained 13-star/21-night sample, after visual selection for well-sampled RM, is representative enough for leave-one-star-out generalization.
    Section 2: the selection is visual and may remove the very cases (weak RM, high uncertainty) where the method would fail, inflating apparent performance.
  • standard math Standard supervised-learning assumptions (i.i.d. features, stationarity of the RV-diagnostic relation) hold across stars.
    Implicit in applying regression and leave-one-star-out evaluation.

reviewed 2026-08-03 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Predicting Radial Velocities from Rossiter-McLaughlin Time Series Observations." pith.science (2026). https://pith.science/paper/WQJ4GIUP

@misc{pith2026260729232,
  author       = {Pith},
  title        = {Pith review of: Predicting Radial Velocities from Rossiter-McLaughlin Time Series Observations},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WQJ4GIUP}},
  note         = {Machine review of arXiv:2607.29232}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

The Rossiter-McLaughlin (RM) effect produces apparent radial velocity (RV) shifts through line-profile distortions caused by a transiting planet blocking different regions of the rotating stellar surface. Because the underlying orbital RV trend can be estimated from out-of-transit observations, RM sequences provide a controlled laboratory for studying flux-induced RV variations. We compiled a sample of 1171 ESPRESSO observations of 13 targets obtained during 21 RM observing nights and trained machine-learning models to reconstruct a reference RV trend from observed RVs, line-profile diagnostics, and activity indicators. Predictive performance varied substantially among stars and depended on both the strength of the RM signal and the similarity of the target star to the training sample. Applications to Sun-as-a-star observations and Proxima Centauri recovered known periodicities but did not fully remove the activity-induced variability. Larger and more diverse datasets will be required to assess the potential of this approach for mitigating activity-induced RV signals.

Figures

Figures reproduced from arXiv: 2607.29232 by Andre Silva, Artur Hakobyan, Diogo Teixeira, Eduardo Cristo, Garik Israelian, Khaled Al Moulla, Nuno C. Santos, Olivier Demangeon, Roman Chertovskih, Vardan Adibekyan.

Figure 1
Figure 1. Figure 1: Left: RM time series for WASP-118 and WASP-166 after subtraction of the mean RV. Green points indicate the out-of-transit measurements used to estimate the reference RV trend (solid line), while grey points are affected by the RM anomaly and are excluded from the fit. Right: Leave-one-star-out predictions for the same systems. The panels compare the observed RVs (blue), the ML-predicted reference RV trend … view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

8 extracted references · 6 linked inside Pith

  1. [1]

    The HARPS search for southern extra-solar planets. XXXV. The interesting case of HD 41248: stellar activity, no planets?. , keywords =. doi:10.1051/0004-6361/201423808 , archivePrefix =. 1404.6135 , primaryClass =

  2. [2]

    , keywords =

    Prospects for the Characterization and Confirmation of Transiting Exoplanets via the Rossiter-McLaughlin Effect. , keywords =. doi:10.1086/509910 , archivePrefix =. astro-ph/0608071 , primaryClass =

  3. [3]

    Identifying Exoplanets with Deep Learning. VI. Enhancing Neural Network Mitigation of Stellar Activity RV Signals with Additional Metrics. , keywords =. doi:10.3847/1538-3881/ae45fd , archivePrefix =. 2602.17760 , primaryClass =

  4. [4]

    Constraining exoplanet blend scenarios using spectroscopic diagnoses

    PASTIS: Bayesian extrasolar planet validation - II. Constraining exoplanet blend scenarios using spectroscopic diagnoses. , keywords =. doi:10.1093/mnras/stv1080 , archivePrefix =. 1505.02663 , primaryClass =

  5. [5]

    , keywords =

    Revisiting Proxima with ESPRESSO. , keywords =. doi:10.1051/0004-6361/202037745 , archivePrefix =. 2005.12114 , primaryClass =

  6. [6]

    , keywords =

    Three years of Sun-as-a-star radial-velocity observations on the approach to solar minimum. , keywords =. doi:10.1093/mnras/stz1215 , archivePrefix =. 1904.12186 , primaryClass =

  7. [7]

    The Journal of Open Source Software , keywords =

    ACTIN: A tool to calculate stellar activity indices. The Journal of Open Source Software , keywords =. doi:10.21105/joss.00667 , archivePrefix =. 1811.11172 , primaryClass =

  8. [8]

    , keywords =

    SOAPv4: A new step toward modeling stellar signatures in exoplanet research. , keywords =. doi:10.1051/0004-6361/202555184 , archivePrefix =. 2510.08319 , primaryClass =

This paper was first reviewed by deepseek-v4-flash on August 3, 2026.