{"id":"033811b1-d9e8-4939-9145-d433e42e5b5a","arxiv_id":"2607.29232","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Machine-learning models trained on Rossiter-McLaughlin transit observations can partially reconstruct the underlying radial-velocity trend, but performance is uneven and activity correction remains unproven.","lead":"This paper trains machine-learning models on 1,171 ESPRESSO observations of 13 stars to see if they can reconstruct the underlying radial-velocity trend during planet transits, using the Rossiter-McLaughlin effect as a controlled test bed. The models worked for some stars but not others, and the authors themselves say the method is not yet ready to remove activity signals.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim that CCF line-profile diagnostics enable the RV reconstruction is not quantitatively supported: no ablation or baseline is reported, so the apparent predictive skill may come from the observed RVs or activity indicators alone.","rationale":"I focused on the evaluation logic rather than the linear-trend target because the authors explicitly justify the linear approximation for their nearly circular, short-duration systems. The reader's weakest-assumption is a legitimate secondary worry, but the more immediately load-bearing issue is that the reported qualitative figures cannot distinguish 'line-profile diagnostics are informative' from 'the model exploited the observed RVs or activity indicators.' The missing ablation/baseline is something the authors can fix with a simple re-run, and if it fails, the central claim is not supported. This aligns with the reader's conditional verdict; no change is needed, but the condition should explicitly include reporting an ablation.","tokens_in":3588,"tokens_out":9447,"duration_ms":104046,"concrete_test":"Run the same leave-one-star-out pipeline twice: (A) with the full feature set, and (B) with the five CCF line-profile diagnostics (FWHM, BIS, Vspan, Wspan, contrast) removed (or permuted within each sequence) while retaining mean-subtracted RVs and activity indicators. Report per-star and mean RMSE for both configurations. If RMSE in (B) is not substantially larger than in (A), the claim that line-profile diagnostics drive the reconstruction is unsupported; if (A) does not beat a simple 'predict observed RV' baseline, the reconstruction is not meaningful.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that ML models can reconstruct part of the apparent RV variations using line-profile diagnostics derived from CCFs. The reported evaluation does not yet establish this. The model's feature set includes the mean-subtracted observed RVs together with fractional line-profile diagnostics and activity indicators, but no numerical RMSE values are given, and no baseline or ablation is reported. Because the observed RV is the sum of the target linear trend and the RM anomaly, a model using only the observed RVs (or activity indicators) might already produce a reasonable approximation to the trend by smoothing or partial averaging, especially for weak anomalies. Without a comparison to a model that excludes the CCF line-profile diagnostics, the visual agreement in Figure 1 cannot be attributed specifically to those diagnostics. The paper's post-hoc statement that well-predicted stars have RM residuals correlated with diagnostics is suggestive but not a quantitative test. Thus, the strongest claim is underdetermined by the evidence presented; it needs an ablation to show that the line-profile diagnostics are the features doing the work.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents a proof-of-concept machine-learning study aimed at reconstructing the underlying radial-velocity (RV) trend during Rossiter-McLaughlin (RM) events from observed RVs, line-profile diagnostics, and activity indicators. The authors assembled 1171 ESPRESSO observations from 21 RM observing nights of 13 stars, defined a reference RV trend by a linear fit to out-of-transit points, and trained several ML regressors evaluated with leave-one-star-out cross-validation. They report that a voting ensemble performed best, that performance varied among stars, and that applications to Sun-as-a-star and Proxima Centauri data recovered known periodicities but did not convincingly remove activity signals. The central claim is that part of the apparent RV variations arising from flux-induced line-profile distortions can be reconstructed using CCF line-profile diagnostics.","tokens_in":3897,"tokens_out":3156,"duration_ms":33530,"significance":"If the claim were quantitatively established, the paper would offer a novel test bed for studying flux-induced RV distortions and a potential step toward mitigating activity signals in exoplanet searches, since RM events provide a known reference. The study is also useful as a demonstration of leave-one-star-out evaluation in this context. However, as presented, the evidence is qualitative: no numerical metrics, no ablation, and the target is an internal linear trend. The significance therefore remains conditional on additional analyses.","major_comments":[{"comment":"The text states that the 'best overall performance, quantified by the smallest mean RMSE' was obtained with a voting ensemble, but no RMSE values, error bars, or per-star metrics are reported anywhere. Without these numbers, the reader cannot judge the magnitude of the reconstruction error or compare it to the RM anomaly amplitude. Please report the RMSE (or equivalent) for each leave-one-star-out fold, the mean and scatter across folds, and an uncertainty estimate from repeated training runs or bootstrap resampling.","section":"Section 3, 'Model performance was assessed...'"},{"comment":"The input features include the mean-subtracted observed RVs together with line-profile diagnostics and activity indicators, while the target is a mean-subtracted linear trend fitted to out-of-transit RVs. Because the observed RV is the sum of this trend and the RM anomaly, a model using only the observed RVs could already approximate the target by interpolation or smoothing, especially for weak RM signals. No ablation or baseline is reported (e.g., models trained on observed RVs alone, on diagnostics alone, or on a trivial smoother). The central claim that the line-profile diagnostics are responsible for the reconstruction is therefore not established. Add explicit ablation experiments and report their numerical performance.","section":"Section 3, feature set; Section 2, target definition"},{"comment":"The regression target is a linear trend fitted to out-of-transit RVs, and the manuscript acknowledges that a Keplerian model would be ideal but argues the difference is small for nearly circular systems. This is plausible, but the target is still an internal construct derived from the same kind of data used as input features. If the out-of-transit RVs are contaminated by stellar activity (a central motivation of the paper), the fitted trend is biased, and the model is trained to predict that biased trend rather than the true orbital RV. The leave-one-star-out scheme prevents direct overfitting but does not address this label circularity. Please provide a validation using synthetic RM signals with known injected trends, or compare predictions against a Keplerian trend when available, to show that the method recovers the true underlying RV rather than an artifact of the fitting procedure.","section":"Section 2, 'To define the regression target...'"},{"comment":"The sample was reduced by visual inspection from 55 observing nights (41 stars) to 21 nights (13 stars), with no quantitative selection criteria. This subjective filtering could bias the sample toward clean, well-sampled RM events and inflate the apparent predictive performance. Please state the explicit criteria used for retention and, if possible, repeat the analysis on the full sample or with objective quality metrics. In addition, the Sun-as-a-star and Proxima tests are explicitly described as lying far outside the training domain (large Mahalanobis distances) and the recovered periodicities are contaminated by rotational modulation; these tests do not currently support the method's utility and should be presented only as null results or omitted.","section":"Section 2, sample selection; Section 3, applications to Sun and Proxima"}],"minor_comments":[{"comment":"The phrase 'controlled laboratory' in the abstract oversells the setup, since the target is a fitted linear trend rather than a known ground truth. Consider rephrasing to 'a test bed with an independently estimable reference trend.'","section":"Abstract and Introduction"},{"comment":"The caption states green points indicate out-of-transit measurements used to estimate the reference trend, but the right panel legend (blue, orange, green) is not described in the caption. Please clarify the meaning of each color in both panels.","section":"Figure 1"},{"comment":"The claim that 'no consistent improvement was obtained' is not quantified. If this negative result is retained, provide at least a representative metric; otherwise it is unfalsifiable.","section":"Section 3, last paragraph on transfer learning"},{"comment":"No data availability statement or repository link for the code and trained models is provided. For reproducibility, please include access to the sample, feature values, and implementation details (hyperparameters, feature preprocessing).","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The reader's conditional verdict and the skeptic's concern about missing ablations are, in my reading, well aligned with the manuscript's actual content. The central claim is not quantitatively supported as written. However, the paper is a proof-of-concept, and the required additions (RMSE tables, ablations, and a synthetic validation) are feasible within the scope of a revision. The paper may also benefit from an explicit discussion of the circularity inherent in using observed RVs as features to predict a smoothed version of those same RVs; the leave-one-star-out scheme mitigates but does not eliminate this concern."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely new framing — using RM sequences as a supervised calibration lab for flux-blocking RV distortions — and the leave-one-star-out evaluation is fair and honestly reported. But the paper currently argues more from visual agreement than from numbers. There are no RMSE values, no error bars, and — the stress-test point is right — no ablation that shows the CCF diagnostics are what does the work. Since the model gets the observed RVs as input and the target is a linear trend fitted to the same sequences, you need a baseline where the diagnostics are removed to know whether you are learning anything new. The paper's own post-hoc correlation argument is suggestive but not a test.\n\nWhat is good: the authors chose a well-posed setup, used a real ESPRESSO sample, excluded poorly-sampled stars instead of keeping everything, and are careful in the conclusions. They explicitly say the Sun and Proxima applications are outside the training domain and do not claim to have removed activity. That is the right scientific posture. The citation pattern is fine; the ML-in-RV references are current, and the paper does not oversell itself relative to them.\n\nSoft spots, in proportion: (1) The regression target is a fitted linear trend, not an independent truth. For the short transits of near-circular orbits that is defensible, but it still means the model learns to predict an internal construct. (2) Visual filtering from 55 nights to 21, with no stated criteria, invites selection bias. (3) No code or data release, which for an ML paper makes the results hard to check. (4) No baseline or ablation, which is the most serious gap. Without it, the central sentence in the conclusions overstates what is shown.\n\nWho this is for: exoplanet RV folks working on activity mitigation. They will want to know about this framing. It deserves a serious referee, but only if the authors add the missing quantitative elements. I would accept it for review, not for publication as is.","headline":"New framing with a fair test design, but the evidence is still qualitative: no metrics, no ablation, and an internal linear trend as target.","tokens_in":4363,"tokens_out":2833,"would_cite":true,"duration_ms":29260,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"During a Rossiter–McLaughlin transit, machine learning can partially recover the star's true radial-velocity trend from the distorted line profiles that create the anomaly.","keywords":["Rossiter–McLaughlin effect","radial velocities","machine learning","cross-correlation functions","line-profile diagnostics","stellar activity","ESPRESSO","exoplanets"],"falsifier":"Find an RM sequence for a system with an eccentric orbit or a transit long enough that the true Keplerian RV curve deviates from the local linear fit by more than the measurement uncertainty; if the trained model reproduces the linear reference instead of the true Keplerian curve, the reference-target assumption is refuted.","tokens_in":3499,"feed_emoji":"🔭","tokens_out":5232,"duration_ms":48679,"temperature":0.7,"pith_summary":"This paper asks whether a machine-learning model can recover the radial-velocity curve a star would have shown during a planet transit if the Rossiter–McLaughlin distortion had not occurred. The authors train regressors on 1,171 ESPRESSO observations of 13 stars across 21 transit nights, using the distorted RVs together with line-profile diagnostics and activity indicators. They show that part of the apparent RV shift caused by flux-blocking can be reconstructed from those diagnostics, so the model's predictions are closer to the uncontaminated trend than the raw observed RVs. Performance is uneven: it works best when the RM-induced residuals correlate strongly with line-shape diagnostics, and it degrades for weak signals or stars unlike the training set. Applied to Sun-as-a-star and Proxima Centauri, the model retains known periodicities but does not fully remove activity-induced variability, so the authors position the result as a proof of concept rather than a finished correction tool.","feed_headline":"ML rebuilds the star's true motion hidden by transits","feed_subtitle":"Line-profile diagnostics undo part of the Rossiter–McLaughlin shift on 21 ESPRESSO nights.","key_machinery":"The load-bearing object is the cross-correlation function (CCF) of each ESPRESSO spectrum, from which the apparent RV and a set of line-shape diagnostics are extracted. The regression target is the mean-subtracted linear trend fitted to out-of-transit RVs, which stands in for the star's true orbital velocity during the transit. The model that carries the argument is a voting ensemble of tree-based regressors (Random Forest, Extremely Randomised Trees, XGBoost, LightGBM, CatBoost) whose hyperparameters are tuned with Optuna and whose generalization is tested with a leave-one-star-out scheme. Feature selection, Mahalanobis-distance checks, and transfer-learning experiments with synthetic SOAP","core_discovery":"The central claim is that the apparent RV anomaly produced by a transiting planet is not just a nuisance; it is partially predictable from the same line-profile distortions that cause it. Using a voting ensemble of tree-based regressors trained on mean-subtracted RVs, fractional variations of CCF diagnostics (FWHM, BIS, Vspan, Wspan, contrast), and the mean and standard deviation of the Ca activity index, the model predicts a reference RV trend defined by a linear fit to out-of-transit data. In leave-one-star-out tests, the predicted trends match the reference for stars whose RM residuals correlate with the diagnostics, demonstrating that flux-induced line-profile deformation carries recover","pith_inferences":["The paper's reliance on a locally linear reference trend could be relaxed: for eccentric orbits or long transits, a Keplerian reference would test whether the model learns the true physical velocity or merely the linear approximation — an experiment the authors do not run but their data would support.","Because the RM effect physically mimics only the flux-imbalance component of starspots, combining this ML mapping with diagnostics sensitive to convective blueshift (such as the asymmetry of the CCF) could close part of the gap toward full activity correction; that combination is a natural, testable extension.","The strong dependence on training-set similarity suggests a practical pipeline: compute the Mahalanobis distance of a new star to the training distribution, and only trust the predicted correction when that distance is small; the paper stops short of proposing such a flag but its results justify it.","The failure of SOAP-based transfer learning hints that simulated line-profile distortions do not yet capture the full diversity of real RM and spot signals; a systematic comparison of simulated and observed diagnostics would pinpoint which missing physics matters."],"forward_implications":["If the reconstruction works in general, RV time series of transiting planets can be corrected for the RM distortion, yielding cleaner mass and orbit measurements from transit-night data.","The method establishes RM sequences as a calibration ground for flux-induced RV signals, which could help build empirical models of the spot-induced component of stellar activity.","The finding that performance depends on similarity between target and training stars implies that future applications should first screen targets by Mahalanobis distance, turning the method into a diagnostic tool as much as a correction tool.","The partial success on Sun-as-a-star and Proxima Centauri suggests the approach can preserve planetary periodicities while leaving rotational modulation largely intact — a caution that flux-blocking corrections alone are not enough for activity mitigation.","Scaling up the training sample to more stars and more diverse activity levels is the direct next step, and the paper points to spot-dominated stars as the most promising regime."],"fun_headline_variants":["ML predicts RV anomaly from line-profile distortions in RM effect","Line-profile diagnostics predict radial velocity shifts in transits","Machine learning reconstructs RV trends from RM effect observations","Predicting RV variations from Rossiter-McLaughlin line shapes","AI uses line-profile distortions to forecast RV shifts in RM data"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the straight line fitted to the out-of-transit RVs is a faithful stand-in for the star's true orbital velocity during the transit; if the true velocity curve bends or is contaminated by stellar activity inside the transit, the model is trained to predict the wrong reference.","fun_headline_variants_meta":{"raw":{"variants":["ML predicts RV anomaly from line-profile distortions in RM effect","Line-profile diagnostics predict radial velocity shifts in transits","Machine learning reconstructs RV trends from RM effect observations","Predicting RV variations from Rossiter-McLaughlin line shapes","AI uses line-profile distortions to forecast RV shifts in RM data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000226,"raw_usage":{"total_tokens":1278,"prompt_tokens":691,"completion_tokens":587,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":435,"completion_tokens_details":{"reasoning_tokens":519}},"tokens_in":435,"tokens_out":587,"duration_ms":6599,"temperature":1.0,"reasoning_tokens":519,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T11:04:57.262309+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Find an RM sequence for a system with an eccentric orbit or a transit long enough that the true Keplerian RV curve deviates from the local linear fit by more than the measurement uncertainty; if the trained model reproduces the linear reference instead of the true Keplerian curve, the reference-target assumption is refuted.","supporting_citations":[],"review_version":1}