{"id":"b5588be9-e161-40cc-a308-f9630717ed8d","arxiv_id":"1908.07560","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"An iSHELL K-band forward-modeling pipeline with a methane gas cell achieves 3 to 5 m/s radial velocity precision on three cool dwarf stars.","lead":"This paper describes a data pipeline that measures the tiny back-and-forth motions of cool, low-mass stars using infrared spectra from the iSHELL spectrograph. It reports velocity precisions of a few meters per second, good enough to follow up planet candidates found by NASA's TESS mission.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Self-built template and post-hoc order selection leave the m/s claim internally consistent but externally unvalidated; recovering the known GJ 15 Ab signal would test whether real RV signals survive.","rationale":"I read the paper as a methods demonstration whose central assertion is that iSHELL plus the PySHELL forward-modeling approach reaches a few m/s precision on K and M dwarfs, sufficient for TESS planet confirmation at K > 3 m/s. The strongest evidence is the small long-term RMS around zero for three stars, the approximate photon-noise consistency for GJ 15 A and 61 Cyg A, the order-to-order consistency of telluric optical depths, and the qualitative agreement between two independently retrieved seasonal templates. The reduction code is public and the data are archived, which supports reproducibility. The weak point is not any single equation: the forward model is elaborate and the solver is plausibly adequate. The weakness is that the quoted precision is measured against a reference spectrum built from the same data, and the best-case numbers are selected from a powerset of order combinations. The paper itself is candid about both limitations, and it does provide a useful bound by noting that with at least eight orders, a 5-7 m/s precision is typical rather than a lucky draw. That candor is a point in the paper's favor. Even so, the abstract's 'demonstrated 5 m/s precision' and '3 m/s over a month' are best-case values, not a pre-specified recipe, and no independent signal has been recovered to prove that the template does not absorb real velocity variations. The reader's verdict of CONDITIONAL already captures this: the method is promising, but the precision claim should be verified against a known planet ephemeris or an independent RV standard. My stress-test does not move that verdict; it sharpens the condition by specifying GJ 15 Ab as the natural, immediately available test case. I therefore recommend no change to the reader's conditional assessment.","tokens_in":22735,"tokens_out":7882,"duration_ms":660673,"concrete_test":"Re-run the pipeline on GJ 15 A using a fixed, pre-registered order set (e.g., orders 8-10, or all usable orders 6-17) instead of powerset-selected orders, and fit the nightly RVs with a Keplerian at the known ephemeris of GJ 15 Ab (P=11.44 d, K=2.9 m/s, Howard et al. 2014) plus a constant. Require the recovered semi-amplitude and phase to agree with the published ephemeris within the quoted uncertainties. If the signal is not recovered or is suppressed below roughly 2 m/s, the claimed 3 m/s precision and the TESS follow-up capability are not established. A stronger version uses two independently built templates from disjoint subsets of epochs; agreement of the recovered planet parameters would simultaneously rule out template absorption of real signals.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—5 m/s over one year and 3 m/s over a month—rests on RVs measured relative to a template derived from the same observations (Sec. 4.3) and on 'best case' order subsets chosen from a powerset (Sec. 5.1). The paper explicitly acknowledges the template risk in Sec. 4.3: 'residual correlated noise can gradually get repeatedly added into the stellar template from missed bad pixels, or from non-stellar spectral features that are not well fit.' The seasonal-template comparison in Sec. 6 checks consistency between two templates, but it does not show that the pipeline preserves real, time-variable stellar velocity signals at the few m/s level. A pipeline can show low RMS while a genuine planet signal is partially absorbed into the template or removed by the order-zero-point/detrending steps. This matters directly for the stated goal of confirming TESS planets with K > 3 m/s. The reported chi^2_red values of 0.5-0.8 (Table 5) are not decisive because the uncertainties come from the same forward model. The load-bearing gap is the absence of any external check against an independent RV time series or a known ephemeris.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript describes a new data-analysis pipeline for extracting radial velocities from K-band (2.18-2.47 μm) spectra taken with iSHELL at the IRTF. The pipeline forward-models each echelle order with 48 parameters: a 13CH4 gas-cell transmission, four telluric absorbers, two fringing sources (OS filter and AR coating), a residual blaze, a Hermite-polynomial LSF, a wavelength solution with spline corrections, and an iteratively retrieved stellar template. RVs are combined across orders using weighted statistics or a TFA-like detrending minimization. Applying the pipeline to Barnard's Star, GJ 15 A, and 61 Cygni A, the authors report best-case long-term RV RMS values of 4.3, 2.7, and 3.8 m/s, respectively, after selecting a subset of orders by a powerset search and, for 61 Cyg A, discarding one night with a +1 km/s outlier. The paper claims 5 m/s precision over one-year baselines for two stars and 3 m/s over one month for GJ 15 A, and argues that this enables TESS planet confirmation around K and M dwarfs.","tokens_in":23078,"tokens_out":7504,"duration_ms":72852,"significance":"The manuscript's contribution, if the claimed precision is externally validated, is significant: it would make iSHELL one of the few instruments delivering few-m/s RVs in the K band using a Cassegrain-mounted spectrograph and a methane isotopologue gas cell, with public data. The forward-model architecture is detailed, and the inclusion of multiple fringing sources and iterative template retrieval is technically ambitious. The paper also makes honest statements about its limitations, including template corruption, uncharacterized telluric error, and order-selection freedom. The main weakness is that the headline numbers are internal, best-case metrics rather than predictive, externally anchored benchmarks: the stellar template is derived from the same spectra, the order subset is chosen post hoc, and no known planetary signal is recovered. Consequently, the paper currently demonstrates internal consistency and a plausible precision floor, but not that real RV signals at the few-m/s level survive the pipeline. If the recommended external-validation tests are added, this would be a valuable instrument-paper for TESS follow-up.","major_comments":[{"comment":"The headline precision values are minima of a powerset search over order combinations. For Barnard's Star (high SNR), the reported 4.33 m/s is the minimum of 4083 combinations; for GJ 15 A, 2.72 m/s comes from only three orders (8, 9, 10) and six nights; for 61 Cyg A, 3.77 m/s comes from five orders. A minimum over a large set is an optimistically biased statistic and is not a prediction of the precision obtained when the order set is fixed in advance. The authors partially acknowledge this at the end of Section 5.2 ('when observing stars with unknown RVs, we do not have this freedom'), but the Abstract and Section 5.1 still present the best-case values as the demonstrated precision. Additionally, the best single-order precisions in Table 4 are quoted at order-dependent 'best iteration' values, and it is unclear whether the multi-order subsets use different iterations per order, which adds further post-hoc freedom. I recommend reporting the full distribution of σ over the powerset or quoting a pre-specified fixed-order precision (e.g., the 5-7 m/s for at least 8 orders mentioned in Section 5.2) as the primary claim.","section":"Section 5.1, Table 5, Fig. 8"},{"comment":"The stellar template is constructed iteratively from the same target spectra, and the authors themselves state that 'residual correlated noise can gradually get repeatedly added into the stellar template from missed bad pixels, or from non-stellar spectral features that are not well fit.' This creates a circularity for the precision claim: low scatter of RVs relative to a template that may have absorbed some of the correlated signal does not prove that real, time-variable stellar velocities are preserved. The seasonal template comparison in Section 6 checks consistency of deep stellar lines, but it does not test whether signals at the few-m/s level survive. Because GJ 15 A is listed in Table 1 as hosting a planet with K=2.9 m/s and P=11.44 d, the paper should attempt to recover this known signal (e.g., by fitting the six nights to the published ephemeris or showing an RV periodogram) as an external validation. Barnard's Star's 233-day, 1.2 m/s signal is smaller but could also be checked. Without such a test, the m/s claim remains an internal precision metric rather than a demonstrated capability to detect or confirm planets.","section":"Section 4.3, Eq. (1), Table 5"},{"comment":"Telluric error is explicitly uncharacterized. The text states 'we do not characterize this' regarding water vapor variability and, in Section 5.2, 'Determining telluric induced error on RVs is the subject of a future investigation.' The K-band orders contain deep, variable water and methane lines; the forward model has four telluric species with a shared velocity shift; and order 14 is flagged as an outlier for all three targets, suggesting a telluric or gas-cell template problem. Since telluric absorption is a major potential contributor to the m/s error budget, the paper needs at least an upper-limit estimate, for example comparing RVs from orders with high versus low telluric absorption or injecting synthetic telluric variations. Until then, the claim that the achieved precision is below the telluric noise floor is unsupported.","section":"Section 2 and Section 5.2"},{"comment":"The night JD 263.01044249 with RV = 1403 m/s is discarded on the assumption that it is an observational error or a flare, with no independent evidence. This exclusion is load-bearing for the '3.8 m/s for 61 Cyg A' claim. The paper should state the RMS with and without this night and ideally investigate the cause (e.g., checking target acquisition, flat fields, or telluric residuals). Also, the phrase 'over one year timescales' for 61 Cyg A is based on 10 nights spanning about 254 days after the exclusion; the effective baseline and number of epochs should be stated explicitly.","section":"Section 5.1, Table 5 (61 Cyg A)"}],"minor_comments":[{"comment":"Row 10 labels the 'OS Filter Fringing Finesse' as 'FAR' but the symbol for the OS fringing finesse should be 'FOS' to match Eq. (5); 'FAR' is already used for the AR fringing finesse in row 15.","section":"Table 3"},{"comment":"The first paragraph contains the duplicated phrase 'which can which can'; this should be corrected.","section":"Section 4.3"},{"comment":"The reduced chi-squared values of 0.5-0.8 in Table 5 are computed using uncertainties that themselves come from the forward model; a sentence clarifying that these values do not independently validate the error bars would prevent over-interpretation.","section":"Section 5.2, Eq. (17)"},{"comment":"The comparison of the two seasonal templates is qualitative only; reporting a quantitative metric (e.g., RMS difference over pixels with line depth >2%) would make the assessment more reproducible.","section":"Section 6"},{"comment":"The caption should explicitly state that the yellow histogram includes combinations of 2-12 orders while the green histogram is restricted to 10-12 orders; the current caption relies on the main text for this information.","section":"Figure 8"},{"comment":"The core pipeline PySHELL is listed as 'Available upon request' rather than deposited in a public repository; making it available publicly, as was done for the reduction code, would strengthen the reproducibility of the results.","section":"Reproducibility"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's central claim rests on internal consistency metrics, and the two largest risks are the post-hoc order-selection bias and the absence of an external signal-recovery test. I would treat a successful recovery of the known GJ 15 Ab signal (or a clearly explained non-detection) as a necessary condition for acceptance. The powerset and outlier-exclusion issues are fixable by reporting fixed-order precisions and including the outlier in the reported statistics, but the template circularity concern requires an external anchor that the current data presentation does not provide."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a genuine methods paper, the first published precision numbers for iSHELL K-band RVs, and the authors are unusually candid about what could go wrong. The headline precision is real internal scatter, but it is not externally validated, and the paper's own Section 4.3 names the load-bearing risk: the stellar template is built from the same spectra being fit, so correlated noise can be absorbed into the template and make the velocities look better than the measurement truly supports.\n\nWhat's new and good: the whole-order forward model, the custom anti-reflection fringing model, the iterative template retrieval applied to iSHELL, and the TFA-style detrending are all sensible extensions of Gao et al. (2016). The data are public, the reduction code is on GitHub, and the paper includes honest exploration of parameter correlations, LSF alternatives, and spline choices. They also report typical precision with many orders (5-7 m/s) rather than only the best-case powerset numbers, which is more honest than most.\n\nSoft spots, in proportion: the paper acknowledges the template risk but only checks two seasonal templates against each other, not against an external reference. That is a consistency check, not a fidelity check. The best-case precision comes from a powerset over order subsets, and one outlier night for 61 Cyg is discarded with a plausible but unproven explanation. Telluric errors are explicitly uncharacterized, and PySHELL isn't released. None of these are fatal for a methods demonstration, but together they mean the 5 m/s and 3 m/s claims should be read as upper bounds on internal precision, not as a validated floor. The cleanest fix would be recovering a known signal—GJ 15 Ab has a published 2.9 m/s ephemeris, and the companion paper (Plavchan et al., submitted) apparently does detect planets with this pipeline, but that evidence isn't in this paper.\n\nFor whom: RV specialists working on NIR follow-up of M dwarfs will want this as the reference for iSHELL capabilities. It deserves peer review, but I'd ask for an external check—either a known planet ephemeris or an independent RV time series—before accepting the headline precision. Serious refs should treat it as a methods paper, not a discovery paper.","headline":"First iSHELL K-band RV precision paper: solid pipeline work, honest limitations, but the m/s claims are internal scatter and need an external check before I'd trust them for TESS follow-up.","tokens_in":23671,"tokens_out":2706,"would_cite":true,"duration_ms":28433,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A K-band spectrograph with a methane-isotopologue gas cell and a 48-parameter forward model achieves 3–5 m/s radial velocities on cool K and M dwarfs, enough to confirm TESS planet candidates with semi-amplitudes above about 3 m/s.","keywords":["radial velocities","M dwarfs","K-band spectroscopy","iSHELL spectrograph","gas cell calibration","forward modeling","stellar template","exoplanet confirmation"],"falsifier":"Inject a known synthetic Doppler shift, for example 10 m/s, into one night's raw spectra before running the pipeline; if the recovered shift differs from the injected value by more than the quoted nightly uncertainty, the stellar template has absorbed correlated noise rather than the true stellar spectrum.","tokens_in":22538,"feed_emoji":"🔭","tokens_out":15240,"duration_ms":128836,"temperature":0.7,"pith_summary":"The paper claims that radial velocities of cool low-mass stars can be measured to 5 meters per second over a year and 3 meters per second over a month using K-band spectra from the iSHELL spectrograph, with a methane-isotopologue gas cell providing the wavelength reference. The central advance is that the stellar reference spectrum is not taken from atmosphere models; it is built from the target observations themselves by iteratively co-adding barycenter-shifted residuals into a template. On three stars previously known to be stable in radial velocity, the method reports best-case long-term RMS values of 4.3 m/s for Barnard's Star, 2.7 m/s for GJ 15 A, and 3.8 m/s for 61 Cygni A. If the precision holds, the same instrument can confirm and weigh planet candidates from the TESS transit survey around bright K and M dwarfs, and can search for planets around moderately active and young cool stars where visible-wavelength velocities are corrupted by spots.","feed_headline":"Near-infrared pipeline reaches 3 m/s on cool dwarf stars","feed_subtitle":"A methane gas cell and a self-built stellar template let one telescope weigh planets around K and M dwarfs.","key_machinery":"The load-bearing mechanism is the iterative stellar template retrieval: starting from a flat guess, the pipeline forward-models each spectrum, subtracts the model, shifts the residuals into the star's barycentric rest frame, median-combines them with inverse-RMS-squared weights, and adds the result back into the template; after 5–40 iterations the extracted velocities stabilize and the template approaches the deconvolved stellar spectrum. Two calibration devices carry the precision: the methane isotopologue ($^{13}$CH$_4$) gas cell, whose FTS-measured transmission provides a common optical-path wavelength reference and constrains the line-spread function, and explicit Fabry-Perot models for the two fringing sources, the order-selection filter and the anti-reflection coating of the silicon immersion grating. The 48-parameter model is optimized with a custom Nelder-Mead solver that alternates full simplex calls with two-dimensional subspace calls, because standard simplex optimization did not converge in this parameter space.","core_discovery":"On its own terms, the paper demonstrates that a 48-parameter forward model can reproduce K-band (2.18–2.47 µm) spectra of cool dwarfs well enough to extract relative radial velocities at the few-m/s level on baselines from one month to one year. The model accounts for the Doppler-shifted stellar spectrum, the methane gas cell transmission, Doppler-shifted telluric water, methane, nitrous oxide, and carbon dioxide, the residual blaze function, a quadratic-plus-spline wavelength solution, the spectrograph line-spread function, and two separate quasi-sinusoidal fringing patterns from the order-selection filter and the immersion-grating anti-reflection coating. The stellar template is derived iteratively: beginning from a flat guess, the pipeline forward-models every spectrum, shifts the residuals into the star's barycentric rest frame, median-combines them with $\\mathrm{RMS}^{-2}$ weighting, and adds the result back into the template, repeating for 41 iterations. The quoted best-case multi-order long-term RMS values are 4.3 m/s for Barnard's Star (high-SNR orders), 5.13 m/s using all Barnard data, 2.72 m/s for GJ 15 A, and 3.77 m/s for 61 Cygni A.","pith_inferences":["The same iterative-template scheme could in principle be run on archival iSHELL K-band data, producing long-baseline RV time series for a much larger sample without any new observations.","The paper's finding that telluric optical depths are consistent across orders suggests that a future joint fit over all orders, sharing telluric and fringing parameters, could shrink the 48-parameter freedom and push precision below the current 3–5 m/s floor.","Since the template is built from barycenter-shifted residuals, a testable requirement is that each target be observed at enough epochs spread over the year; one could derive a minimum-epoch criterion from the convergence behavior, something the paper does not quantify."],"forward_implications":["Planet candidates from the TESS transit survey that orbit K and M dwarfs brighter than K magnitude 9 and have velocity semi-amplitudes above roughly 3 m/s can be confirmed and their masses measured; the paper estimates on the order of 100 such candidates are amenable to iSHELL follow-up.","Near-infrared activity jitter should be reduced relative to the optical by roughly the frequency ratio, so a star with 5 m/s optical activity would show less than about 1.5 m/s in the K band, opening searches around moderately active and young cool stars.","Because the stellar template is empirical and built from the target itself, the method avoids relying on synthetic stellar atmosphere models, which are known to be deficient for late M dwarfs with complex molecular opacities.","Combining at least eight echelle orders should deliver long-term precision of 5–7 m/s for typical K and M dwarfs with sufficient RV content, because single-order precision improves as $N^{-1/2}$ with the number of orders."],"supporting_citations":[{"why":"Supplies the CSHELL forward-modeling framework this pipeline adapts, including the finding that unmodeled fringing induces >50 m/s errors.","marker":"Gao et al. (2016)"},{"why":"Provides the iterative deconvolution method used to build the stellar template from the target observations themselves.","marker":"Sato et al. (2002)"},{"why":"Provides the FTS-measured methane gas cell spectrum used as the wavelength reference in the forward model.","marker":"Anglada-Escudé et al. (2012)"},{"why":"Establishes the gas-cell calibration technique for precise near-infrared radial velocities that iSHELL's methane cell implements.","marker":"Plavchan et al. (2013)"},{"why":"Supplies the telluric transmission templates for water, methane, nitrous oxide, and carbon dioxide used in the forward model.","marker":"Bertaux et al. (2014)"},{"why":"Gives the photon-noise precision formula used to compare measured nightly uncertainties with the expected noise floor.","marker":"Bouchy et al. (2001)"},{"why":"Provides the barycentric correction code used to shift residuals into the stellar rest frame during template retrieval.","marker":"Wright & Eastman (2014)"},{"why":"Provides the 4.515 m/s per year secular acceleration subtracted from Barnard's Star velocities before computing long-term precision.","marker":"Choi et al. (2012)"}],"fun_headline_variants":["K-band RVs for cool dwarfs hit 3 m/s","Methane cell + self-built template: 3 m/s RVs","iSHELL measures cool dwarf RVs to 3 m/s","Few m/s precision for M dwarfs in K band","Barnard's Star RV precision reaches 3 m/s"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes the template built from the star's own spectra converges to the true stellar spectrum, rather than gradually soaking up fringing, telluric, or bad-pixel artifacts that would make the measured velocities look more precise than they really are.","fun_headline_variants_meta":{"raw":{"variants":["K-band RVs for cool dwarfs hit 3 m/s","Methane cell + self-built template: 3 m/s RVs","iSHELL measures cool dwarf RVs to 3 m/s","Few m/s precision for M dwarfs in K band","Barnard's Star RV precision reaches 3 m/s"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000248,"raw_usage":{"total_tokens":1623,"prompt_tokens":1095,"completion_tokens":528,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":711,"completion_tokens_details":{"reasoning_tokens":439}},"tokens_in":711,"tokens_out":528,"duration_ms":5846,"temperature":1.0,"reasoning_tokens":439,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:06:36.563800+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inject a known synthetic Doppler shift, for example 10 m/s, into one night's raw spectra before running the pipeline; if the recovered shift differs from the injected value by more than the quoted nightly uncertainty, the stellar template has absorbed correlated noise rather than the true stellar spectrum.","supporting_citations":[{"cited_title":"2016, PASP, 128, 104501","cited_arxiv_id":null,"evidence_quote":"Supplies the CSHELL forward-modeling framework this pipeline adapts, including the finding that unmodeled fringing induces >50 m/s errors."},{"cited_title":"2002, PASJ, 54, 873","cited_arxiv_id":null,"evidence_quote":"Provides the iterative deconvolution method used to build the stellar template from the target observations themselves."},{"cited_title":"P., Anglada-Escude, G., White, R., et al","cited_arxiv_id":null,"evidence_quote":"Establishes the gas-cell calibration technique for precise near-infrared radial velocities that iSHELL's methane cell implements."},{"cited_title":"T., & Eastman, J","cited_arxiv_id":null,"evidence_quote":"Provides the barycentric correction code used to shift residuals into the stellar rest frame during template retrieval."},{"cited_title":"2012, in American Astronomical Society Meeting Abstracts, V ol","cited_arxiv_id":null,"evidence_quote":"Provides the 4.515 m/s per year secular acceleration subtracted from Barnard's Star velocities before computing long-term precision."}],"review_version":1}