{"id":"721c05e1-7ee6-4270-94a7-3586001358fc","arxiv_id":"2504.13262","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"TESS stellar rotation periods are accurate for most stars spinning faster than once per 10 days, but become unreliable beyond about 12 days, and stitching multiple sectors does not fix the problem.","lead":"The authors tested the accuracy of star rotation periods measured from TESS satellite data by comparing them to periods from the earlier K2 mission for about 23,000 overlapping stars, using K2 as a benchmark. They found TESS periods are reliable out to about 10 days and lose reliability beyond roughly 12 days, and they released a tool for choosing which measurements to trust.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Benchmark dependence: reliability numbers inherit any RH20 K2 period errors; the internal re-analysis covers only a favorable subset, leaving the bias from benchmark mistakes unquantified.","rationale":"The reader's weakest-assumption identification matches mine: the RH20 benchmark is the load-bearing element of the calibration. I agree with the CONDITIONAL verdict. The argument is not failing; Appendix A's independent benchmark and the authors' own subset re-analysis provide real supporting evidence, and the underlying conclusions are probably right. However, the benchmark error rate is only bounded in a favorable subset, while the headline reliability numbers are quoted without that caveat for the full sample, including faint, long-period, low-power stars. The proposed test is expensive but decisive and is simply the same check the authors already ran on a subset. I also flag the abstract-versus-Equation 1 precision discrepancy because it is an internal inconsistency in a headline number; it does not change the verdict but should be corrected. I do not see a more load-bearing concern than the benchmark error rate: the CPM detrending assumption is secondary, the sample selection is well documented, and the stitching conclusion is consistent with the presented data. The public code and ROC analysis strengthen the paper. Verdict remains CONDITIONAL because the benchmark-error quantification is missing and the abstract precision statement is internally inconsistent.","tokens_in":23784,"tokens_out":2037,"duration_ms":19610,"concrete_test":"Repeat for a random sample of at least 1,000 RH20 stars the full re-analysis that the authors ran only on mismatches, spanning the entire benchmark range (all LS powers, T magnitudes, and periods out to 40 days). Use independent K2 light curves (e.g., K2SFF) and two period-finders (Lomb-Scargle and autocorrelation) to re-derive each period without assuming RH20. If the RH20-vs-re-derived mismatch rate exceeds 1-2% in the low-power, faint, or long-period bins, the reliability curves in Figures 6-7 must be corrected by that amount and the central claim needs that caveat. Separately, recompute the abstract precision statement from Equation 1: at Prot = 10 days it gives about 5.8%, not below 3%, so the '<3%' claim should be restricted to Prot < 5 days or the abstract should be revised to match the equation.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central calibration treats RH20 K2 periods as ground truth (Sections 2, 4.1, and 7). If a measurable fraction of RH20 periods are wrong or systematically biased, every reliability and completeness fraction shifts, and the empirical uncertainty relation in Equation 1 absorbs those errors as if they were TESS noise. The paper's own re-analysis (Section 4.2) found RH20 wrong for roughly 1% of the overlap sample, but that check was performed on a random subset of mismatches in a high-LS-power, short-period, bright-star regime. It does not constrain the RH20 error rate among faint, low-power, or long-period stars—precisely the stars that dominate the reported reliability drop beyond 10 days. Appendix A partially mitigates the concern by reproducing the same qualitative trend with a Kepler benchmark (Reinhold et al. 2013), but the paper does not quantify the Kepler-vs-TESS mismatch rate in the same reliability framework. Thus the headline '70-80% reliable out to 10 days' is a conditional upper limit whose uncertainty is asserted rather than measured. A secondary internal inconsistency: the abstract states uncertainties are 'typically below 3% for periods < 10 days,' but Equation 1 yields about 5.8% at 10 days; the abstract claim is only accurate for periods below about 5 days.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper quantifies the reliability, completeness, and precision of stellar rotation periods derived from TESS light curves, using a cross-matched sample of roughly 23,000 stars observed by both TESS and K2. Light curves are extracted with the unpopular causal pixel model (CPM) and rotation periods are measured with a Lomb-Scargle periodogram, with RH20 K2 periods treated as the benchmark truth. The main results are an empirical fractional period uncertainty relation (Equation 1), a match-based reliability metric (Equation 2), and completeness estimates as functions of period, Lomb-Scargle power, TESS magnitude, and signal-to-noise ratio. The authors report that single-sector TESS periods are roughly 70-80% reliable out to 10 days, with uncertainties below 3% for periods under 5 days, and that reliability and completeness drop sharply beyond about 12 days. They also find that stitching consecutive TESS sectors reduces period uncertainties but does not improve reliability or completeness. The paper includes an application to three young associations and an appendix repeating the reliability analysis with a Kepler benchmark (R13).","tokens_in":24038,"tokens_out":2691,"duration_ms":24556,"significance":"If the results hold, this paper provides a useful empirical calibration of TESS rotation period measurements that many stellar rotation and gyrochronology studies can adopt. The study has notable strengths: a large overlap sample, a carefully documented pipeline based on CPM light curves and Lomb-Scargle periodograms, a second independent benchmark check against Kepler rotation periods in Appendix A, and release of code for computing reliability and completeness for user-defined cuts. The framework connecting reliability and completeness to explicit quality cuts is directly applicable to ongoing TESS-based surveys. The central qualitative conclusions, especially the sharp drop in reliability beyond roughly 10-12 days and the modest gain from stitching sectors, are physically expected and appear robust. The quantitative headline numbers, however, need correction and additional sensitivity testing, as detailed below.","major_comments":[{"comment":"The abstract's claim that uncertainties are \"typically below 3% for stars with periods < 10 days\" is inconsistent with Equation (1), which gives roughly 5.8% at 10 days; the text itself states that uncertainties are below 3% only for Prot < 5 days. Please revise the abstract and any summary statements so that all quoted uncertainty numbers agree with Equation (1) and Figure 4.","section":"Abstract and Section 4.1, Equation (1)"},{"comment":"The match criterion in Equation (2) uses a 3-sigma window derived from the empirical uncertainty relation of Equation (1), which is itself fitted to the same TESS-K2 comparison data. This makes the reported reliability fractions partly self-referential: outliers that inflate the fitted sigma widen the matching window, potentially masking failures. Please quantify the sensitivity of the headline reliability values to alternative match definitions, such as a fixed fractional tolerance (e.g., 10% or 20%) or uncertainties taken from the TESS-TESS comparison, and report how much the 70-80% reliability claim changes.","section":"Section 4.2, Equation (2)"},{"comment":"The analysis treats RH20 K2 periods as ground truth, but the internal re-analysis reported in Section 4.2, which found RH20 to be wrong for roughly 1% of the overlap sample, was performed on a random subset of mismatches in a regime of high Lomb-Scargle power, short periods, and bright stars. It does not constrain the RH20 error rate among faint, low-power, or long-period stars, which are precisely the stars that dominate the reliability drop beyond 10 days. Appendix A reproduces the qualitative trend with a Kepler benchmark but does not quantify the mismatch rate in the same reliability framework. Please add a sensitivity test that perturbs a plausible fraction of benchmark periods, or restricts the analysis to the highest-confidence RH20 subset (e.g., HPeak > 0.5), and report how the headline reliability numbers change.","section":"Sections 2 and 4.2"}],"minor_comments":[{"comment":"The caption ends with \"using the relation of .\" followed by a blank; this appears to be an incomplete reference and should be fixed.","section":"Figure 10 caption"},{"comment":"Several references are duplicated in the bibliography, including Curtis et al. (2020), Douglas et al. (2019), and Rampalli et al. (2021a); please consolidate duplicate entries.","section":"References"},{"comment":"The phrase \"It also exudes rapidly-rotating stars\" should read \"It also excludes rapidly-rotating stars.\"","section":"Section 7.2"},{"comment":"The name \"Vowell\" is typeset as \"V owell\" in the reference list; please correct the spacing.","section":"References"},{"comment":"The legend labels such as \"1 (TESS-RH20)\" appear to contain a typographical artifact; the label should presumably read \"sigma (TESS-RH20)\" for clarity.","section":"Figure 4"}],"recommendation":"major_revision","confidential_remarks":"The paper is well within the scope of a journal like AJ or ApJ and addresses a practical need in the TESS rotation community. The central conclusions are plausible, but the abstract inconsistency and the self-referential match criterion require attention before publication. The requested sensitivity tests should be feasible with the released code and sample."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, this is a useful calibration paper. It gives the community what it needs: empirical reliability and completeness maps for TESS rotation periods on a 23,000-star K2-TESS overlap, a public code to compute per-target reliability, and a second benchmark (Kepler, Appendix A) that reproduces the main trend. That the main conclusions survive a switch from K2 to Kepler is the strongest evidence in the paper that the drop in reliability beyond ~12 days and the lack of gain from stitching sectors are real.\n\nThe paper's central claim—single-sector TESS rotation periods are 70-80% reliable out to 10 days without extra cuts, and reliability collapses beyond ~12 days—is convincing. The sample is large, the light-curve extraction is standard (unpopular/CPM), and the authors are candid about what they cannot measure.\n\nSoft spots, in order of importance. First, the abstract states uncertainties are 'typically below 3% for stars with periods < 10 days.' Equation 1 gives about 5.8% at 10 days; the text elsewhere says below 3% only under 5 days. The abstract overstates precision and should be corrected. Second, the reliability numbers are conditional on the K2 benchmark (RH20) being correct. If a measurable fraction of RH20 periods are wrong or biased, the reliability fractions shift. The authors re-checked a random subset of mismatches and found RH20 wrong in ~1% of those, but that subset is bright, short-period, high-LS-power—exactly the favorable regime. The Kepler appendix confirms the trend qualitatively but does not quantify the Kepler-vs-TESS mismatch in the same reliability framework. So the '70-80% reliable' headline is a conditional upper limit. I would not call it load-bearing—the Kepler cross-check plus the period-dependent trend make the qualitative conclusion robust—but the reported fractions carry unknown benchmark uncertainty. Third, there is mild circularity: the match criterion uses the uncertainty relation fitted from the same TESS-K2 comparison. That inflates the self-consistency of the fractions; it doesn't invalidate the period-dependent drop, but it should be acknowledged. Fourth, a missing citation in Section 5 ('the relation of .') is a trivial fix.\n\nThis paper is for anyone using TESS rotation periods for gyrochronology, cluster membership, or occurrence studies. It deserves a serious referee. I would send it to review, with the abstract fix and a request to either propagate RH20 benchmark uncertainty or broaden the Kepler analysis into the same framework.","headline":"A solid, practically useful calibration of TESS rotation periods against K2 and Kepler, with the main caveat that the headline reliability numbers inherit any errors in the K2 benchmark and the abstract overstates precision at 10 days.","tokens_in":24606,"tokens_out":2427,"would_cite":true,"duration_ms":20948,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"TESS rotation periods are 70–80% reliable out to 10 days, then collapse beyond 12 days.","keywords":["stellar rotation","TESS","K2","rotation period reliability","Lomb-Scargle periodogram","causal pixel model","gyrochronology","light curves"],"falsifier":"Verify the benchmark itself: take a random subset of the K2-TESS overlap stars, measure their rotation periods from long-baseline ground-based photometry or an independent K2 pipeline, and compare. If the independently verified stars show that below-10-day TESS periods match truth less than about 70 percent of the time—or that more than a few percent of the RH20 periods are wrong—the central reliability claim would need substantial revision.","tokens_in":1565,"feed_emoji":"🔭","tokens_out":2450,"duration_ms":79991,"temperature":0.7,"pith_summary":"This paper asks when a stellar rotation period measured from TESS light curves can be trusted, and it answers with an empirical calibration against K2, a prior space mission with longer observing windows. Using roughly 23,000 stars observed by both missions, it finds that a standard Lomb-Scargle period analysis of single-sector TESS data recovers the K2 rotation period 70–80 percent of the time for periods out to 10 days, even without quality cuts, with fractional uncertainties below 3 percent for periods under 5 days. Beyond about 12 days—roughly half a TESS sector—reliability collapses, many detections become half-period aliases, and stitching consecutive sectors does not restore accuracy. The payoff is a quantitative map of reliability and completeness as a function of rotation period, signal strength, brightness, and signal-to-noise ratio, which lets studies of stellar ages and young associations decide which periods to trust and how to count non-detections.","feed_headline":"TESS rotation periods are 70-80% reliable out to 10 days","feed_subtitle":"23,000 K2 stars show where TESS rotation periods can be trusted and where the 27-day window breaks them.","key_machinery":"The load-bearing apparatus is an empirical calibration set: 22,986 stars observed by both TESS and K2, with K2 rotation periods from Reinhold & Hekker (2020) treated as ground truth; after cuts on binaries, contamination, and completeness, 16,752 stars carry the analysis. On the TESS side, the pipeline is a causal pixel model (CPM) light curve, built with a non-parametric model of instrumental systematics using pixels outside the target aperture, followed by a Lomb-Scargle periodogram. The match criterion—a TESS period within $3\\sigma$ of the K2 period using the fitted fractional uncertainty relation—defines reliability, and completeness counts matches against the full sample. These definitions turn raw period measurements into reliability and completeness maps as functions of period, Lomb-Scargle power, TESS magnitude, and signal-to-noise ratio.","core_discovery":"On the paper's own terms, the discovery is that TESS rotation periods extracted with a causal pixel model and a Lomb-Scargle periodogram are empirically calibrated quantities: they are accurate to about 70–80 percent reliability below 10 days, degrade sharply near 12 days, and are barely better than random beyond 15 days. The fitted single-sector fractional uncertainty is below 3 percent for periods under 5 days and grows roughly linearly to about 6 percent at 12 days, following $\\sigma(\\%) = 0.005577\\,P_{\\rm rot} + 0.001768$. There is a systematic bias of about 10 percent toward too-short periods in the 10–14 day range, attributed to the 27-day sector window. Stitching sectors reduces period uncertainty by up to a factor of two at long periods but does not improve reliability or completeness, because persistent systematics such as the 13.7-day scattered-light signal are reinforced by merging.","pith_inferences":["If the reliability map transfers to TESS-only samples, then catalogs of TESS rotation periods should report more than a single period: each star needs a reliability and a completeness value, and age or membership analyses should marginalize over aliased and missed periods.","A natural extension is to use TESS continuous-viewing-zone stars observed across many sectors to map how systematics grow with sector count; the paper does not do this, but such a map could predict exactly when stitching starts to hurt.","The 13.7-day scattered-light signal and its 6.85-day half-alias likely explain part of the inflated error at 6–8 and 10–14 days, so testing whether removing scattered-light contamination before stitching raises reliability would be a direct follow-up.","Because machine-learning period finders are improving long-period recovery, the reliability maps in this paper provide a clean benchmark for deciding whether such methods genuinely beat a Lomb-Scargle periodogram beyond 12 days."],"forward_implications":["Single-sector TESS rotation periods can be used as gyrochronology inputs for periods below 10 days, with Lomb-Scargle power thresholds setting the trade-off between reliability and completeness.","Periods measured beyond about 12 days should not be treated as secure detections, since many are half-period aliases; long-period TESS-only rotation statistics need priors or independent confirming data.","Stitching TESS sectors buys precision but not accuracy, so studies seeking slow rotators should analyze sectors separately and keep the highest-power period rather than merging light curves.","The fitted uncertainty relation and the released code let any user assign per-star period errors and compute reliability or completeness for arbitrary cuts on signal power, brightness, and signal-to-noise ratio.","Applied to young associations, the reliability map identifies roughly 7–14 unreliable rotation measurements per cluster and predicts about 5–17 missed detections per cluster, so cluster rotation sequences built from TESS alone should carry these probabilities."],"supporting_citations":[{"why":"Supplies the benchmark K2 rotation periods and the HPeak quality threshold used to define the ground-truth sample.","marker":"Reinhold & Hekker 2020"},{"why":"Provides the unpopular causal pixel model code that builds the variability-preserving TESS light curves.","marker":"Hattori et al. 2021"},{"why":"The Lomb-Scargle periodogram is the period-measuring method whose reliability and completeness are being calibrated.","marker":"Lomb 1976; Scargle 1982"},{"why":"TESS-point determines which K2 targets fall in TESS sectors, constructing the overlap sample.","marker":"Burke et al. 2020"},{"why":"In the appendix, its Kepler rotation periods serve as an independent benchmark to confirm the same reliability trends.","marker":"Reinhold et al. 2013b"}],"fun_headline_variants":["TESS rotation periods: reliable to 10 days, not beyond 12","TESS spins: accurate under 10 days, unreliable after 12","K2-TESS overlap: rotation periods trusted only to ~10 days","TESS stellar rotation: drops off after 12-day mark"],"cache_read_input_tokens":26752,"weakest_assumption_plain":"The calibration treats the K2 rotation periods of Reinhold & Hekker (2020) as the true rotation periods; if a substantial share of those benchmark values are wrong or systematically biased, every reliability and completeness number in the paper shifts.","fun_headline_variants_meta":{"raw":{"variants":["TESS rotation periods: reliable to 10 days, not beyond 12","TESS spins: accurate under 10 days, unreliable after 12","K2-TESS overlap: rotation periods trusted only to ~10 days","TESS stellar rotation: drops off after 12-day mark"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00072,"raw_usage":{"total_tokens":3270,"prompt_tokens":1020,"completion_tokens":2250,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":636,"completion_tokens_details":{"reasoning_tokens":2172}},"tokens_in":636,"tokens_out":2250,"duration_ms":15577,"temperature":1.0,"reasoning_tokens":2172,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:12:58.402920+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Verify the benchmark itself: take a random subset of the K2-TESS overlap stars, measure their rotation periods from long-baseline ground-based photometry or an independent K2 pipeline, and compare. If the independently verified stars show that below-10-day TESS periods match truth less than about 70 percent of the time—or that more than a few percent of the RH20 periods are wrong—the central reliability claim would need substantial revision.","supporting_citations":[],"review_version":1}