{"id":"39f0bd21-9713-4634-9d3a-131f15cbdc17","arxiv_id":"2501.05602","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A Gaussian-process likelihood-ratio test, calibrated by simulations from the noise-only model, detects (quasi-)periodic signals in irregularly sampled red-noise light curves.","lead":"The paper presents a Gaussian-process method for finding periodic signals in unevenly sampled astronomical light curves while modeling the underlying red noise. It reanalyzes several claimed detections and finds lower significances than originally reported, and the method is released as a public Python package.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Calibration simulations skip the AICc triage and segment selection used on real data; reported p-values therefore condition on a data-chosen null, and a full-pipeline null simulation is needed.","rationale":"The reader's weakest assumption already includes the requirement that the AICc-selected null model is close enough to the truth for simulations from its posteriors to reproduce the null distribution of the LRT statistic, and the reader's rationale explicitly flags the unquantified effect of data-driven segment selection and AICc kernel triage. My stress-test pass reaches the same conclusion from the procedure itself: the calibration simulations in Section 3.1 use a known, pre-specified DRW null rather than the full model-selection and segment-selection pipeline applied in Section 3.2. This is not an internal inconsistency or a claim of misconduct; it is a missing calibration step. The paper deserves credit for public code, for showing that a pre-specified DRW null gives uniform p-values, and for honestly acknowledging the NGC 4945 segment-selection optimism. But the headline claim of well-calibrated significances for the full method as applied to real data is not yet established, because the p-values do not include the variability induced by choosing the null model and the segment from the data. This is exactly the kind of condition that a conditional acceptance should require: a full-pipeline simulation study under misspecified null models. My recommendation is therefore UNCHANGED relative to the reader's CONDITIONAL verdict; the concern strengthens the condition rather than overturning the paper.","tokens_in":35436,"tokens_out":5346,"duration_ms":56343,"concrete_test":"Generate roughly 1000 null light curves from red-noise processes deliberately outside the tested celerite family (e.g., a power-law PSD with beta = 1.5 or a broken power law), with the same cadence, gaps, and segment-selection rule used for NGC 4945 or P13. For each, run the complete pipeline: AICc kernel triage (Section 2.5), segment selection, then the LRT calibration from the selected null posteriors, and record the posterior predictive p-value. Test the empirical p-value distribution for uniformity, especially at thresholds 0.05 and 0.01. If the false-positive fraction exceeds the nominal level, the headline calibration claim fails because the selection step is not part of the simulated null.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the LRT reference distribution built from null-posterior simulations reproduces the sampling distribution of T_LRT under the null. For that to hold, the simulations must mirror the whole procedure used on the data. In the real applications the null model is not pre-specified: Section 2.5 selects it by AICc triage over many celerite combinations, and Sections 3.2.2 and 3.2.3 additionally select a stationary-looking segment and, for NGC 4945, a background fraction. The calibration runs in Section 3.1 instead simulate from the posteriors of a known DRW and add a Lorentzian; they never repeat the AICc triage, segment choice, or background assumption on each synthetic dataset. The paper itself shows the triage can pick a different combination than brute force (Section 2.5), reports competing nulls within Delta-AICc of about 2 (Tables 3-7), and notes the NGC 4945 significance is optimistic because the segment was chosen a posteriori (Section 3.3). Consequently, the quoted posterior predictive p-values are conditional on a data-chosen null and can be anti-conservative: the selection step adds trials and model uncertainty that the simulated T_LRT distribution does not contain. This is the weakest load-bearing link between the simulations and the claimed well-calibrated significances.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a Gaussian-process (GP) based method for detecting (quasi-)periodic components in irregularly sampled time series. The null hypothesis is an aperiodic GP model built from celerite kernels (DRW, SHO with Q=1/sqrt(2), Matérn-3/2, and jitter), while the alternative model adds a Lorentzian kernel. Significance is assessed with a likelihood-ratio test T_LRT (Eq. 1), whose reference distribution is constructed by simulating light curves from the posteriors of the null model and then refitting both the null and alternative models to each simulated data set. The recipe is applied to simulated DRW+Lorentzian light curves with varying cadence and baseline, to false-positive simulations from pure DRW noise, and to four real datasets (NGC 1365, NGC 7793 P13, NGC 4945, and B0537-441). The authors report good recovery rates in the simulated signal-present cases, p-values consistent with uniformity in the pre-specified-null false-positive simulations, and generally lower significances than several published Lomb-Scargle-based period claims.","tokens_in":35671,"tokens_out":7996,"duration_ms":76694,"significance":"If the calibration claim is correct, this paper offers a practical time-domain alternative to Lomb-Scargle period searches for irregularly sampled data with red noise. The strengths include a clear step-by-step recipe, a publicly available implementation (mind_the_gaps on GitHub/Zenodo), explicit treatment of heteroscedastic errors and irregular sampling, and a principled use of null-posterior simulations following Protassov et al. (2002). The applications to real datasets are instructive, and the authors are appropriately cautious in several of their claims, notably regarding NGC 4945. However, the calibration of the posterior predictive p-values is not yet established for the full analysis pipeline used on real data, because the simulations do not reproduce the model-selection and segment-selection steps. The central claim of well-calibrated significances therefore requires additional work before the method can be adopted as a standard tool.","major_comments":[{"comment":"The false-positive calibration is not a full-pipeline test. In the §3.1 simulations, each light curve is fitted with a pre-specified DRW null and a DRW+Lorentzian alternative; the AICc triage described in §2.5 is not repeated, nor are the segment-selection and background-fraction choices used in the real-data applications. Because the real-data null model is selected after seeing the data, the sampling distribution of T_LRT conditioned on that selection includes additional trials and model uncertainty that the simulations do not contain. The uniform p-values in Table 2 therefore support calibration only for a fixed null model, not for the data-driven null selection used in §3.2. A simulation that runs the complete pipeline (kernel triage, ΔAICc=2 competitors, segment choice, background assumption) on each synthetic data set is needed to support the central claim of well-calibrated significances.","section":"§3.1, §2.5, Table 2"},{"comment":"The NGC 4945 significance of ~98.7% is derived from null-posterior simulations of the SHO Q=1/sqrt(2) model on the selected 192-day segment. The segment was chosen a posteriori because the claimed signal was strongest there, and the authors explicitly note that the result may be optimistic (§3.3). This post-selection inflation is not included in the reference distribution. The remedy suggested in §3.3 — simulating full-length light curves and selecting the segment that maximizes the LRT for each simulation — is the appropriate procedure, but it is not implemented. Until it is, the quoted p-value should be presented as exploratory rather than calibrated.","section":"§3.2.3, §3.3"},{"comment":"For the P13 X-ray data, the best-fitting model has a standardised-residual KS p-value of 0.001, and the authors note that the count-rate distribution is non-Gaussian (KS p=0.008). Despite this, the hierarchical LRT significances (99.99%, 99.2%, 91.5%) are computed from simulations that assume the GP null is a correct description of the noise. The use of a lognormal PDF in the simulations (Appendix A) changes only the flux distribution, not the conditional GP likelihood used in the fits. The paper states that 'the model may not capture the full complexity of the data', but the p-value calibration requires the null model to be adequate. A calibration study for non-Gaussian or misspecified noise is needed before these significances can be taken at face value.","section":"§3.2.2, Table 5"},{"comment":"The model-selection procedure is applied inconsistently for the blazar B0537-441. The lowest-AICc models in Table 7 (e.g., Matérn-3/2+SHO Q=1/sqrt(2), AICc=3526.2) are discarded because their residual diagnostics reject Gaussianity, and the adopted Lorentzian+DRW alternative has AICc=3573.9, nearly 48 units worse. Meanwhile the DRW-only null has AICc=3639.9 and an acceptable residual p-value of 0.27. The paper does not specify how residual diagnostics and AICc are to be combined in the model-selection step. Since the claimed QPO significance (99.98%) is conditional on this particular null choice, and that choice is not the AICc-preferred model, the reported significance is not covered by the calibration simulations in §3.1.","section":"§3.2.4, Table 7"}],"minor_comments":[{"comment":"There are numerous typographical errors, including 'standarized' for 'standardized', 'Lorentizan' for 'Lorentzian', 'alternations' for 'alternatives', 'skweness' for 'skewness', and 'fullfilled' for 'fulfilled'. A careful proofreading pass is needed.","section":"Throughout"},{"comment":"The caption of Figure 12 says the segment is 'shown in Figure12' but the relevant light curve is shown in Figure 11; the cross-reference should be corrected.","section":"Figure 12 caption"},{"comment":"The caption refers to 'ΔAIC' while the text and table body use 'ΔAICc'; the notation should be made consistent.","section":"Table 3 caption"},{"comment":"The description of the iterative AICc routine would benefit from a small pseudocode block or a more precise statement of when the routine terminates, since the paper reports cases where it fails to find the overall best combination and the real-data applications use a mixture of AICc and residual diagnostics.","section":"§2.5"},{"comment":"The sentence beginning 'We found the Lorenzian component to be significant' should read 'Lorentzian'; this typo appears in several places and can cause confusion with the Lorenz curve.","section":"§3.2.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is potentially a valuable methods contribution, and the authors have made the code publicly available. The main obstacle is the missing full-pipeline calibration: the simulations in §3.1 do not include the AICc triage, segment selection, or background-fraction choices that are applied to real data. This is fixable within the scope of a revision by adding a simulation study that repeats the entire analysis pipeline on synthetic data, including the model-selection and post-selection steps. I do not see a need for rejection, but the central calibration claim cannot be accepted without this additional evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThis paper does something concrete and useful: it replaces Lomb-Scargle periodogram significance tests with a time-domain Gaussian-process likelihood-ratio test, calibrated by simulating from the null-model posteriors. That is a sensible adaptation of Vaughan (2010) and Protassov et al. (2002) to unevenly sampled data. The method is clearly described, the code is public, and the authors test it on a range of simulated cadences and real objects, including reanalyses of claimed QPOs. Those reanalyses are the most valuable part: they find NGC 1365 is only ~1.7 sigma, NGC 4945 ~2.5 sigma in the chosen segment, and they support the P13 UV/X-ray period difference. The false-positive simulations with DRW noise and gaps give uniform p-values, which is encouraging.\n\nThe main soft spot is exactly what the stress-test note says: the calibration simulations don't include the full analysis pipeline. On real data, the null model is selected by AICc triage, segments are chosen by eye, and for NGC 4945 a background fraction is assumed. The Section 3.1 simulations fit a known DRW and add a Lorentzian; they never repeat the model selection or segment choice on each synthetic dataset. So the quoted posterior-predictive p-values condition on a data-chosen null and can be anti-conservative. The paper itself acknowledges the segment-selection problem for NGC 4945 and notes the AICc triage can miss the best combination, but it doesn't quantify the effect on p-values. That is a load-bearing gap for the claim of well-calibrated significances.\n\nMinor issues: the P13 period difference is called significant based on non-overlapping uncertainties (roughly 2.4 sigma), which is suggestive but not a formal test. The residual diagnostics reject the GP model for the full NGC 4945 light curve and for the P13 X-ray data, and the paper is honest about that but still uses those fits. The lognormal appendix is a nice robustness check; parameter recovery holds up even for F_var = 0.6, which is reassuring.\n\nOverall, this is a serious methods paper with real potential. It deserves a proper referee. The referee should ask for a full-pipeline null simulation - repeat AICc triage and segment selection on synthetic data - to either validate or correct the p-value calibration. If that comes out okay, this will be a standard reference for GP-based periodicity searches.","headline":"Useful GP-based LRT recipe for periodicity in unevenly sampled red-noise light curves, with valuable reanalyses; the p-value calibration is plausible but conditional on the data-chosen null, and the paper deserves review.","tokens_in":36274,"tokens_out":1861,"would_cite":true,"duration_ms":17310,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["95.75.Wx"],"model":"deepseek-v4-flash","headline":"Periodicity searches in unevenly sampled light curves can be done in the time domain with Gaussian-process likelihoods, fitting the red noise from the data and calibrating signal significance by simulation rather than relying on…","keywords":["quasi-periodic oscillations","Gaussian processes","Lomb-Scargle periodogram","unevenly sampled time series","likelihood ratio test","red noise","celerite kernels","posterior predictive p-value"],"falsifier":"Generate a large ensemble of noise-only light curves from a process the kernel family cannot represent (for example, a broken power-law PSD with high-frequency slope shallower than $-2$, or a non-stationary process whose PSD slope drifts over the baseline) and apply the method to each: if the recovered p-values deviate from a uniform distribution, the calibration is not valid for that noise type.","tokens_in":35162,"feed_emoji":"🔭","tokens_out":15231,"duration_ms":120314,"temperature":0.7,"pith_summary":"This paper claims that the significance of a (quasi-)periodic signal in an irregularly sampled light curve can be computed reliably by working entirely in the time domain with Gaussian processes. The method fits the aperiodic variability (the null hypothesis) with GP kernels, adds an exponentially decaying sinusoid as the periodic alternative, and calibrates the likelihood-ratio improvement by simulating light curves from the null model's posteriors, the same logic as the established frequency-domain likelihood-ratio test of Vaughan (2010) but valid when the periodogram's statistical properties are unknown. Demonstrated on simulations, the test recovers injected signals in a majority of light curves once enough cycles are observed and yields uniform false-positive rates even when gaps are present. Re-analysing four published QPO claims, the paper supports the ~64-day and ~65.6-day superorbital periods of NGC 7793 P13 and a ~4.8-day oscillation in the blazar B0537-441, but finds only ~1.7σ and ~2.5σ significance for the NGC 1365 and NGC 4945 detections. If right, this gives a generalizable recipe for period searches in the sparsely and unevenly sampled monitoring data that dominates much of time-domain astrophysics.","feed_headline":"Two of four claimed QPO detections survive a Gaussian-process recheck","feed_subtitle":"Fitting red noise from the data itself gives trustworthy period significances for unevenly sampled light curves.","key_machinery":"The machinery is the likelihood-ratio test statistic $T_{\\rm LRT} = -2\\ln(L_0/L_1)$ computed from Gaussian-process likelihoods directly in the time domain, with its null distribution obtained empirically rather than analytically. The null hypothesis is a GP model of the aperiodic variability built from sums of celerite kernels, the damped random walk, the SHO with $Q=1/\\sqrt{2}$, the Matérn-3/2 approximation, and a Jitter white-noise term, whose PSDs are bending power laws; the alternative adds an exponentially decaying sinusoid (a Lorentzian PSD) for the (quasi-)periodic component. Model selection is done with the AICc, and the calibration step follows the posterior-predictive scheme of Protassov et al. (2002): draw hyperparameters from the null posterior, simulate light curves with the observing cadence and Poisson noise of the data, refit with the null and alternative models, and compare the observed $T_{\\rm LRT}$ to the simulated distribution. The celerite formalism is what makes the repeated refitting feasible, reducing the GP cost from $O(N^3)$ to $O(NJ^2)$.","core_discovery":"The central claim is that a likelihood-ratio test computed from Gaussian-process fits in the time domain, with the aperiodic noise inferred from the data as the null hypothesis and calibrated by simulations drawn from that null model's posteriors, yields well-calibrated significances for an additional (quasi-)periodic component in unevenly sampled light curves. The paper states the procedure 'can be considered equivalent to the LRT approach proposed by Vaughan (2010), but adapted to deal with irregularly sampled data'. The null hypothesis is a combination of celerite kernels (damped random walk, damped harmonic oscillator with $Q=1/\\sqrt{2}$, an approximate Matérn-3/2, and a Jitter white-noise term) selected by the AICc; the alternative adds a Lorentzian component, an exponentially decaying sinusoid. Significance is obtained by drawing kernel parameters from the null posteriors, simulating light curves with the same sampling and noise properties, refitting each with the null and alternative models, and locating the observed $T_{\\rm LRT} = -2\\ln(L_0/L_1)$ in the resulting reference distribution. On simulated data the false-positive rates are consistent with uniform p-values, and on four published QPO claims the method confirms two (P13 and B0537-441) while downgrading two (NGC 1365 and NGC 4945) relative to their published significances.","pith_inferences":["The calibration's validity for any new source hinges on the kernel family's coverage of the true noise; light curves whose PSD is not representable by bending-power-law kernels (for example, broken power laws with high-frequency slopes shallower than $-2$) should be stress-tested before the recipe is applied.","The iterative model-selection routine is acknowledged in the paper to occasionally miss the globally preferred combination, so a more exhaustive search would be needed before using the method for blind large-scale surveys, where the trials factor is no longer automatically absorbed by the simulations.","Because the procedure tests whether an added Lorentzian improves the fit, a non-sinusoidal periodicity (like the P13 harmonics) is tested one component at a time; tying the harmonics into a single alternative model would give a more powerful combined test.","The documented failures of the GP residual diagnostics on some real datasets (the full NGC 4945 light curve and the P13 X-ray data) mark the practical boundary of the method: when residuals are non-Gaussian or the process is non-stationary, the quoted p-values should be treated cautiously even if PSD recovery remains acceptable."],"forward_implications":["Published QPO significances based on Lomb-Scargle peaks above a red-noise continuum can be substantially overestimated; in the paper's re-analysis, two of the four claims drop below the conventional detection threshold.","Period searches in sparsely and unevenly sampled monitoring campaigns (ULXs, AGN, TESS light curves) gain a well-defined likelihood, making significance statements and parameter uncertainties computable without binning the data.","Because the noise is inferred from the data rather than assumed, the recipe extends to any source whose variability can be described as a Gaussian process, including other choices of mean function and priors.","False-positive tests on noise-only light curves, with and without gaps, yield uniformly distributed p-values, indicating the calibration controls spurious detections at the explored cadences.","Requiring roughly five or more observed cycles for a reliable detection matches earlier results and sets a practical limit on what sparse monitoring can claim."],"supporting_citations":[{"why":"Supplies the likelihood-ratio test framework for periodic components that this method adapts to irregular sampling.","marker":"Vaughan (2010)"},{"why":"Provides the posterior-predictive calibration scheme: simulating from the null posteriors to build the reference distribution of the test statistic.","marker":"Protassov et al. (2002)"},{"why":"Supplies the celerite kernels and the $O(NJ^2)$ computational shortcut that make repeated GP refits feasible.","marker":"Foreman-Mackey et al. (2017)"},{"why":"Provides the inverse-Fourier light-curve simulation algorithm used to generate the null-hypothesis datasets.","marker":"Timmer & Koenig (1995)"},{"why":"Extends the simulations to light curves with arbitrary (e.g. lognormal) flux PDFs and a given PSD.","marker":"Emmanoulopoulos et al. (2013)"},{"why":"Provides the Gaussian-process formalism: the likelihood, kernels, and the Fourier link between covariance and PSD.","marker":"Rasmussen & Williams (2006)"},{"why":"Documents the statistical pathologies of Lomb-Scargle periodograms under irregular sampling that motivate the time-domain approach.","marker":"VanderPlas (2018)"},{"why":"Precedent for time-domain PSD modelling with CARMA and source of the standardized-residual and ACF diagnostics.","marker":"Kelly et al. (2014)"}],"fun_headline_variants":["GP likelihood test outperforms periodogram for uneven data","Only half of QPO claims survive GP-based reanalysis","Gaussian processes deliver reliable period significances","Period searches get a GP upgrade for messy light curves","Uneven sampling? GP method finds true periodicities"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The calibration stands or falls on the assumption that the true aperiodic variability is a stationary Gaussian process whose covariance is captured by the tested kernel family, so that light curves simulated from the null posteriors reproduce the null distribution of the likelihood-ratio statistic.","fun_headline_variants_meta":{"raw":{"variants":["GP likelihood test outperforms periodogram for uneven data","Only half of QPO claims survive GP-based reanalysis","Gaussian processes deliver reliable period significances","Period searches get a GP upgrade for messy light curves","Uneven sampling? GP method finds true periodicities"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000282,"raw_usage":{"total_tokens":1738,"prompt_tokens":1086,"completion_tokens":652,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":702,"completion_tokens_details":{"reasoning_tokens":577}},"tokens_in":702,"tokens_out":652,"duration_ms":8029,"temperature":1.0,"reasoning_tokens":577,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:14:01.233699+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate a large ensemble of noise-only light curves from a process the kernel family cannot represent (for example, a broken power-law PSD with high-frequency slope shallower than $-2$, or a non-stationary process whose PSD slope drifts over the baseline) and apply the method to each: if the recovered p-values deviate from a uniform distribution, the calibration is not valid for that noise type.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the inverse-Fourier light-curve simulation algorithm used to generate the null-hypothesis datasets."},{"cited_title":"E., Williams C","cited_arxiv_id":null,"evidence_quote":"Provides the Gaussian-process formalism: the likelihood, kernels, and the Fourier link between covariance and PSD."}],"review_version":1}