{"id":"18d1191f-ce74-4453-bf96-6dcad11e257d","arxiv_id":"1909.01385","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"New error formulas for the energy-dependent cross-spectrum and rms spectrum account for the common reference band and intrinsic coherence, fixing over-fitting from older single-spectrum errors.","lead":"This paper derives new analytic error formulas for the energy-dependent cross-spectrum, a standard X-ray timing statistic, and shows that older formulas cause over-fitting when applied to energy-dependent measurements. The corrected errors matter for any study fitting spectral-timing models to black hole and neutron star X-ray data.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Modulus/phase error formulae rely on a small-signal linearization and likely fail for low intrinsic coherence, despite the claim of validity for any coherence.","rationale":"The paper's main contribution is a genuinely useful correction to error formulae for energy-dependent cross-spectra, and the real/imaginary-part and rms formulae (Eqs. 18/20/22) appear correct under the stated Gaussian model, as the Monte Carlo tests confirm. However, the strongest claim (accepted by the reader) explicitly says the formulae are valid 'for any value of intrinsic coherence', including for modulus and phase. The derivation of the modulus error is not exact: it linearises the radial direction on the complex plane, which is valid only when |G| is large compared to the statistical scatter. At low or zero coherence the distribution of the modulus is Rayleigh/Rician, and its standard deviation differs from σ by a factor of order 1.5 or more; the phase error diverges while the true phase becomes uniform. The paper's simulations and observational example both use high-coherence, high signal-to-noise cases, so this regime is never tested. This is an internal limitation of the argument rather than a disagreement with external consensus: even if every Gaussian assumption is granted, the modulus/phase claim fails for small |G|/σ. Therefore the reader's ACCEPT should be conditional: either the claims should be restricted to the high-coherence/high-S/N regime where linearisation holds, or the paper should add a validation and possibly a corrected or qualified formula for the low-coherence modulus. The real/imag and rms parts of the central claim, and the practical recommendation for high-coherence X-ray data, remain intact.","tokens_in":19005,"tokens_out":19422,"duration_ms":205621,"concrete_test":"Run the Section 3 Monte Carlo with γ = 0, 0.05, and 0.1 while keeping all other inputs fixed; compare reduced-χ² histograms for the modulus error bars in Eq. (18) against the expected χ²(49) distribution.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim covers the modulus and phase for any intrinsic coherence, but the derivation of d|G| = dRe = dIm in Section 4.2 uses a circular-contour argument valid only when the true cross-spectrum is large compared with its statistical scatter. For low or zero intrinsic coherence, |G| is small (or zero) and the complex estimate is centred near the origin; the modulus then follows a Rayleigh/Rician distribution whose standard deviation is not equal to the standard deviation of the real part. Under the paper's own Gaussian model with γ = 0, G is a zero-mean complex Gaussian: the modulus has mean ≈ 1.253σ and standard deviation ≈ 0.655σ, where σ = dRe = dIm from Eq. (18). Thus Eq. (18) overestimates the modulus scatter by about a factor 1.5, and subtracting the bias b² in quadrature makes the debiased modulus even tighter (standard deviation ≈ 0.4σ). Consequently a constant modulus model fitted to 50 channels would give reduced χ² well below unity, i.e. the same over-fitting the paper claims to cure. The phase is also problematic at low coherence: Eq. (19) diverges as γ → 0, while the actual phase distribution becomes uniform on [0, 2π), for which a linear 1σ error is not meaningful. The Monte Carlo validation in Section 3 uses only γ = 0.6 with |G|/σ ≈ 27, a high signal-to-noise regime that does not probe this failure; the Cygnus X-1 example likewise has high coherence. The claim 'valid for any value of intrinsic coherence' is therefore not supported, and the modulus/phase error formulae are at best approximations valid when |G| is several times σ.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents analytic error formulae for the energy-dependent cross-spectrum (real part, imaginary part, modulus, phase) and for the rms spectrum, explicitly accounting for the use of a common reference band and for correlations between subject bands. The central claim is that equations (18), (19) and (22) give correct 1σ uncertainties for any intrinsic coherence, in contrast to the Bendat & Piersol (2010) formulae, which overestimate errors in the energy-dependent limit and cause overfitting. The paper validates the formulae with 50,000 Monte Carlo realizations under a Gaussian signal-plus-noise model, demonstrates the method on RXTE Cygnus X-1 data, discusses the optimal reference band and the bias correction for coherence, and provides a Fortran implementation.","tokens_in":19307,"tokens_out":9179,"duration_ms":91202,"significance":"If the real/imag and rms formulae are correct, they fill a practical gap: fitting energy-dependent timing spectra is increasingly common, and the paper correctly identifies why single-spectrum error bars cause overfitting. The Monte Carlo validation is extensive, the observational fit is consistent, the correction to the Uttley et al. (2014) rms error formula is valuable, and the treatment of broad reference bands is elegant. The main weakness is that the modulus and phase formulae are derived from a small-scatter linearization that fails at low intrinsic coherence, while the paper claims validity for any coherence; this needs to be fixed before publication.","major_comments":[{"comment":"The equality d|G| = dRe = dIm is derived from concentric circular error contours, which is only an approximation for a Gaussian cloud that does not contain the origin. Under the paper's own model with γ = 0, G̃ is a zero-mean complex Gaussian and |G̃| has a Rayleigh distribution with standard deviation ≈ 0.655 σ, where σ = dRe from Eq. (18); equation (18) therefore overestimates the modulus error by roughly a factor 1.5, and a constant-modulus fit to 50 channels would yield reduced χ² ≈ 0.43 rather than 1. The Section 3 simulation uses γ = 0.6 with |G|/σ ≈ 27, so it does not probe this regime, and the claim that the formulae are valid 'for any value of intrinsic coherence' (abstract and conclusions) is thus not supported.","section":"§4.2, Eq. (18)"},{"comment":"The phase error diverges as γ → 0 because the denominator |G̃|² − b̃² tends to zero, whereas the true phase distribution under the Gaussian model tends to a uniform distribution on [0, 2π), for which a linear 1σ error is not meaningful. Equation (44), taken from Bendat & Piersol, is a small-angle linearization and is not valid when the Gaussian cloud contains the origin; the paper should either restrict Eq. (19) to |G|/dRe ≫ 1 or provide a proper wrapped/Rician treatment.","section":"§4.2, Eq. (19)"},{"comment":"The assertion that the 'other 35 terms' have zero k covariance does not follow from the fact that the total k covariance equals the first term's contribution; the remaining terms could cancel. Since the conclusion Cov_n{Re[G̃], Im[G̃]} = 0 rests on this shortcut, the derivation needs a direct evaluation or an explicit symmetry argument for the cross terms.","section":"§4.2, around Eq. (46)"}],"minor_comments":[{"comment":"The phrase 'valid for any value of intrinsic coherence' should be qualified, since the modulus and phase formulae are only demonstrated for high coherence; the real/imag and rms formulae appear more general.","section":"Abstract and Section 7"},{"comment":"Please report the ratio |G|/σ for the simulated and observed data; without this, the reader cannot easily see that the validation is confined to the high-coherence, high-signal-to-noise regime.","section":"Section 3"},{"comment":"There are two typos: 'This is not indented as a means' should be 'intended', and 'logarithmicaly' should be 'logarithmically'.","section":"Section 5"},{"comment":"It would help to state explicitly that d|G| in Eq. (18) is a large-|G|/σ approximation rather than an exact equality for all coherence values.","section":"Section 2.4"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe headline: this is a genuinely useful methods paper that corrects a real over-fitting problem in X-ray spectral-timing, but the modulus and phase error formulas are not valid for arbitrarily low intrinsic coherence as claimed. The real and imaginary part formulas and the rms formula are solid; the modulus/phase formulas are fine only in the high-coherence, high-signal-to-noise regime.\n\nWhat's new: Ingram derives analytic errors for the energy-dependent cross-spectrum that correctly account for a common reference band and correlations between subject bands, which the standard Bendat & Piersol formulas don't. He also points out a factor-of-two error in the widely quoted rms error formula in Uttley et al. 2014. The broad-reference-band recipe with noise subtraction is elegant and removes the need to subtract each subject band from the reference. The Monte Carlo validation is extensive (50,000 realizations), and the chi-square distributions match theory for the cases tested. The Cygnus X-1 example is a nice demonstration. Code is provided. That's all real and worth having.\n\nWhere it's soft: the claim \"valid for any value of intrinsic coherence\" is not supported for modulus and phase. Section 4.2 derives d|G| = dRe = dIm by arguing that circular Gaussian error contours imply the same width in any direction. That is only true when the true cross-spectrum is several sigma from the origin. When gamma is low and |G| is comparable to or smaller than the scatter, the modulus follows a Rician/Rayleigh distribution whose standard deviation is noticeably smaller than the real/imaginary part error — for gamma=0 it's about 0.655 sigma versus sigma. The phase formula diverges as gamma goes to 0 while the actual phase distribution becomes uniform, so quoting a linear 1-sigma phase error is meaningless there. The Monte Carlo tests use gamma=0.6 with |G|/sigma ~ 27, which never probes this regime. The observational example is also high-coherence. So the abstract's \"any value of intrinsic coherence\" overclaims; the modulus and phase formulas are high-coherence approximations.\n\nThat said, the real and imaginary part formulas (Eqs. 18) and the rms formula (Eq. 22) appear correct for any coherence — the derivation for Re/Im is a straightforward variance sum and doesn't rely on the problematic contour argument. And in most practical X-ray timing applications, coherence is high enough that the modulus/phase formulas are fine. But anyone fitting low-coherence bands (e.g., high-frequency or faint channels) should be cautious.\n\nVerdict: worthy of serious peer review as a methods paper. The core statistical correction is important and the derivations are mostly sound. I'd ask the author to either prove the modulus/phase formulas for low coherence (which I doubt they can) or explicitly restrict the claim to |G| of order a few sigma and above, and to add a simulation at gamma=0 or gamma=0.1 to show the limits.\n\nBring it to reading group; it'll spark a good discussion about when Gaussian error propagation breaks down.","headline":"Useful and mostly correct fix for an over-fitting bug in X-ray spectral-timing, but the modulus/phase formulas overclaim 'any coherence' — they are high-coherence approximations.","tokens_in":19836,"tokens_out":4014,"would_cite":true,"duration_ms":37696,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The standard error bars used for energy-dependent X-ray cross-spectra are too large, and this paper derives the correct analytic replacements.","keywords":["cross-spectrum","error analysis","X-ray timing","coherence","rms spectrum","phase lag","Monte Carlo simulation","accreting compact objects"],"falsifier":"Simulate light curves with log-normal, non-Gaussian flux variability, apply equations (18), (19), and (22) with the same N=500 averaging, and check whether 1-sigma error bars bracket the true model in about 68 percent of energy channels and whether the reduced chi-squared distribution matches the theoretical chi-squared with Ne-1 degrees of freedom; systematic deviations would show that the Gaussian assumption is load-bearing.","tokens_in":18794,"feed_emoji":"📊","tokens_out":5920,"duration_ms":49305,"temperature":0.7,"pith_summary":"This paper derives analytic error formulae for the energy-dependent cross-spectrum and rms spectrum, the Fourier statistics used to study rapid X-ray variability from accreting compact objects. It shows that the textbook errors of Bendat & Piersol (2010), which were designed for a single cross-spectrum, overestimate uncertainties when many subject bands are referenced to one common band, causing systematic over-fitting. The new formulae cover the real part, imaginary part, modulus, phase, and rms spectrum, and remain valid for any intrinsic coherence between energy bands. They are no harder to implement than the old ones, and a worked Cygnus X-1 example shows they bring a reverberation-model fit from reduced chi-squared 0.76 to an acceptable value.","feed_headline":"Correct error bars for X-ray cross-spectra stop overfitting","feed_subtitle":"New formulas for phase, modulus, and rms work at any coherence, so model fits match their expected scatter.","key_machinery":"The central object is the energy-dependent cross-spectrum with a common reference band, modeled as complex Gaussian Fourier transforms with additive, uncorrelated noise. The argument splits the variance into a 'k variance' over frequency realizations and an 'n variance' over energy channels; in the energy-dependent limit, the shared reference band contributes zero n-variance to the real and imaginary parts, while subject-band and noise terms contribute $1/(2N)$ variances. This reduction makes the new errors on the real, imaginary, and modulus parts identical, and it sets the phase-error and rms-error formulae. The broad-reference-band correction subtracts the subject band's own Poisson noise contribution from the cross-spectrum.","core_discovery":"The central claim is that equations (18), (19), and (22) give the correct 1-sigma errors for the energy-dependent cross-spectrum and rms spectrum, in the limit where averaged Fourier estimates are Gaussian. The key is that the same reference band appears in every subject-band cross-spectrum, and that subject bands are correlated with one another, so reference-band and shared-signal fluctuations act as a common systematic rather than independent scatter. Consequently, fitting the correct model with these errors yields reduced chi-squared with expectation value unity, whereas Bendat & Piersol (2010) errors yield over-fitting. The formulae also handle a broad reference band built from the sum of subject bands by subtracting a small noise term, eliminating the need to excise the current subject band from the reference.","pith_inferences":["If the Gaussian-plus-uncorrelated-noise model is adequate, the same variance-splitting argument should apply to other correlated Fourier statistics, such as energy-dependent covariance spectra, by simple rescaling; the paper already gives the rescaling for complex covariance.","The formulae expose a clean separation between statistical scatter among energy channels and a normalization systematic coming from the shared reference band, which matters when model normalizations are tied across frequency ranges.","A natural stress test is to apply the formulae to simulated log-normal, non-linear variability; if coverage departs from nominal, the Gaussian assumption would need a multiplicative correction or a different noise model.","The new understanding of which fluctuations are systematic rather than statistical could sharpen frequentist model comparison in multi-band timing fits, not just single-model chi-squared."],"forward_implications":["Energy-dependent spectral-timing fits that use the new uncertainties will no longer be systematically over-fit: the expected reduced chi-squared for a correct model is unity, not below it.","A broad reference band formed by summing all subject bands can be used directly by subtracting the subject band's own Poisson noise contribution; the common practice of excising the subject band from the reference is unnecessary.","For rms spectra, the formula interpolates between the old single-band error at $γ=0$ and a corrected $γ=1$ limit that differs from a commonly quoted version by a factor $√2$.","The example fit to Cygnus X-1 with the reltrans reverberation model moves from reduced chi-squared 0.76 to 42.9/36, with best-fit parameters consistent within 1 sigma.","The new formulae also improve the expected sensitivity of X-ray polarimetry-timing searches for quasi-periodic oscillations in polarization degree and angle."],"supporting_citations":[{"why":"Provides the standard single cross-spectrum and power-spectrum error formulae; the paper's new formulae correct their over-fitting behavior in the energy-dependent limit.","marker":"Bendat & Piersol (2010)"},{"why":"Simulation algorithm used to generate complex Gaussian light curves for the Monte Carlo validation.","marker":"Timmer & Koenig (1995)"},{"why":"Standard review defining rms and covariance spectra and the common practice of removing the subject band from the reference band; its gamma=1 rms error formula is corrected by a factor sqrt(2).","marker":"Uttley et al. (2014)"},{"why":"Defines the complex covariance statistic whose error follows from rescaling equation (18).","marker":"Mastroserio et al. (2018)"},{"why":"The reltrans fit to Cygnus X-1 whose reduced chi-squared 0.76 is explained and corrected by the new formulae.","marker":"Mastroserio et al. (2019)"},{"why":"Justifies treating averaged Fourier estimates as Gaussian for N greater than about 40.","marker":"Huppenkothen & Bachetti (2018)"},{"why":"Earlier empirical gamma=1 rms error formula; the paper compares and corrects its conversion to the frequency domain.","marker":"Vaughan et al. (2003)"},{"why":"Earlier covariance-spectrum error formula; the paper identifies an extra term that should be omitted.","marker":"Wilkinson & Uttley (2009)"},{"why":"Dead-time model used to estimate the noise power on the Cygnus X-1 data in the observational test.","marker":"Zhang et al. (1995)"},{"why":"The reverberation model reltrans used for the re-fit of the Cygnus X-1 energy-dependent cross-spectrum.","marker":"Ingram et al. (2019)"}],"fun_headline_variants":["New error formulae stop overfitting in X-ray cross-spectra","Correct 1-sigma errors for energy-dependent cross-spectra","Analytic errors for cross-spectra end overfitting","Correct errors at any coherence stop overfitting"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The formulae assume that the measured Fourier transforms behave like complex Gaussian random variables with signal and noise adding independently, and that averages over at least about 40 realizations are Gaussian; real variability from accreting sources is known to be log-normal and non-linear.","fun_headline_variants_meta":{"raw":{"variants":["New error formulae stop overfitting in X-ray cross-spectra","Correct 1-sigma errors for energy-dependent cross-spectra","Analytic errors for cross-spectra end overfitting","Correct errors at any coherence stop overfitting"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000795,"raw_usage":{"total_tokens":3473,"prompt_tokens":890,"completion_tokens":2583,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":506,"completion_tokens_details":{"reasoning_tokens":2518}},"tokens_in":506,"tokens_out":2583,"duration_ms":16560,"temperature":1.0,"reasoning_tokens":2518,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T05:18:31.744580+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate light curves with log-normal, non-Gaussian flux variability, apply equations (18), (19), and (22) with the same N=500 averaging, and check whether 1-sigma error bars bracket the true model in about 68 percent of energy channels and whether the reduced chi-squared distribution matches the theoretical chi-squared with Ne-1 degrees of freedom; systematic deviations would show that the Gaussian assumption is load-bearing.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Simulation algorithm used to generate complex Gaussian light curves for the Monte Carlo validation."}],"review_version":1}