{"id":"c6100aac-525d-43dc-8354-668a82b8b812","arxiv_id":"2505.17346","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A Bayesian fitting framework with a nonlinear split-photodiode model improves extraction of two closely spaced decay constants in disk resonator ringdowns.","lead":"This paper proposes a Bayesian framework for extracting decay rates from mechanical ringdown measurements of coated disk resonators, using a refined model of how a Gaussian laser spot moves across a split photodiode. It reports up to 25% accuracy improvements in simulations and shows that ringdowns previously discarded as unfittable can be analyzed.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The likelihood in Eq. (24) treats FFT-peak ringdown samples as independent Gaussians; the authors' own Appendix B.1 shows this yields biased τ estimates in 4/50 simulations, including a τ1/τ2 swap, so the claimed superior accuracy is not robust for low-amplitude tails.","rationale":"The paper is a solid methodological contribution: it ships bayesbeat, validates on 50 simulated ringdowns, and demonstrates real-data fits for cases the prior method could not handle. The most load-bearing condition for the central claim is that Eq. (24) adequately describes the statistics of the FFT-derived data points. This condition is least secure because each point is a maximum of an FFT power spectrum, not a Gaussian sample, and the model omits the FFT noise floor. The authors' own results in Appendix B.1 provide direct evidence of failure: four biased estimates, including one complete τ1/τ2 swap, in cases where the signal falls below the noise floor. This is not an external disagreement with consensus; it is an internally documented limitation that directly targets the claimed superiority of the method for real ringdowns with low-amplitude tails. A concrete remedy—adding a noise-floor term or using a proper extreme-value likelihood—can be tested on the same simulated injections. If it removes the bias, the paper needs only a scope statement; if not, the abstract's quantitative claim requires substantial qualification. The reader's weakest assumption captures the same issue, and the conditional verdict remains appropriate; no change is needed.","tokens_in":20784,"tokens_out":5248,"duration_ms":43216,"concrete_test":"Re-analyze the 50 simulated ringdowns of §5 with a corrected likelihood that adds an additive noise-floor term (or uses the actual extreme-value distribution of the FFT peak) while keeping M3 and T=7 fixed. Check whether the four biased cases in Appendix B.1, including the τ1↔τ2 swap, move inside the 3-σ credible interval and whether the relative error reduction versus M1 remains consistent with the abstract's 'up to 25%' claim. If the bias persists, the central accuracy claim must be restricted to signal levels above the noise floor.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central accuracy claim depends on Eq. (24), which models each data point d(t_i) as an independent Gaussian with variance s(t_i)^2 ξA^2 + ξS^2. However, each d(t_i) is the maximum of an FFT power spectrum over a 0.2 s window (Eq. (2)). Its statistics are therefore those of an extreme value (Rician or χ²-type), not Gaussian, and they do not decay to zero when the signal vanishes: the FFT peak retains a noise floor. The authors document the consequence in Appendix B.1: in four of 50 simulated ringdowns the inferred τ values are biased when the signal falls below the RIN noise floor, with one case producing a full swap of τ1 and τ2. Section 7 explicitly concedes that the model does not account for the FFT noise floor. Because the real-data motivation includes six previously 'unfittable' ringdowns, many of which are single-decay and low-amplitude, the claimed superior estimation accuracy is least secure precisely in the regime where the statistical model is known to fail. The qualifier 'especially for larger oscillation amplitudes' narrows the scope, but the abstract's unconditional 'superior estimation accuracy' is not supported for the low-SNR tail, and the paper does not quantify the 'up to 25%' figure as an estimation-accuracy improvement over M1.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents a Bayesian framework for extracting the two decay constants tau1 and tau2 from GeNS ringdown measurements of coated disk resonators used in gravitational-wave coating loss studies. Three signal models are considered: a traditional two-exponential model (M1), a numerical split-photodiode model (M2) that includes the Gaussian beam profile and photodiode gap nonlinearity, and an analytic power-series approximation (M3) that enables fast likelihood evaluation. The likelihood in Eq. (24) combines stationary and amplitude-proportional (RIN-like) Gaussian noise. The models are validated on 50 simulated ringdowns generated with M2 and applied to 47 real ringdowns, with model comparison performed via nested sampling (nessai) and Bayes factors. The paper claims up to 25% improvement in estimation accuracy over traditional methods, especially for larger oscillation amplitudes, and reports that six previously unfittable ringdowns can now be analyzed without adjustment.","tokens_in":21088,"tokens_out":12493,"duration_ms":92589,"significance":"The framework addresses a genuine and consequential problem: photodiode nonlinearity and RIN noise create model misspecification that biases conventional exponential fits, and the paper offers a physically motivated correction with principled model comparison. The analytic M3 construction, the two-component noise model with standardized residual diagnostics, and the public implementation (bayesbeat) with a data release are notable strengths, as is the unusually candid documentation of failure modes (four biased simulations, one tau1/tau2 swap, and prior-width sensitivity in Appendix B.1). If the accuracy claims hold after the requested quantification, this is a useful contribution to the mechanical-loss and coating thermal noise community and to the instrumentation readership of Classical and Quantum Gravity.","major_comments":[{"comment":"The abstract's headline claim of 'improvements in estimation accuracy by up to 25%' is not supported by the accuracy metrics reported in the paper. The only 25% figure in the manuscript is in Section 3.2 (Figure 4), where it quantifies the difference between the predicted amplitudes of the T=1 and T=7 signal models, not the error in the inferred decay constants. The accuracy comparison in Figure 10 is reported qualitatively ('M3 is consistently closer to the true value'), with no percentage improvement, and the text also states that at lower amplitudes neither model is favoured. Please report a quantitative accuracy metric (e.g., median or 90th percentile of |tau_hat - tau_true|/tau_true across the 50 injections, binned by amplitude) or reword the abstract so that the 25% figure is attributed to model-output differences rather than to estimation accuracy.","section":"Abstract; Section 3.2 (Figure 4); Section 5.2 (Figure 10)"},{"comment":"The likelihood in Eq. (24) treats each ringdown sample as an independent Gaussian with variance s(t_i)^2 xi_A^2 + xi_S^2, but each d(t_i) in Eq. (2) is the maximum of an FFT power spectrum over a 0.2 s window; its distribution is an extreme-value statistic with a positive noise floor that does not vanish as s -> 0. The authors themselves document the consequence: Appendix B.1 reports biased tau estimates in four of the 50 simulated ringdowns when the signal falls below the FFT/RIN noise floor, including one case with a tau1/tau2 swap, and Section 7 concedes that the model does not account for the FFT noise floor. This is precisely the regime relevant to the claim in Section 6 that six previously discarded real ringdowns, described in Appendix B.3 as low-amplitude and often single-decay signals, 'can now be reliably analysed.' As it stands, the unconditional 'superior estimation accuracy' of the abstract is not established for these low-SNR tails. I request either a quantitative assessment of the bias for the affected real ringdowns (e.g., a posterior predictive check that includes a noise-floor term, or an analysis demonstrating that those ringdowns do not enter the low-SNR regime) or a scope-limited wording of the claim.","section":"Eq. (24); Eq. (2); Appendix B.1; Section 7"},{"comment":"The simulated-data validation is largely a self-consistency check within one model family: the injections are generated with M2 (Eq. (13)) and analyzed with M3 (Section 3.2), an analytic approximation of the same photodiode model. The recovery of the injected tau values therefore demonstrates internal consistency of the signal-model family rather than the adequacy of the signal model against an independent ground truth; the comparison that matters for the paper's improvement claim is M3 versus M1, and both share the same, possibly misspecified, Gaussian likelihood of Eq. (24). A concrete strengthening would be to inject simulated data using the extreme-value statistic of Eq. (2) including the noise floor, or to validate M3 against a measured photodiode response such as the scan in Figure 2b, which would test the absolute accuracy of the inferred decay constants rather than only the relative improvement over M1.","section":"Section 5.2; Eq. (13); Section 3.2"}],"minor_comments":[{"comment":"The two noise components are mislabeled in the text: 'The first, nA is stationary Gaussian noise' and 'The second, nA is amplitude dependent noise' both refer to nA, whereas in Eq. (22) the multiplicative amplitude-dependent component is nA and the stationary component is nS; the text should be corrected to match Eq. (22).","section":"Section 4.2, after Eq. (21)"},{"comment":"The condition 'xi_A < 0' appears twice ('allowing xi_A < 0 has minimal effects of the signal fit' and 'with only stationary noise versus (xi_A = 0) stationary and amplitude dependent noise (xi_A < 0)'); the intended condition is xi_A > 0, and the sign flip will confuse readers.","section":"Section 5.1 and Figure 5 caption"},{"comment":"As typeset, the first and third terms on the right-hand side of Eq. (17) are identical and opposite in sign, so the 3*omega_1 contribution cancels exactly; presumably one sign is a typographical error. Since the full expressions for M3 are only given in the bayesbeat code, this displayed equation should be corrected so that the derivation is checkable from the paper alone.","section":"Eq. (17)"},{"comment":"The phrase 'at lower amplitudes neither model is favoured (log10 B = 1)' is inconsistent with the Jeffreys scale defined in Section 4, where log10 B > 1 is strong evidence in favor of one model; the scatter in Figure 8 suggests values near zero at low amplitude, so this should presumably read 'log10 B is close to 0'.","section":"Section 5.2, near Figure 8"},{"comment":"The in-text claim that 'the landmark detection of GW150914_095045 in 2015 [1]' is supported by reference [1], which is the 2016 Living Reviews article by Abbott (arXiv:1304.0670), not the GW150914 discovery paper; the correct citation is Abbott et al., Phys. Rev. Lett. 116, 061102 (2016). In addition, the sentence ending '...calculated (see Figure 1). to the isotropy of the substrate...' appears to be missing the word 'Due' at the start of the second clause.","section":"Section 1 and Reference [1]"},{"comment":"The beat-frequency prior is set using the same data that are subsequently analyzed ('This is achieved by applying a fourth-order low-pass butter filter, computing the FFT of the data, and then finding any frequencies...'), which is a double use of the data in the evidence calculation; given the enormous Bayes factors the conclusions are unlikely to change, but this data-driven prior should be acknowledged or fixed independently of the data.","section":"Appendix A.1"},{"comment":"Two small presentation issues: in Figure 10 the caption has an unclosed parenthesis ('M3 with T = 7 (orange when analysing 50 simulated ringdowns'), and in Appendix C the text reads 'the analyses with M3 taken longer' instead of 'take longer'.","section":"Appendix C and Figure 10 caption"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is honest and the core idea is sound; the framework, code release, and candid documentation of limitations are genuine strengths. However, the headline 'up to 25%' accuracy claim is not directly supported by the reported accuracy metrics, and the documented FFT-noise-floor misspecification affects exactly the newly recovered real-data regime. Both issues are fixable within the scope of the manuscript by adding quantitative accuracy statistics and a noise-floor assessment or by narrowing the claims, so I recommend major revision rather than rejection. The citation error for reference [1] and the xi_A sign typos should also be corrected before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know about this paper. First, the abstract's 'improvements in estimation accuracy by up to 25%' is not supported by the evidence. The 25% appears in Figure 4 as the difference between the T=1 and T=7 signal models' predicted amplitudes, not as a measured improvement in tau estimation. The simulation results in Figure 10 do show M3 recovering the true decay constants more accurately than M1 in 46 of 50 injections, but the paper never reports a single numeric accuracy gain, and the abstract overstates what is shown. Second, the paper is unusually transparent about its main weakness: the likelihood in Eq. (24) treats each FFT-peak sample as an independent Gaussian, but those samples are maxima over 0.2 s windows and retain a noise floor when the signal decays. The authors document in Appendix B.1 that this yields biased tau estimates in four of 50 simulations, including one tau1/tau2 swap, and Section 7 concedes the model does not account for the FFT noise floor. This means the claimed 'superior accuracy' is least secure exactly in the low-amplitude tail that motivates the paper's six previously unfittable ringdowns.\n\nWhat is actually new and good here: the M3 analytic photodiode model, with its Taylor expansion and the harmonic series in Eq. (19), is a real advance over the simple two-sinusoid model used by Vajente et al. The two-component stationary-plus-amplitude-dependent Gaussian noise likelihood is also new relative to earlier Bayesian ringdown fits. The paper ships reproducible code (bayesbeat on PyPI, with a data release), validates on 50 simulated ringdowns, and demonstrates that six real ringdowns the old method could not fit are now analysable without special adjustments. That is concrete, reproducible methodological work, and it deserves credit.\n\nThe soft spots, in proportion. The 25% claim is the most serious issue because it misstates what was measured. The FFT noise-floor limitation is real but documented, so it is a scope limitation rather than a hidden flaw. The real-data comparison lacks independent ground truth for absolute mechanical loss, so the different loss estimates from M3 versus M1 are not externally validated. The authors should also report prior sensitivity for the beat-frequency width, since Appendix B.1 implies the tau-swap bias can be avoided by narrowing priors, which is post-hoc.\n\nWho gets value from this: anyone doing ringdown-based mechanical loss measurements for coating Brownian noise, and anyone applying Bayesian inference to nonlinear detector models. The paper deserves a serious referee; the methodological core is sound and the flaws are reparable. I would accept for peer review, with requested revisions to reword the abstract, add or explicitly defer a noise-floor term, and address prior sensitivity.","headline":"Genuinely useful Bayesian framework for ringdown analysis with a real novelty in the photodiode model, but the abstract's 25% claim is unsupported and the likelihood has a documented noise-floor limitation.","tokens_in":21641,"tokens_out":4191,"would_cite":true,"duration_ms":31179,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Bayesian modeling of the split-photodiode nonlinearity recovers disk-resonator decay constants up to 25% more accurately than simple exponential fitting, and fits ringdowns the old method discarded.","keywords":["mechanical loss","Bayesian inference","nested sampling","ringdown measurement","split photodiode","coating Brownian thermal noise","gravitational wave detectors","Gentle Nodal Suspension"],"falsifier":"Generate 50 simulated ringdowns with the full numerical photodiode model (M2) and add a hard FFT noise floor so that the signal drops below it in the second half of each record; then run the M3, $T=7$ inference. If the posterior medians for $\\tau_1$ and $\\tau_2$ are systematically displaced from the injected values, or if the 3-$\\sigma$ intervals exclude the truth in more than a few percent of cases, the two-component Gaussian likelihood fails for low-amplitude tails and the claimed accuracy gain does not generalise.","tokens_in":20582,"feed_emoji":"🔬","tokens_out":11256,"duration_ms":77972,"temperature":0.7,"pith_summary":"This paper seeks to establish that the standard analysis of mechanical ringdowns of coated disk resonators—fitting the beat between two decaying modes as a pair of simple exponentials—is systematically biased by the nonlinear response of the split-photodiode readout and by amplitude-dependent laser noise, and that a Bayesian treatment removes both problems. The authors introduce three signal models of increasing complexity, the most complete being an analytic power-series approximation (M3, with $T=7$) of the true photodiode response, paired with a likelihood whose noise variance grows with signal amplitude. Using 50 simulated ringdowns and 47 real measurements, they report that the refined model recovers the decay constants $\\tau_1$ and $\\tau_2$ with up to 25% better accuracy at large oscillation amplitudes, and that six real ringdowns that previously failed to fit can be analyzed without modification. This matters because these decay constants set the mechanical loss of coating materials, which controls coating Brownian thermal noise—currently the sensitivity limit of gravitational-wave detectors.","feed_headline":"Bayesian fits cut decay-constant error by up to 25 percent","feed_subtitle":"Nonlinear photodiode modeling plus amplitude-aware noise lets failed ringdowns sharpen coating-loss estimates.","key_machinery":"The load-bearing object is the analytic signal model M3 together with the two-component noise likelihood. M3 starts from a Gaussian-beam model of a split photodiode with a finite gap, where the differential voltage is a difference of error functions; expanding that response as a Taylor series in the beam position $\\mu$ to order $T$ gives $V_{S,T} = \\sum_{k=0}^{T} C_k \\mu^k$, with coefficients fixed by the gap and beam size. Demodulating at the mode-pair midpoint frequency and low-pass filtering casts the squared readout as a sum of beat harmonics, $s_3^2(t) = X_0 + X_1 \\cos(\\Delta\\omega t + \\Delta\\varphi) + \\cdots + X_T \\cos(T\\Delta\\omega t + T\\Delta\\varphi)$, whose coefficients carry powers of the decaying amplitudes $A_1, A_2$ and the decay constants $\\tau_1, \\tau_2$. The likelihood treats each ringdown sample as Gaussian with variance $\\xi_i^2 = s(t_i)^2 \\xi_A^2 + \\xi_S^2$, adding an amplitude-proportional relative-intensity-noise term to a stationary electronic-noise term, and nested sampling with normalizing flows is used to compute posteriors and Bayes factors for model comparison.","core_discovery":"The central discovery, on the paper's own terms, is that the bias in ringdown parameter estimation comes from two effects the old analysis ignored: the Gaussian laser spot sweeping across the photodiode gap produces an amplitude-dependent, harmonic-generating response, and the readout noise is not stationary but scales with the instantaneous signal. The paper shows that modeling both effects with M3 and a two-component Gaussian likelihood recovers the true injected $\\tau_1$ and $\\tau_2$ in simulations where the simple model M1 is often biased—more than half of M1's posteriors exclude the true value, while M3 excludes it in only four ringdowns, all of which dip below the relative-intensity-noise floor. On real data the same framework yields tighter, more consistent mechanical-loss estimates across mode families and eliminates the spurious outliers and outright failures of the original method, including fits to six ringdowns that had previously been unanalysable.","pith_inferences":["Because each ringdown data point is the maximum of a 0.2 s FFT power spectrum, replacing the Gaussian likelihood with an extreme-value model for that maximum is a direct way to remove the noise-floor bias the paper itself documents in Appendix B.1; this is my inference, not a claim the paper makes.","The same Taylor-expansion-of-the-error-function trick should transfer to any optical-lever or split-detector measurement where a Gaussian spot scans across a gap, so the framework is likely applicable beyond disk resonators to other mechanical loss measurements.","Switching the readout from FFT-bin maxima to heterodyne or lock-in demodulation, which the authors mention, would simplify the likelihood and reduce the computational cost of the nested-sampling analysis while sidestepping the noise-floor problem.","The model-comparison logic could be used diagnostically in reverse: a low-amplitude ringdown that strongly prefers M3 may signal a misaligned beam, an unexpected beam size, or an asymmetric photodiode response, giving a data-driven check on apparatus alignment."],"forward_implications":["Previously discarded single-decay ringdowns—six in this dataset—can be fitted without special handling, yielding posterior distributions for both $\\tau_1$ and $\\tau_2$ instead of failing.","Mechanical-loss estimates become more tightly clustered across repeated measurements of the same mode family, and the spurious outliers in $1/Q_2$ seen with the original method disappear.","The Bayes factor between M1 and M3 tells the experimenter when a given measurement needs the nonlinear readout model and when the simple exponential model suffices, so model complexity can be matched to oscillation amplitude.","Up to 25% better accuracy in $\\tau_1$ and $\\tau_2$ translates directly into smaller uncertainties in the inferred coating loss $\\phi(f) = 1/(\\pi f \\tau)$, and hence in predicted coating Brownian thermal noise.","Residuals and posterior widths from the Bayesian fit expose amplitude-dependent distortions in the apparatus that the old analysis hid, turning the fitter into a diagnostic tool."],"supporting_citations":[{"why":"It supplies the original fitting method and the M1 signal model that this work extends and compares against.","marker":"[19]"},{"why":"It is the source of the original analysis method used on real ringdowns against which the new results are compared.","marker":"[7]"},{"why":"It defined the stationary-noise likelihood that the new two-component noise model generalizes.","marker":"[9]"},{"why":"It introduces the Gentle Nodal Suspension technique that produces the low-damping ringdown measurements analysed here.","marker":"[17]"},{"why":"It describes the normalizing-flow nested sampler used to obtain posteriors and evidence values.","marker":"[18]"},{"why":"It gives the nested-sampling algorithm that computes the Bayesian evidence underlying every Bayes factor in the paper.","marker":"[24]"},{"why":"It supplies the calibrated logarithmic scale used to label evidence as substantial, strong, or decisive.","marker":"[22]"},{"why":"It sets out the Bayesian framework and model-comparison conventions used throughout the analysis.","marker":"[21]"},{"why":"It provides the code implementation of the M1, M2, and M3 signal models and the likelihood, making the analysis reproducible.","marker":"[20]"}],"fun_headline_variants":["Bayesian inference cuts ringdown error 25%, recovers failed data","Nonlinear photodiode model boosts ringdown accuracy 25%","Bayesian fits rescue discarded ringdowns, cut error 25%","Amplitude-aware Bayesian model sharpens ringdown fits 25%","Bayesian framework recovers lost ringdowns, cuts error 25%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that every ringdown data point is an independent Gaussian random variable whose variance is the signal squared times an amplitude-noise term plus a stationary-noise term; in reality each point is the peak of a 0.2-second FFT power spectrum, so the statistics are extreme-value rather than Gaussian and the model breaks down once the signal falls to the noise floor.","fun_headline_variants_meta":{"raw":{"variants":["Bayesian inference cuts ringdown error 25%, recovers failed data","Nonlinear photodiode model boosts ringdown accuracy 25%","Bayesian fits rescue discarded ringdowns, cut error 25%","Amplitude-aware Bayesian model sharpens ringdown fits 25%","Bayesian framework recovers lost ringdowns, cuts error 25%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000563,"raw_usage":{"total_tokens":2679,"prompt_tokens":960,"completion_tokens":1719,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":576,"completion_tokens_details":{"reasoning_tokens":1623}},"tokens_in":576,"tokens_out":1719,"duration_ms":9618,"temperature":1.0,"reasoning_tokens":1623,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:48:31.708942+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate 50 simulated ringdowns with the full numerical photodiode model (M2) and add a hard FFT noise floor so that the signal drops below it in the second half of each record; then run the M3, $T=7$ inference. If the posterior medians for $\\tau_1$ and $\\tau_2$ are systematically displaced from the injected values, or if the 3-$\\sigma$ intervals exclude the truth in more than a few percent of cases, the two-component Gaussian likelihood fails for low-amplitude tails and the claimed accuracy gain does not generalise.","supporting_citations":[{"cited_title":"A high throughput instrument to measure mechanical losses in thin film coatings","cited_arxiv_id":null,"evidence_quote":"It supplies the original fitting method and the M1 signal model that this work extends and compares against."},{"cited_title":"Demonstration of the Multimaterial Coating Concept to Reduce Thermal Noise in Gravitational-Wave Detectors","cited_arxiv_id":null,"evidence_quote":"It defined the stationary-noise likelihood that the new two-component noise model generalizes."},{"cited_title":"The Theory of Probability","cited_arxiv_id":null,"evidence_quote":"It supplies the calibrated logarithmic scale used to label evidence as substantial, strong, or decisive."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It sets out the Bayesian framework and model-comparison conventions used throughout the analysis."},{"cited_title":"Williams and Joe Bayley.mj-will/bayesbeat","cited_arxiv_id":null,"evidence_quote":"It provides the code implementation of the M1, M2, and M3 signal models and the likelihood, making the analysis reproducible."}],"review_version":1}