{"id":"63f68f7d-1e4d-42ab-8155-6d83fd61b43f","arxiv_id":"2603.03443","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"The BAO sound-horizon scale extracted from simulated galaxy clustering differs from the standard integral definition enough to bias DESI Year-5-level cosmological constraints on Ω_m and N_eff unless corrected.","lead":"This paper quantifies how the BAO 'standard ruler' in cosmology can be slightly miscalibrated: the sound-horizon scale computed from theory differs from the scale actually imprinted in galaxy clustering. For near-future surveys like DESI Year 5, the mismatch can bias inferred cosmological parameters like matter density and neutrino number at a level comparable to the statistical error.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Quantitative thresholds rely on the untested assumption that reconstruction and realistic survey effects do not alter the BAO shift; Carter et al. found larger biases, so the |ΔΩ_m|=0.03 thresholds may not transfer.","rationale":"The central claim is that the BAO scale imprinted in clustering differs from the integral sound horizon, and that this mismatch introduces systematic bias in cosmological interpretation for DESI-like surveys, with specific thresholds on Ω_m and N_eff. The qualitative argument is sound: even in linear theory, the BAO oscillation frequency is scale-dependent and differs from k·r_int; the ratio s_obs/s_int is not identically 1. The paper's computations for several methods and parameter variations are internally consistent, and the comparison with previous work (Thepsuriya & Lewis, Carter et al.) supports the existence of the effect. The main vulnerability is the leap from simplified forecasts to quantitative statements about DESI Y5. The paper explicitly lists missing elements (reconstruction, masking, realistic covariance) and relies on [32] to argue they do not change conclusions. However, [32] included those effects and found systematically larger biases in α for the same parameter variations, suggesting that realism may amplify rather than erase the effect. The thresholds in the abstract (0.03/0.3) also disagree with the conclusions (0.05/0.4), which, while possibly a typographical slip, underscores that the precise numbers are not robust enough to be taken as a definitive prediction. The reader's choice of reconstruction/pipeline dependence as the weakest assumption is therefore accurate and load-bearing: if those omitted effects alter the cosmology dependence of the BAO shift, the central quantitative claim changes. The proposed test—running the realistic pipeline on mocks with reconstruction and the true covariance—would directly settle whether the simplified thresholds hold. Until such a test is performed, a CONDITIONAL verdict is appropriate, and I see no reason to strengthen or weaken the reader's assessment.","tokens_in":30172,"tokens_out":8979,"duration_ms":83878,"concrete_test":"Run the Sec. 4 DESI-like pipeline on N-body mocks (e.g., Aemulus or AbacusSummit) at the fiducial cosmology and at Ω_m=0.31±0.03 and N_eff=3.046±0.3, with standard BAO reconstruction and the DESI Y5 window/covariance. Compare the recovered Δα/α to the predictions from Table 4. If the deviation exceeds the statistical uncertainty, the quoted thresholds are not robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline thresholds (|ΔΩ_m|≈0.03, |ΔN_eff|≈0.3 for DESI Y5) are computed with a simplified pipeline that omits BAO reconstruction, survey masking, and the real DESI covariance (Sec. 4). The authors import the claim that these omissions do not change the conclusions from Carter et al. [32] (also [18]). But [32] used full simulations including these effects and reported biases about twice as large as the paper's at |ΔN_eff|=1 (0.4% vs 0.2%, App. A). That discrepancy goes in the direction of making the effect larger, not negligible, but it also means the simplified pipeline is not validated against the realistic one for the specific quantities (Δα/α as a function of Ω_m and N_eff) that set the thresholds. If reconstruction or the survey window changes the effective BAO scale's cosmology dependence, the 0.03 and 0.3 thresholds could shift, potentially moving the regime where the bias is 'significant' relative to DESI Y5 errors.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies whether the BAO scale extracted from galaxy clustering is identical to the integral sound horizon r_d^int used in standard compressed-parameter interpretations. It compares several estimators (r_peak, r_P(k) BAO, r_P(k) full, r_ξ) against r_d^int for variations in Ω_b h^2, Ω_m, N_eff, early-Universe recombination models, early dark energy, and massive neutrinos, using DESI-like survey volumes and a Gaussian covariance. The central claim is that the bias Δα/α becomes a significant fraction of the statistical uncertainty for deviations |ΔΩ_m|≈0.03 and |ΔN_eff|≈0.3 for DESI Y5. The paper also implements a more realistic non-linear, redshift-space pipeline and proposes mitigation strategies, including a second-order Taylor expansion in a reduced parameter set.","tokens_in":30405,"tokens_out":6223,"duration_ms":55820,"significance":"If the central claim holds, this is an important systematic for DESI Y5 and other Stage-IV surveys, particularly when fitting extended cosmological models that allow Ω_m or N_eff to drift far from the fiducial values. The paper's strengths are its deterministic, non-stochastic methodology; explicit cross-checks against Ref. [14]; an extended parameter space including pre-recombination physics; a non-linear RSD pipeline in Sec. 4; and concrete correction strategies. The claimed thresholds, however, are not yet the robust result they appear to be: they depend on the simplified likelihood and on an extrapolation about the effect of BAO reconstruction that is not demonstrated here.","major_comments":[{"comment":"The thresholds quoted in the Conclusions, |ΔΩ_m|≈0.05 and |ΔN_eff|≈0.4, contradict the abstract and Secs. 3.3/3.4, which state |ΔΩ_m|=0.03 and |ΔN_eff|=0.3, and also Table 4, which gives 0.029/0.30 for DESI Y5 total. This is not a rounding difference and changes the practical guidance. Please reconcile the numbers and state the exact criterion (e.g., 1/5 σ_α) in one place.","section":"Section 6 (Conclusions)"},{"comment":"The quantitative thresholds are computed with a pipeline that omits BAO reconstruction, masking, and the real survey covariance. The claim that Ref. [32] shows these simplifications do not affect the findings is not supported by the comparison in App. A, where Ref. [32] finds ~0.4% bias at |ΔN_eff|=1 versus 0.2% here. Since the bias-vs-parameter surface is what sets the thresholds, please either run a reconstructed version of the Comp. 1/Comp. 2 test or provide an analytic argument that reconstruction affects only the amplitude/errors and not the cosmology dependence of the BAO scale.","section":"Sec. 4 and Conclusions"},{"comment":"The Taylor-expansion parameter set θ={Ω_b+cdm h^2, N_eff, Ω_b h^2} omits Σmν, but Table 4 gives a Y5-total allowed deviation of +0.482 eV, comparable to the headline thresholds and within the KATRIN bound. The proposed correction is therefore incomplete for a viable parameter direction. Include Σmν in the expansion or justify the omission explicitly.","section":"Sec. 5, Eq. (5.1)"}],"minor_comments":[{"comment":"Refs. [18] and [32] are the same paper (Carter et al. 2020); please consolidate to avoid duplicate citation.","section":"References"},{"comment":"'Thompson scattering' should be 'Thomson scattering'; also in Eq. (2.7) there is a missing space in 'noiseN in'.","section":"Eq. (2.3) and text"},{"comment":"Consider labeling the x-axis consistently as h Mpc^-1 and noting the unit convention in the caption.","section":"Fig. 1 and Sec. 2.2"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely and plausible systematic, but the internal inconsistency between the abstract/Sec. 3 and the Conclusions, together with the undemonstrated reconstruction robustness, makes the headline numbers unreliable as written. If the authors can reconcile the thresholds and provide at least a targeted validation of the reconstruction assumption for the compensated cosmologies, the paper would be suitable for publication in JCAP."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid, useful paper on a known but underappreciated systematic. The core mismatch between the integral sound horizon and the scale actually imprinted in clustering has been known since Thepsuriya & Lewis (2015) and Carter et al. (2020), but this paper extends it to pre-recombination dark energy, varying electron mass, and compensated Neff–h cosmologies, and quantifies the impact for DESI Y5-like surveys. The Taylor correction framework is a practical addition—the coefficients are transparently derived and checked on held-out corners.\n\nWhat's good: the methodology is clear, the synthetic forecasts are documented, and the paper cross-checks against earlier work. The conclusions are stated carefully: for most parameters and for DESI Y1 the effect is negligible; it matters mainly for Omega_m and N_eff in extended models. The paper is honest about its simplifications.\n\nSoft spots, in proportion. First, there is an internal inconsistency: the abstract and Sec. 3 quote |Delta Omega_m| ~ 0.03 and |Delta N_eff| ~ 0.3 as the thresholds for DESI Y5, while Sec. 6 says 0.05 and 0.4. That has to be reconciled. Second, the thresholds come from a simplified pipeline without reconstruction, masking, or the real covariance. The paper argues, following Carter et al., that these don't change the conclusions, and its own Sec. 4 check on two compensated cosmologies is reassuring. But the stress-test note is right that Carter et al. reported larger biases (0.4% vs 0.2% at |Delta N_eff| = 1). That goes in the direction of making the effect larger, not smaller, so it strengthens the paper's warning rather than undermining it, but it also means the exact thresholds should be treated as approximate. The Taylor coefficients are calibrated on the same synthetic data, though the held-out validation mitigates that. None of these is load-bearing; they affect the precise numbers, not the existence of the effect.\n\nWho it's for: BAO analysts interpreting compressed alpha measurements, especially in extended models, and DESI Y5 teams doing systematics budgets. It deserves a serious referee. My recommendation: send it to peer review. It will need a revision to fix the inconsistency and soften the reconstruction claim, but the central argument holds up.","headline":"Useful, transparent extension of a known BAO systematic to new models and DESI Y5; headline numbers have an internal inconsistency and the reconstruction caveat is real, but the work should be refereed.","tokens_in":30975,"tokens_out":4013,"would_cite":true,"duration_ms":35532,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"BAO's standard ruler is not perfectly standard: computing the sound horizon by integral instead of the clustering-imprinted scale biases Ω_m by ~0.03 and N_eff by ~0.3 at DESI Y5 precision.","keywords":["baryon acoustic oscillations","BAO shift parameter","sound horizon","standard ruler","fiducial cosmology dependence","DESI Year 5 forecast","effective number of relativistic species","compressed BAO parameters"],"falsifier":"Run a DESI-Y5-like mock analysis that includes BAO reconstruction and a realistic survey window, and measure the recovered α shift for models at |ΔΩ_m| = 0.03 and |ΔN_eff| = 0.3; if the shift is much smaller or larger than the 1/5-σ level reported here, the thresholds and the need for correction change. Also compare the paper's second-order Taylor-expansion prediction for a three-parameter corner cosmology against a full template minimization; the paper reports a worst-case 8% residual, so any substantially larger residual would falsify the adequacy of that correction.","tokens_in":30003,"feed_emoji":"📏","tokens_out":5500,"duration_ms":49843,"temperature":0.7,"pith_summary":"BAO cosmology assumes the 'standard ruler' — the sound horizon at baryon drag — can be computed by a one-line integral and used to interpret compressed BAO measurements. This paper argues the assumption is not exact: the scale actually imprinted in galaxy clustering is slightly different from the integral, and the difference shows up as a systematic bias when the fitted cosmology moves away from the fiducial one. At DESI Year-5 precision the bias becomes a significant fraction (about one fifth) of the statistical error once Ω_m is off by 0.03 or N_eff by 0.3. The bias survives a realistic non-linear pipeline and can be corrected with a cheap Taylor expansion, so the authors recommend corrections or a systematic budget for such explorations.","feed_headline":"Sound-horizon mismatch biases DESI Y5 at ΔΩ_m≈0.03","feed_subtitle":"The textbook integral underestimates the BAO scale; the bias passes one-fifth of the statistical error when Ω_m or N_eff stray from the fidu","key_machinery":"The central object is the ratio s = r_d^fid / r_d for the cosmology being fitted, compared with the same ratio computed by the integral definition r_int. The BAO shift α — the measured size of the acoustic feature relative to a fiducial template — is what cosmology fits actually constrain, and any difference between the integral ratio and the 'observed' ratio propagates directly into α. The paper constructs four increasingly realistic extractions of the observed BAO scale (peak positions; de-wiggled BAO-only oscillations; a full power-spectrum template with broadband nuisance polynomials; and a correlation-function template), plus a DESI-like EFT pipeline with multipoles, and measures the bi","core_discovery":"The paper's central claim is that the scale conventionally read off BAO data is not the same as the scale defined by the textbook sound-horizon integral. For compressed BAO analyses — where the shift α is interpreted through the ratio r_d^fid / r_d — using the integral r_int in the interpretation step while the data extraction sees the actual oscillation scale r_obs creates a systematic mismatch. Quantifying this across a wider model space than earlier work (including pre-recombination physics, early dark energy, varying electron mass, massive neutrinos, and 'compensated' cosmologies), the authors find that for DESI Year-5 volume and noise the mismatch becomes a significant fraction of the s","pith_inferences":["Because the bias scales with statistical precision, any survey that beats DESI Y5 errors will hit the 1/5-σ threshold for smaller parameter deviations; the effect is likely to become a standard component of BAO systematics budgets for future experiments.","A cleaner long-term fix suggested by the paper's own comparison is to stop using the integral in the interpretation step altogether and instead calibrate the effective BAO scale from the same template family used in the fit; the paper mentions this avoidance route but focuses on corrections because of computational cost.","The quoted thresholds are for deviations from the authors' adopted Planck-like ΛCDM fiducial; analyses using a different fiducial can reuse the same machinery, but the numerical threshold values would shift.","The reliance on an earlier simulation study for reconstruction means the DESI Y5 thresholds should be re-checked once real Y5 reconstruction is available; the direction of any change is unknown."],"forward_implications":["For DESI Y5-like analyses that vary Ω_m or N_eff beyond about 0.03 or 0.3 from the fiducial, ignoring the mismatch adds a systematic comparable to 0.2σ or more; the paper recommends either correcting or inflating the error budget.","A second-order Taylor expansion in {Ω_cdm h², Ω_b h², N_eff} reproduces the bias to within roughly 10% of the statistical error in the worst tested two- or three-parameter cases, and typically much better.","The bias slopes are nearly independent of survey volume and noise, so a correction precomputed for a high-precision idealized survey can be applied to realistic surveys.","DESI Year-1 and smaller surveys are not significantly affected; the thresholds for Y1 are roughly three times larger.","The effect matters most in models that open degeneracies: compensated cosmologies that keep α_int = 1 by shifting N_eff and h in tandem still show >1σ bias for DESI Y5."],"fun_headline_variants":["BAO standard ruler bias threatens DESI Y5 precision","Textbook sound-horizon integral biases BAO at DESI Y5","BAO scale mismatch: standard ruler not so standard for DESI","Systematic bias in BAO ruler: DESI Y5 needs corrections","Sound-horizon integral vs actual BAO scale: bias for DESI"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the simplified, reconstruction-free template fits used throughout the paper respond to cosmology the same way a full DESI BAO pipeline with reconstruction, masking, and realistic covariance does; the paper leans on an earlier simulation-based study for this, and if that transfer fails the quoted thresholds shift.","fun_headline_variants_meta":{"raw":{"variants":["BAO standard ruler bias threatens DESI Y5 precision","Textbook sound-horizon integral biases BAO at DESI Y5","BAO scale mismatch: standard ruler not so standard for DESI","Systematic bias in BAO ruler: DESI Y5 needs corrections","Sound-horizon integral vs actual BAO scale: bias for DESI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000719,"raw_usage":{"total_tokens":3078,"prompt_tokens":770,"completion_tokens":2308,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":514,"completion_tokens_details":{"reasoning_tokens":2228}},"tokens_in":514,"tokens_out":2308,"duration_ms":15910,"temperature":1.0,"reasoning_tokens":2228,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T19:05:50.559884+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a DESI-Y5-like mock analysis that includes BAO reconstruction and a realistic survey window, and measure the recovered α shift for models at |ΔΩ_m| = 0.03 and |ΔN_eff| = 0.3; if the shift is much smaller or larger than the 1/5-σ level reported here, the thresholds and the need for correction change. Also compare the paper's second-order Taylor-expansion prediction for a three-parameter corner cosmology against a full template minimization; the paper reports a worst-case 8% residual, so any substantially larger residual would falsify the adequacy of that correction.","supporting_citations":[],"review_version":1}