{"id":"cb2913be-0e80-4c44-a581-d02c9c3a7bb4","arxiv_id":"2601.20293","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"An Ok-diagnostic analysis combining DESI DR2 BAO with Pantheon+ or DES Y5 supernovae yields curvature values consistent with a flat universe at the 2-4 sigma level, with the sign depending on the dataset.","lead":"The authors use supernova distances and DESI galaxy clustering data to run a model-independent test of whether the universe follows the standard FLRW geometry and whether space is flat. They find no strong, consistent evidence against flatness, but the inferred curvature depends on which supernova catalog and redshift range is used.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The §3.2 p-value filter assumes χ²_Ok has ν=NBAO−1 independent degrees of freedom, but the Ok points share one smoothing reconstruction and the SNIa covariance; without a calibration of this χ² null, the 95% FLRW-consistency gate and the resulting Ωk medians are not validated.","rationale":"The reader's weakest assumption is well placed. The p-value filter is a novel addition in this paper and is the exact gate that defines the headline sample, yet no calibration is shown. I also note the paper's own caveats that the DESI BAO pipeline assumes flatness (Section 2.2) and that the Δχ²_tot selection is biased toward flatness (Section 2.4); these qualify the 'model-independent' language but are disclosed, so I do not treat them as hidden flaws. The smoothing-scale dependence in Appendix A is transparently reported and would justify a conditional verdict on its own, but the p-value calibration is the more directly testable statistical gap. A Monte Carlo null test can settle it without rederiving the DESI pipeline. Verdict remains CONDITIONAL: the broad conclusion of consistency with flatness within a few sigma may survive, but the specific median Ωk values should not be trusted until the filter is calibrated.","tokens_in":13936,"tokens_out":10748,"duration_ms":97262,"concrete_test":"Generate at least 1000 mock realizations from a flat ΛCDM fiducial using the actual Pantheon+/DES Y5 redshift distributions, covariance matrices, and the DESI DR2 BAO windows used in the paper; run the full iterative smoothing and Ok pipeline identically. For each realization, compute χ²_Ok for the constant fit and compare the empirical distribution to χ²_{NBAO−1} (and to χ²_{ν_eff}, where ν_eff is estimated from the eigenvalue spectrum of the Ok covariance). If the empirical rejection rate at the nominal α=0.05 exceeds the Monte Carlo error (≈1.4% for 1000 realizations), recalibrate the p-values and recompute Table 1; if the headline medians shift by more than the quoted spread, the selection is the bottleneck.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is the p-value filter introduced in Section 3.2. The paper assumes that, for a constant fit to the reconstructed Ok(z), the statistic χ²_Ok follows a χ²_ν distribution with ν=NBAO−1, and it rejects reconstructions with p<0.05 before interpreting the fitted constant as Ωk,0. This filter defines which reconstructions are 'FLRW-consistent', so the medians in Table 1 and the abstract are conditional on it. However, the Ok values at the different BAO redshifts are not independent: they are built from the same iterative-smoothing realization of D and D', which inherits the full SNIa covariance, and the transverse and radial BAO modes are correlated within each bin (as the paper notes in Section 2.2). The effective number of degrees of freedom is therefore plausibly smaller than NBAO−1, and the nominal 95% filter is uncalibrated. The concern is not merely formal: under the Δχ²_SNIa<0 selection only 1.12% of the Pantheon+ & DESI reconstructions pass, and under Δχ²_SNIa+dM<0 the pass fraction is 82.34%, so the rejected 18% can shift the reported median. The paper provides no Monte Carlo or analytic calibration of this χ² null, and the appendix shows the median itself is sensitive to the smoothing scale (e.g., Pantheon+ low-z Ωk goes from 0.172 at Δ=0.2 to 0.011 at Δ=0.4 in Table A.1), so the conditional-on-filter numbers should not be taken at face value until the p-value calibration is established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper performs a model-independent test of the FLRW metric and spatial curvature by reconstructing the dimensionless comoving distance D(z) and its derivative from Type Ia supernovae (Pantheon+ or DES Y5) with an iterative smoothing algorithm, and combining these reconstructions with DESI DR2 BAO transverse and radial mode measurements. Using the O_k diagnostic, which is constant and equal to Omega_k0 in an FLRW universe, the authors fit a constant to the reconstructed O_k(z) at BAO redshifts, apply a p-value filter to select reconstructions consistent with FLRW, and report the median fitted constant as an estimate of Omega_k0. The headline results are Omega_k0 med = 0.035 (+0.046, -0.079) +/- 0.037 for Pantheon+ & DESI DR2, 0.092 (+0.055, -0.132) +/- 0.064 with the Pantheon+ data cut at z=1.13, and -0.119 (+0.113, -0.047) +/- 0.043 for DES Y5 & DESI DR2. The paper concludes that the O_k diagnostic does not rule out the FLRW metric and that most reconstructions are consistent with flatness within 3 sigma, with a slight preference for positive curvature for Pantheon+ and negative for DES Y5.","tokens_in":14361,"tokens_out":15160,"duration_ms":120945,"significance":"If the methodology were fully validated, this would be a valuable addition to the literature on null tests of the FLRW metric, providing a data-driven, Dark-Energy-model-independent measurement of spatial curvature using current DESI DR2 BAO and SNIa data. The paper is transparent about many limitations: it acknowledges the flatness assumption in the DESI BAO pipeline (Section 2.2), warns that the Delta chi2_tot selection is not a genuine litmus test (Section 2.4), and provides an appendix documenting sensitivity to the smoothing scale. The use of two SNIa compilations and multiple Delta chi2 selection criteria is a strength. However, the central claim rests on a p-value filter whose statistical null distribution is assumed rather than calibrated, and the headline numbers in the abstract are not reproduced in Table 1. These issues must be addressed before the results can be taken at face value.","major_comments":[{"comment":"The p-value filter assumes that the statistic chi^2_Ok for a constant fit to the reconstructed O_k(z) follows a chi^2 distribution with nu = N_BAO - 1. This assumption is not justified. The O_k values at different BAO redshifts are constructed from the same iterative-smoothing reconstruction of D(z) and D'(z), which inherits the full SNIa covariance matrix, and the transverse and radial BAO modes are correlated within each bin. The effective number of independent degrees of freedom is therefore likely smaller than N_BAO - 1, yet no analytic or Monte Carlo calibration of the null distribution is provided. The filter is load-bearing: for Pantheon+ & DESI DR2 with Delta chi^2_SNIa < 0, only 1.12% of reconstructions survive; for Delta chi^2_SNIa+dM < 0, 82.34% survive, and the reported medians in Table 1 and the abstract are conditional on this selection. Without a calibrated null distribution, the 95% consistency claim is unsubstantiated. I recommend adding a simulation-based calibration of the p-values (e.g., mock SNIa+BAO realisations from a known FLRW model through the full pipeline) and, if necessary, using the full covariance of the O_k estimates rather than treating the points as independent.","section":"Section 3.2"},{"comment":"The abstract quotes Omega_k0 med = 0.035 (+0.046, -0.079) +/- 0.037 for Pantheon+ & DESI DR2 and 0.092 (+0.055, -0.132) +/- 0.064 for the low-z cut, but Table 1 contains no such values. The closest entries are 0.033 (+0.047, -0.072) +/- 0.037 and 0.091 (+0.055, -0.131) +/- 0.063 in the Delta chi^2_SNIa+dM < 0 rows, and 0.058 (+0.043, -0.107) +/- 0.038 and 0.098 (+0.048, -0.134) +/- 0.064 in the Delta chi^2_SNIa < 0 rows. The central values and asymmetric errors differ beyond rounding. Since the abstract is the primary statement of the result, this inconsistency must be resolved before publication; the authors should either update the abstract to quote the Table 1 values or explain the exact selection criterion and reconstruction set used for the abstract numbers.","section":"Abstract vs Table 1"},{"comment":"The results depend strongly on the smoothing scale Delta. For example, for Pantheon+ & DESI DR2 (low-z) with Delta chi^2_SNIa+dM < 0, Omega_med varies from 0.172 at Delta = 0.2 to 0.011 at Delta = 0.4, compared with 0.091 at Delta = 0.3; for DES Y5 & DESI DR2, it varies from -0.176 to -0.066. These variations are as large as or larger than the reported 'spread' uncertainties and the median 1-sigma errors. The paper nevertheless describes the results as 'robust' (Section 4). The authors should either include the smoothing-scale variation as a systematic uncertainty in the headline results, or soften the robustness claim to reflect the demonstrated sensitivity.","section":"Appendix A and Section 4"},{"comment":"The cut at z < 1.13 is introduced after observing that the high-redshift Pantheon+ points drive the O_k reconstructions away from FLRW, and the abstract presents the resulting values (e.g., 0.092) as one of the main results. Because the cut is motivated by the outcome it removes, the resulting Omega_k0 medians should be framed as an exploratory consistency test rather than a primary measurement. The authors should prespecify the redshift range or treat the cut as a robustness check with an explicit discussion of the selection effect.","section":"Section 3.2 and Discussion"}],"minor_comments":[{"comment":"The sentence 'Only 1.12% the Pantheon+ & DESI DR2 reconstructions pass the p-value test. The remaining reconstructions are consistent with the FLRW metric' is logically contradictory; it should read 'these reconstructions' instead of 'the remaining reconstructions'.","section":"Section 3.2"},{"comment":"The statistic chi^2_Ok is not explicitly defined. Please provide the formula and specify whether the full covariance matrix of the O_k estimates is used or only the diagonal errors.","section":"Section 3.2"},{"comment":"The first uncertainty on Omega_k0 med is described as the spread (max - min) of the central values over reconstructions, not a standard 1-sigma confidence interval. Please clarify the notation (e.g., use brackets for the range) and state explicitly that the asymmetric errors are not 1-sigma errors.","section":"Table 1 and text"},{"comment":"The percentage of reconstructions passing the p-value test is given as 1.12% in Section 3.2 and 1.11% in Section 4; please make these consistent.","section":"Section 4 vs Section 3.2"},{"comment":"The statement that the litmus test results 'do not depend on the values of H0 and rd' should be qualified by the fact that the BAO pipeline assumes a flat fiducial cosmology, as acknowledged in Section 2.2; the H0/rd independence is exact only at the level of the O_k construction.","section":"Section 1 and 2.2"},{"comment":"Reference [16] appears incomplete; please provide the full bibliographic details for the DES Y5 cosmology paper.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The abstract's inconsistency with Table 1 and the unvalidated p-value filter are the two issues that, in my view, most need attention before publication. The paper is otherwise a reasonable application of known techniques, but the central numbers are not yet reliable as stated. I would suggest the authors run mock-based calibration of the p-value filter and fix the headline-number discrepancy; these are fixable within the paper's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. The paper extends the Ok null test to DESI DR2 BAO combined with Pantheon+ and DES Y5 supernovae, and its broad conclusion—FLRW and flatness are not ruled out—is probably right. But the specific curvature medians in the abstract and Table 1 are conditional on a p-value filter whose calibration isn't demonstrated, and the values shift noticeably with reasonable analysis choices.\n\nWhat's genuinely new is modest but real: same iterative-smoothing plus Ok machinery from L'Huillier and Shafieloo and later work, applied to the newest BAO data, with a p-value gate intended to select reconstructions whose Ok is constant, plus a post-hoc low-redshift cut for Pantheon+. The paper is careful about its own limitations: it admits the DESI BAO pipeline assumes flatness, treats the Delta chi^2_tot selection as a consistency check rather than a genuine test, and includes a smoothing-scale appendix.\n\nThe soft spot is the Section 3.2 p-value filter. The paper assumes chi^2_Ok has a chi-squared distribution with N_BAO - 1 degrees of freedom. But the Ok points at different redshifts come from the same smoothing reconstruction and share SNIa covariance, so they're correlated; the effective degrees of freedom are smaller. No Monte Carlo or analytic calibration is given. That matters because under the Delta chi^2_SNIa selection, only 1.12% of Pantheon+ & DESI DR2 reconstructions pass, so the median is defined by a tiny, possibly arbitrary subset. Under Delta chi^2_SNIa+dM the pass rate is 82%, but the filter still affects the result. The appendix reinforces the fragility: for Pantheon+ low-z, the median Omega_k jumps from 0.172 at smoothing scale 0.2 to 0.011 at 0.4—a range not reflected in the quoted error bars.\n\nThe low-z cut is post hoc, but the paper is transparent about it. The Ok diagnostic itself is measured, not assumed, so there's no circularity in the main numbers. The citation pattern looks fine; they reference the relevant DESI DR2 BAO papers and earlier null-test work.\n\nIf you work on curvature tests or DESI, this is a useful consistency check, and the broad flatness result is probably robust. I wouldn't quote the specific medians. I'd send it to a serious referee, with the request that the p-value filter be either calibrated or dropped, and that the dependence on the smoothing scale be presented as part of the headline result.","headline":"A careful null test of FLRW with DESI DR2, but the headline curvature medians depend on an unvalidated p-value filter and shift with smoothing scale; the broad flatness-consistency conclusion is likely solid.","tokens_in":14904,"tokens_out":4071,"would_cite":false,"duration_ms":33864,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Using model-independent supernova distance reconstructions and DESI DR2 BAO data, this paper tests whether the FLRW metric and spatial flatness hold without assuming any dark energy model.","keywords":["FLRW metric","spatial curvature","iterative smoothing","DESI DR2 BAO","Pantheon+","DES Y5","Ok diagnostic","model-independent cosmology"],"falsifier":"Run the full pipeline on mock data generated from a flat FLRW model with the same redshift distribution, masks, and covariance matrices: if the p-value filter rejects more than 5 percent of reconstructions, or if the accepted median $\\Omega_{k,0}$ deviates from zero by more than the reported uncertainty, the filter is miscalibrated. Equivalently, compute the eigenvalue spectrum of the correlation matrix of the $\\mathcal{O}_k$ points and count the modes above noise to measure the effective number of independent degrees of freedom and compare it to $N_{\\rm BAO}-1$.","tokens_in":1894,"feed_emoji":"🌌","tokens_out":8394,"duration_ms":96230,"temperature":0.7,"pith_summary":"Flatness of the universe and the Friedmann-Lemaitre-Robertson-Walker metric are foundational assumptions of standard cosmology. This paper asks whether current data, taken on their own, can rule them out. The authors reconstruct supernova distances and their derivatives with an iterative smoothing algorithm, then combine these with DESI DR2 baryon acoustic oscillation measurements to build the $\\mathcal{O}_k$ diagnostic, a quantity that must be constant and equal to $\\Omega_{k,0}$ in an FLRW universe. They find that after filtering out reconstructions whose $\\mathcal{O}_k$ is not consistent with a constant, the median curvature is close to zero for all three data combinations: $0.035^{+0.046}_{-0.079}\\pm0.037$ for Pantheon+ & DESI DR2, $0.092^{+0.055}_{-0.132}\\pm0.064$ for the same supernovae cut at $z=1.13$, and $-0.119^{+0.113}_{-0.047}\\pm0.043$ for DES Y5 & DESI DR2. So the FLRW metric is not ruled out and most reconstructions are consistent with flatness within about $3\\sigma$.","feed_headline":"FLRW metric survives DESI DR2 curvature test","feed_subtitle":"Median curvature sits within 3 sigma of flat for most reconstructions; sign flips with supernova sample.","key_machinery":"The central object is the $\\mathcal{O}_k$ diagnostic, defined as $\\mathcal{O}_k(z) = (\\Theta^2(z)-1)/D^2(z)$ with $\\Theta(z)=h(z)D'(z)$, where $D(z)$ is the dimensionless comoving distance, $D'(z)$ its derivative, and $h(z)$ the dimensionless Hubble parameter; in an FLRW universe this quantity is identically $\\Omega_{k,0}$. The reconstruction side uses iterative smoothing of the supernova distance modulus to obtain $D$ and $D'$ at the BAO redshifts, while the BAO side supplies transverse and radial mode ratios $d_M/r_d$ and $d_H/r_d$. The key trick is writing $\\Theta(z)$ as the ratio of these two BAO mode ratios times $D'/D$, which cancels the unknown $H_0$ and sound-horizon scale, allowing a genuinely model-independent reconstruction. A p-value filter based on the chi-squared of fitting a constant to the $\\mathcal{O}_k$ points decides which reconstructions are treated as FLRW-consistent and therefore interpretable as measurements of $\\Omega_{k,0}$.","core_discovery":"The paper claims that a data-driven reconstruction of the expansion history, independent of any dark energy model, produces an $\\mathcal{O}_k$ diagnostic that is consistent with the FLRW prediction of a constant value for most reconstructions. In an FLRW universe, $\\mathcal{O}_k(z)\\equiv\\Omega_{k,0}$, and the authors treat deviations from a constant as evidence against the metric. After applying a 95\\% p-value filter to select reconstructions that are FLRW-consistent, and keeping only those that improve the fit relative to the best flat $\\Lambda$CDM model, the median curvature values are the ones quoted above: a slight positive preference from Pantheon+ data and a slight negative preference from DES Y5 data, with uncertainties comfortably overlapping flatness. The high-redshift Pantheon+ data, where the sample is sparse, are identified as the source of apparent FLRW inconsistency, and the authors find that truncating the supernovae at $z=1.13$ restores full consistency.","pith_inferences":["The p-value filter assumes that the chi-squared of the constant fit has $N_{\\rm BAO}-1$ independent degrees of freedom, but the $\\mathcal{O}_k$ points share correlated errors from the same smoothing reconstruction and the same supernova covariance matrix; if the effective degrees of freedom are smaller, the 95\\% filter would be miscalibrated, and this specific test is not performed here.","A direct way to test that calibration would be to run the full pipeline on mock realizations of a flat FLRW universe with the same redshift masks and covariances, checking whether exactly about 5\\% of reconstructions are rejected and whether the accepted median $\\Omega_{k,0}$ remains unbiased; this would also reveal whether the reported spread is inflated or deflated.","The dataset-dependent sign of the curvature preference suggests that residual systematics in the two supernova compilations, especially high-redshift selection effects, could easily shift the median by amounts comparable to the quoted uncertainties, meaning the sign is not cosmologically informative until those systematics are better understood."],"forward_implications":["The FLRW metric, a core assumption of the concordance model, is not rejected by this combined SNIa+BAO dataset, so dark energy models built on that metric remain viable.","The results place model-independent constraints on spatial curvature that are independent of the Hubble constant and the sound-horizon scale, complementing parametric fits.","The apparent high-redshift deviation from FLRW in Pantheon+ is driven by sparse and less reliable supernovae at $z>1$, rather than by a genuine breakdown of the metric.","The sign of the preferred curvature depends on the supernova sample, with Pantheon+ leaning positive and DES Y5 leaning negative, yet both are statistically consistent with flatness for most reconstructions.","The recovery of near-zero curvature under the flatness-assuming selection criterion serves as an internal consistency check of the method."],"supporting_citations":[{"why":"Defines the $\\mathcal{O}_k$ diagnostic and establishes it as a litmus test of the FLRW metric and spatial curvature.","marker":"[7, 8]"},{"why":"Provides the methodology for reconstructing $\\mathcal{O}_k$ from SNIa distance reconstructions combined with BAO measurements, and the error-propagation approach used here.","marker":"[9-11]"},{"why":"Introduces the iterative smoothing algorithm used to reconstruct the distance modulus and its derivative from supernova data.","marker":"[17, 18]"},{"why":"Supplies the DESI DR2 BAO measurements, including the correlation between transverse and radial modes, which provide the BAO side of the $\\mathcal{O}_k$ diagnostic.","marker":"[4]"},{"why":"Supplies the Pantheon+ Type Ia supernova compilation, the primary dataset for one set of distance reconstructions.","marker":"[15]"},{"why":"Supplies the DES Y5 Type Ia supernova compilation, the primary dataset for the second set of distance reconstructions.","marker":"[16]"},{"why":"Provides the fixed Hubble constant value used in the flat $\\Lambda$CDM comparison and the Planck value of $c/(H_0 r_d)$ used for cross-checks.","marker":"[3]"}],"fun_headline_variants":["FLRW metric passes DESI DR2 curvature test","Curvature sign flips between supernova samples","Model-free test of DESI DR2 keeps Universe flat","FLRW metric survives, curvature hints vary","DESI DR2 data: flatness within 3σ for most fits"],"cache_read_input_tokens":16896,"weakest_assumption_plain":"The p-value filter assumes that the chi-squared of fitting a constant to the $\\mathcal{O}_k$ points follows a chi-squared distribution with $N_{\\rm BAO}-1$ degrees of freedom, even though the $\\mathcal{O}_k$ points share correlated errors from the same smoothing reconstruction and the same supernova covariance matrix, so the effective number of independent degrees of freedom is likely smaller.","fun_headline_variants_meta":{"raw":{"variants":["FLRW metric passes DESI DR2 curvature test","Curvature sign flips between supernova samples","Model-free test of DESI DR2 keeps Universe flat","FLRW metric survives, curvature hints vary","DESI DR2 data: flatness within 3σ for most fits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000201,"raw_usage":{"total_tokens":1463,"prompt_tokens":1115,"completion_tokens":348,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":731,"completion_tokens_details":{"reasoning_tokens":268}},"tokens_in":731,"tokens_out":348,"duration_ms":3810,"temperature":1.0,"reasoning_tokens":268,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:39:49.282698+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the full pipeline on mock data generated from a flat FLRW model with the same redshift distribution, masks, and covariance matrices: if the p-value filter rejects more than 5 percent of reconstructions, or if the accepted median $\\Omega_{k,0}$ deviates from zero by more than the reported uncertainty, the filter is miscalibrated. Equivalently, compute the eigenvalue spectrum of the correlation matrix of the $\\mathcal{O}_k$ points and count the modes above noise to measure the effective number of independent degrees of freedom and compare it to $N_{\\rm BAO}-1$.","supporting_citations":[],"review_version":2}