{"id":"da6fd3b2-cc12-4c8e-bbfa-19756ca69e43","arxiv_id":"2412.09409","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":8,"one_line_summary":"An AIC/χ² comparison of exponential F(R) gravity, wCDM, and CPL models against Pantheon+, DESI DR1, cosmic chronometer, and Planck-compressed data claims ΛCDM is excluded at 4σ, but the significance conversion is not statistically justified.","lead":"Cosmologists fit two modified gravity models and two dark energy parameterizations to the latest supernova, galaxy clustering, Hubble clock, and cosmic microwave data, then report that the standard ΛCDM model is excluded at 4σ. The paper matters because if that claim held, the benchmark model of cosmology would need to be replaced by modified gravity or evolving dark energy.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's '4σ exclusion' of ΛCDM has no stated null distribution; Δχ² values of 29–31 are quoted as significance without a formal test, and for F(R) the null sits at a parameter-space boundary.","rationale":"The reader's weakest_assumption correctly identifies the missing statistical calibration as the load-bearing issue. The strongest claim in the paper is exactly the 4σ exclusion of ΛCDM by both exponential F(R) gravity and CPL. That claim cannot be evaluated from the reported min χ² values alone, because no test statistic, null distribution, or boundary correction is provided. For exp(-βR), the ΛCDM limit β→∞ is a non-regular boundary, making standard chi-square asymptotics unreliable; for CPL, the test is nested in principle, but the paper never actually performs the likelihood-ratio test. The reduced χ² values above 1 add a further concern that the covariance matrices may not fully capture systematic errors, which would bias any raw Δχ² significance high. A parametric bootstrap under ΛCDM would settle whether the observed Δχ² values are as rare as claimed. Since the reader already rejected the paper on this basis, the verdict remains unchanged: the underlying AIC preference may be real, but the headline 4σ exclusion is not established by the analysis as presented.","tokens_in":16353,"tokens_out":5586,"duration_ms":61697,"concrete_test":"Run a parametric bootstrap under ΛCDM: using the best-fit ΛCDM parameters and the exact covariance matrices used in the paper (Pantheon+, BAO including DESI, CC, and CMB shift parameters), generate N=1000 synthetic data vectors; fit ΛCDM, exp(-βR), and CPL to each realization, recording Δχ² = χ²_ΛCDM - χ²_alternative. Compare the observed Δχ² values (28.99 for exp(-βR), 31.07 for CPL) to the bootstrap distribution. If the fraction of simulations with Δχ² ≥ observed exceeds about 3×10⁻⁵ for either model, the '4σ' claim is not supported; also report the same calibration for ΔAIC thresholds. This directly accounts for the boundary/non-nested structure of the F(R) comparison and for any non-Gaussian behavior of the profile likelihood.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, 'ΛCDM excluded at 4σ,' is not supported by any stated statistical test. Table II reports min χ² = 2046.79 for ΛCDM, 2017.80 for exp(-βR), and 2015.72 for CPL, giving Δχ² = 28.99 and 31.07. The text then moves directly to 'excluded at 4σ' in Section VI after discussing AIC. But a Δχ² value is not a significance: one must specify the test statistic and its distribution under the null hypothesis. For exp(-βR), ΛCDM is the limit β→∞; the parameter is at the boundary of the model and is unidentifiable under the null, so Wilks' theorem does not apply and the effective number of degrees of freedom is not simply the parameter-count difference. For CPL, ΛCDM is nested at (w0, w1) = (-1, 0), but the paper still never computes a p-value or states a null distribution. Moreover, all models have χ²/dof ≈ 1.14–1.16, indicating the covariance model may be incomplete, which can inflate Δχ². The ΔAIC preference (-25 to -27) is a meaningful model ranking, but AIC does not by itself yield a σ-value; the 'game over' assertion depends entirely on the uncalibrated 4σ language.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper fits the generalized exponential F(R) gravity model (7), its α=1 limit, wCDM, and CPL parametrizations to a combined dataset of Pantheon+ supernovae, DESI DR1 BAO, cosmic chronometers, and Planck CMB shift parameters, and compares them with ΛCDM. The central claim, repeated in the abstract and Section VI, is that ΛCDM is excluded at 4σ by both the exponential F(R) model and the CPL parametrization, with reported Δχ² values of about 29 and 31 relative to ΛCDM and ΔAIC values of about -25 and -27. However, no formal hypothesis test or null distribution is presented; the '4σ' language is not justified.","tokens_in":16691,"tokens_out":4298,"duration_ms":38675,"significance":"If the 4σ exclusion were statistically valid, the paper would report a decisive tension between ΛCDM and both modified gravity and dynamical dark energy using current public data, which would be a major result. The paper does provide a transparent tabulation of best-fit parameters, χ² values, and AIC differences for the models considered, and it correctly computes reduced chi-square values. Yet because the central significance claim is derived from Δχ² and ΔAIC without a calibrated test, the conclusion is not established; the reported improvements are plausibly model ranking rather than exclusion.","major_comments":[{"comment":"The central claim that ΛCDM 'is excluded at 4σ' is not supported by any stated statistical test. The paper quotes Δχ² = 28.99 for exp(-βR) and Δχ² = 31.07 for CPL relative to ΛCDM in Table II, but a difference in minimum chi-square is not a significance level; one must state a test statistic and its distribution under the null hypothesis. For the exponential F(R) model, ΛCDM is recovered in the limit β → ∞, so the null is at the boundary of the parameter space and the parameters are unidentifiable under the null; Wilks' theorem does not apply. For CPL, ΛCDM is nested at (w0, w1) = (-1, 0), but no p-value is computed. The '4σ' language therefore has no probabilistic content in the manuscript.","section":"Abstract, Section VI, Table II"},{"comment":"All models have reduced chi-square appreciably larger than unity: ΛCDM has min χ²/d.o.f. = 2046.79/1769 ≈ 1.157, exp(-βR) has 2017.80/1767 ≈ 1.142, and CPL has 2015.72/1767 ≈ 1.141. This indicates either that the covariance model underestimates the errors or that the models are misspecified, and it can inflate the observed Δχ² between models. The paper neither discusses this nor corrects for it, so the reported Δχ² values cannot be taken at face value as evidence against ΛCDM.","section":"Section IV, Table II"},{"comment":"The Akaike information criterion is used as if its differences were hypothesis-test significance: ΔAIC values of -24.99 for exp(-βR) and -27.07 for CPL are presented together with the claim of 'excluded at 4σ'. AIC is a model-selection criterion with relative weights; it does not provide a p-value or Gaussian sigma. Without a simulation or a properly calibrated evidence measure for the non-nested F(R) versus ΛCDM comparison, the AIC differences do not license the exclusion language.","section":"Section IV, Eq. (28); Section VI"}],"minor_comments":[{"comment":"There is a typo in the abstract: 'tehir power' should be 'their power'.","section":"Abstract"},{"comment":"The text says 'an additional term Finf might be considered in the action (7)' and later writes 'is considered tp become negligible'; 'tp' should be 'to'.","section":"Section II, around Eq. (7)"},{"comment":"The likelihood expression L(θ_j) uses mabs, the absolute minimum of χ², but mabs is not defined before its use in Eq. (27); define it explicitly.","section":"Section IV, Eq. (27)"},{"comment":"There are several grammatical slips in the conclusions, for example 'scenarioa' should be 'scenarios', 'the the large difference' should be 'the large difference', and 'are not connected not with DESI BAO data' should be 'are not connected with DESI BAO data'.","section":"Section VI"},{"comment":"In Table III, the column header for dataset (b) lacks the degrees-of-freedom values that are given for dataset (a); adding the d.o.f. would make the reduced chi-squares directly comparable.","section":"Table III"}],"recommendation":"reject","confidential_remarks":"The paper's underlying fits are to public data and the chi-square tables are useful, but the headline claim is built on an invalid statistical conversion. The issue is not a matter of presentation: the conclusion 'game over' rests on '4σ' language that is never given a null distribution. A rejection is appropriate because the central result would need to be re-derived with a proper significance test, which is beyond the current manuscript's scope. The paper also has heavy self-citation, though I do not see that as the main problem."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper is a clean, careful data-fitting exercise. The genuinely new thing is the result: with Pantheon+ and the latest BAO/CC/CMB data, standard exponential F(R) gravity shifts its best fit to β ≈ 0.75, far from the ΛCDM limit, and the CPL parameterization also moves away from ΛCDM. Earlier fits, including the authors' own 2017 paper, favored ΛCDM, and they are honest about that. The generalization F(R) = R - 2Λ(1 - e^{-βR^α}) lands at α ≈ 1, so the data point to the simpler exponential model. The sensitivity analysis in Table III is also a plus: they check what happens without DESI BAO and conclude Pantheon+ drives the preference. That is a useful, falsifiable statement.\n\nThe math looks internally consistent. The dynamical system is set up correctly, and the fitting procedure is standard. The citation pattern is heavy on their own earlier work, but the central comparison rests on public external data, so I do not see that as circular.\n\nThe soft spot is the headline claim. 'ΛCDM excluded at 4σ' appears in the abstract and Section VI, but the paper never defines the test statistic or its null distribution. They compute Δχ² ≈ 29–31, then jump to '4σ' after citing AIC differences. That does not work. For the F(R) model, ΛCDM sits at β → ∞, a boundary where the parameter is unidentifiable under the null, so Wilks' theorem does not apply. For CPL, the model is nested at (w0, w1) = (-1, 0), but they still do not report a p-value. All models have reduced χ² ≈ 1.14–1.16, which suggests the covariance model may be incomplete and could inflate Δχ². No code or data products are released, so reproducing the fits is not immediate.\n\nThat said, the ΔAIC values of -25 to -27 are substantial. This is not a null result inflated by noise. The paper's real contribution is to show that the latest data give a large model-ranking preference for F(R)/CPL over ΛCDM. What is missing is the calibration of that preference into a significance. A simulation-based likelihood-ratio test or a proper non-nested model comparison would settle it.\n\nWho is this for? Cosmologists comparing dark-energy parameterizations or F(R) gravity to ΛCDM. It is a useful data update, but not a decisive exclusion. The 'game over' rhetoric is unjustified.\n\nRecommendation: send it to peer review. The statistical claim needs to be reworked, but the empirical result is important enough and the analysis is careful enough to justify referee time.","headline":"A competent data-fitting paper whose '4σ exclusion' headline is not supported by the statistics; the real signal is the large ΔAIC, which deserves a proper non-nested test.","tokens_in":17235,"tokens_out":1757,"would_cite":false,"duration_ms":19603,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["04.50.Kd","98.80.-k","95.36.+x"],"model":"deepseek-v4-flash","headline":"The latest Pantheon+ and DESI data exclude the standard cosmological model at about 4σ.","keywords":["F(R) gravity","exponential gravity","dark energy equation of state","CPL parametrization","Pantheon+ supernovae","DESI baryon acoustic oscillations","ΛCDM model","Akaike information criterion"],"falsifier":"Simulate the exponential F(R) models and CPL under the null hypothesis that ΛCDM is true, generating mock Pantheon+ and DESI data with the same covariance, and count how often Δχ²≥28.99 arises; if that fraction exceeds the 4σ tail probability of about 6×10⁻⁵, the exclusion claim fails.","tokens_in":16147,"feed_emoji":"🔭","tokens_out":5071,"duration_ms":47346,"temperature":0.7,"pith_summary":"This paper claims that the newest cosmic-distance data—the Pantheon+ supernova sample, DESI's first baryon-acoustic-oscillation release, cosmic chronometers, and Planck CMB shift parameters—reverse the previous situation in which ΛCDM fit everything best. Confronting a family of exponential F(R) gravity models and the model-independent CPL dark-energy equation of state with the same data, the authors find that ΛCDM is the worst-fitting model, excluded at about 4σ. If correct, the standard cosmological model would need to be replaced by a modified theory of gravity or by dark energy whose equation of state changes with time, and the Hubble-constant tension would deepen because the preferred H0 drops to about 66 km/s/Mpc. The paper's title question 'is the game over?' is answered in the affirmative.","feed_headline":"Standard cosmology excluded at 4σ by newest data","feed_subtitle":"Modified gravity and dynamical dark energy both fit Pantheon+ and DESI observations far better than ΛCDM.","key_machinery":"The central object is the exponential F(R) gravity family F(R)=R−2Λ(1−$e^{{−β R^α}}$) with R=R/(2Λ), where α=1 recovers standard exponential gravity and ΛCDM is recovered in the limit β→∞ or R→∞. The dynamics is solved as a dynamical system in E(a) and R(a) starting from ΛCDM-like initial conditions at high redshift. The comparison statistic is the minimum of the total χ² over SN Ia, BAO, H(z), and CMB data, judged by the Akaike information criterion AIC=min χ²+2N_p, which penalizes extra parameters; the CPL parametrization w=w0+w1(1−a) serves as a model-independent cross-check.","core_discovery":"Using a combined χ² analysis of 1701 Pantheon+ supernovae, DESI DR1 and other BAO data, 32 H(z) cosmic chronometers, and Planck 2018 CMB shift parameters, the paper finds that the standard exponential F(R) model F(R)=R−2Λ(1−$e^{{−βR/(2Λ)}}$) has min χ²=2017.80 versus ΛCDM's 2046.79 (Δχ²=28.99), with an AIC difference ΔAIC=−24.99; the generalized model with α free gives the same min χ² and ΔAIC=−22.99, while CPL gives min χ²=2015.72 and ΔAIC=−27.07. The best-fit β=0.75±0.10 is far from the β→∞ ΛCDM limit, and the best-fit H0≈66 km/s/Mpc with Ωm≈0.314 are mutually excluded with ΛCDM at 1σ to 3σ. The paper concludes that ΛCDM is excluded at 4σ both by exponential F(R) gravity and by the CPL parametrization.","pith_inferences":["The paper does not run a formal hypothesis test: the 4σ significance is read off from Δχ² or ΔAIC, and because ΛCDM sits at the boundary of each alternative (β→∞ or w=−1), a likelihood-ratio calibration with boundary corrections could shrink the claimed significance.","If the Pantheon+ covariance or its zero-point calibration shifts, part of the Δχ²≈29 could be systematic; reanalyzing with an alternative supernova covariance matrix would test this directly.","A concrete forecast: if DESI DR2 data strengthen the CPL w1 preference, the dynamical-dark-energy interpretation gains support; if w1 moves back toward 0, the case for modified gravity weakens.","The same fitting machinery could be applied to an F(R) model with an explicit R² inflation term to check whether the late-time preference survives when early-time cosmological consistency is imposed."],"forward_implications":["If correct, ΛCDM is not merely disfavored but excluded at about 4σ by late-universe data, so the cosmological constant as the sole dark energy is insufficient.","The preferred parameter values H0≈66 km/s/Mpc and Ωm≈0.314 imply the Hubble tension with local distance-ladder measurements worsens, while matching Planck's CMB estimate better.","Dark energy's equation of state must be dynamical: CPL's best fit (w0≈−0.74, w1≈−0.64) deviates strongly from w=−1, consistent with modified gravity mimicking evolving dark energy.","The preference is driven mainly by Pantheon+ supernova data, not by DESI BAO, since removing BAO or using only DESI points leaves ΔAIC nearly unchanged.","A viable theory of gravity beyond general relativity must include non-trivial Ricci-scalar terms at cosmological scales."],"supporting_citations":[{"why":"Supplies the DESI DR1 BAO data used in the fits.","marker":"[2]"},{"why":"Supplies the Pantheon+ supernova sample and its covariance matrix.","marker":"[3]"},{"why":"Introduces the exponential F(R) gravity model.","marker":"[10]"},{"why":"Previous fit with older data that favored ΛCDM, the baseline the new result overturns.","marker":"[17]"},{"why":"Defines the wCDM and CPL equation-of-state parametrizations used as cross-checks.","marker":"[25]"},{"why":"Supplies Planck 2018 CMB parameters and priors.","marker":"[33]"},{"why":"Provides the CMB shift-parameter values and covariance used in the CMB χ² term.","marker":"[38]"},{"why":"Akaike information criterion used for model comparison with parameter-count penalties.","marker":"[39]"}],"fun_headline_variants":["ΛCDM ruled out at 4σ by Pantheon+ and DESI","Modified gravity beats ΛCDM at 4σ in new data","Dark energy not constant: ΛCDM excluded at 4σ","ΛCDM fails 4σ test against F(R) gravity and CPL"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a drop of about 29 in minimum χ² can be read directly as a 4σ exclusion, even though ΛCDM sits at the edge of the alternative models (β→∞ or w=−1) where standard likelihood-ratio rules do not automatically apply.","fun_headline_variants_meta":{"raw":{"variants":["ΛCDM ruled out at 4σ by Pantheon+ and DESI","Modified gravity beats ΛCDM at 4σ in new data","Dark energy not constant: ΛCDM excluded at 4σ","ΛCDM fails 4σ test against F(R) gravity and CPL"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00064,"raw_usage":{"total_tokens":3008,"prompt_tokens":1070,"completion_tokens":1938,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":686,"completion_tokens_details":{"reasoning_tokens":1857}},"tokens_in":686,"tokens_out":1938,"duration_ms":14168,"temperature":1.0,"reasoning_tokens":1857,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T17:05:56.953405+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate the exponential F(R) models and CPL under the null hypothesis that ΛCDM is true, generating mock Pantheon+ and DESI data with the same covariance, and count how often Δχ²≥28.99 arises; if that fraction exceeds the 4σ tail probability of about 6×10⁻⁵, the exclusion claim fails.","supporting_citations":[{"cited_title":"asymptotical","cited_arxiv_id":null,"evidence_quote":"Supplies the Pantheon+ supernova sample and its covariance matrix."},{"cited_title":"This means that at the initial point the factor ǫ = e− β R⟩\\⟩ α should be much smaller than unity","cited_arxiv_id":null,"evidence_quote":"Introduces the exponential F(R) gravity model."}],"review_version":1}