{"id":"1ab84cc0-1074-4861-8c68-3aacd9b34e95","arxiv_id":"2502.10264","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"The apparent preference for evolving dark energy depends strongly on which supernova catalog and which BAO survey are used, and is not robust across all independent data combinations.","lead":"This paper re-runs over 35 combinations of cosmological datasets to test whether hints of evolving dark energy hold up across different surveys. It finds the hint is strong only for certain data choices, especially DESI BAO with DESY5 or Union3 supernovae, and nearly disappears when SDSS BAO is combined with PantheonPlus supernovae.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Prior-boundary rows make the headline significance map unreliable: the Δχ²→χ²_2 conversion is used even where wa is unconstrained and hits the prior edge, so the claimed 'robust across most combinations' synthesis is not yet quantitatively supported.","rationale":"The paper is a systematic, carefully structured re-analysis using public likelihoods, and the qualitative pattern — PantheonPlus and SDSS BAO systematically weaken the hint — is supported by the table even if absolute significances shift. I therefore do not want to overstate the problem: this is a conditional-accept issue, not a rejection. The strongest part is the cross-check across many dataset combinations; the weakest link is the model-comparison statistic when the CPL fit is prior-limited. The SN-catalog overlap raised by the reader is real, but it matters mainly for the word 'independent' in the framing; it does not change any individual fit, whereas the prior-boundary issue directly changes the reported significances that define the paper's central map. The paper discloses the overlap in Section III and never combines two SN catalogs in one likelihood, so I treat the overlap as secondary. No chains or scripts are provided, so independent reproduction is not possible from the text alone; this reinforces the need for the boundary check. Agreement with the reader is partial: they note the Δχ² issue in their rationale but select SN overlap as the weakest assumption; I would put the prior-boundary correction first because it affects the quantitative meaning of every affected Table II entry.","tokens_in":34083,"tokens_out":11805,"duration_ms":122531,"concrete_test":"Take the boundary-limited rows (CMB+DESI, CMB+DESI+CC, DESI+CC, and any other row where wa is reported as an upper limit) and re-derive the model-comparison p-value without Wilks' χ²_2 assumption: run 10^4 Monte Carlo realizations of ΛCDM using the same covariance matrices and CPL flat priors, compute the empirical distribution of Δχ², and read off the p-value (or apply the Chernoff boundary mixture with effective dof equal to the number of interior parameters). If the recalibrated CMB+DESI significance drops below 2σ, the bullet that DDE is 'robust across most dataset combinations' must be revised to a weaker statement and Table II should report the boundary-corrected σ.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's significance ladder (Table II, last column; Section III, Eqs. 8–10) rests on Δχ² = min χ²_CPL − min χ²_ΛCDM being asymptotically χ²_2 under the null. This fails when the CPL best fit lies on the flat prior boundary, which is exactly what Table II reports for several combinations used to argue that DESI+CMB supports DDE: CMB+DESI gives wa < −1.05, CMB+DESI+CC gives wa < −0.991, and DESI+CC gives wa < −0.197 (68% CL), so the posterior is truncated by the prior boundary at wa = −2. For a boundary maximum, Wilks' theorem does not apply; the null distribution is a mixture with a point mass at zero and lower effective degrees of freedom, so converting Δχ² with χ²_2 inflates every affected σ. The 2.3σ and 2.6σ entries for CMB+DESI and CMB+DESI+CC are therefore not trustworthy as stated, and the synthesis that 'the preference for DDE remains robust across most dataset combinations' depends on those inflated entries. The qualitative dataset-dependence may survive, but the specific map of which combinations clear 2σ or 3σ would need to be redrawn.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a systematic analysis of the preference for dynamical dark energy (CPL parametrization) over ΛCDM, using Planck CMB, three Type Ia supernova catalogs (PantheonPlus, Union3, DESY5), DESI and SDSS BAO measurements, and cosmic chronometers, across 35 dataset combinations. For each combination the authors compute Δχ² between the best-fit CPL model and ΛCDM and convert it to a Gaussian significance via a χ²_2 approximation. They report that the preference is strongest for CMB+DESI+DESY5 (3.9σ), is weakened when PantheonPlus or SDSS BAO are used, and is otherwise robust across most combinations. The paper concludes that the current DDE hint is dataset-dependent rather than a universal feature of the data.","tokens_in":1847,"tokens_out":2470,"duration_ms":132357,"significance":"If the significance map were reliable, this would be a valuable synthesis: it organizes a large number of publicly available likelihoods and clearly demonstrates that the DDE preference depends on the choice of SN catalog and BAO survey. The paper uses standard codes and data, reports constraints in a transparent table, and explicitly discloses the overlap between SN catalogs, which is a strength. The central qualitative conclusion—that combining PantheonPlus with SDSS BAO substantially weakens the DDE preference—is plausible and consistent with the broader literature. However, several of the headline significance values are not trustworthy because they are derived from a likelihood-ratio approximation that fails for prior-boundary cases, and the repeated use of 'independent' for overlapping SN catalogs overstates the robustness argument.","major_comments":[{"comment":"The conversion Δχ² → χ²_2 in Eqs. (8)–(10) is invalid for rows where the CPL best fit lies on the flat prior boundary. Table II reports such cases: CMB+DESI has wa < −1.05 (68% CL), CMB+DESI+CC has wa < −0.991, CMB alone has wa < −0.197, and DESI+U3+CC has wa < −0.906, all with the lower end of the posterior truncated by the prior boundary at wa = −2. For boundary maxima, Wilks' theorem does not apply: the null distribution of the likelihood-ratio statistic is not χ²_2 but a mixture with lower effective degrees of freedom. The quoted significances for these rows (e.g., 2.3σ for CMB+DESI, 2.6σ for CMB+DESI+CC, 1.9σ for CMB) are therefore inflated. This is load-bearing for the claim in Section V that 'the preference for DDE remains robust across most dataset combinations,' since the CMB+DESI and CMB+DESI+CC entries are used to support that claim without SN data. The authors should either recompute significances with a boundary-aware null distribution (e.g., profile likelihood over the full prior, or a posterior-based evidence ratio), or explicitly mark these rows as unreliable and redraw the significance map. Reporting the raw Δχ² values would also allow readers to check the conversion.","section":"Section III, Eqs. (8)–(10); Table II"},{"comment":"The paper repeatedly calls PantheonPlus, Union3, and DESY5 'independent' Type Ia supernova catalogs (abstract, Section I, Section V), but its own dataset bullets state that Union3 shares 1363 of its 2087 supernovae with PantheonPlus, and that DESY5 includes 194 low-redshift supernovae overlapping with PantheonPlus. The robustness argument in Section V, which treats agreement among these catalogs as evidence from independent probes, is therefore overstated. The text should replace 'independent' with 'distinct' or 'different compilations,' and should either quantify the impact of the overlap (for example, by rerunning the analysis with the shared supernovae removed) or explicitly state that no quantitative correction for the overlap is made.","section":"Section III dataset bullets; Section V"},{"comment":"The abstract and the final bullet list state that 'SDSS-BAO combined with SN from Union3 and DESY5 (with and without CMB) support the preference for DDE.' Table II gives only 1.9σ for CMB+SDSS+U3 and 1.8σ for CMB+SDSS+U3+CC, both below the conventional 2σ threshold. The wording should be softened to 'weakly favor' or 'show a mild trend' for the Union3+CMB cases, and the distinction between >2σ and <2σ evidence should be made explicit.","section":"Abstract; Section V bullet list"},{"comment":"The statement that 'the only scenario where this preference is significantly weakened is when SDSS BAO and PantheonPlus SN are considered simultaneously' is contradicted by other rows in Table II, including CMB+PP (0.1σ), CMB+SDSS (0.3σ), and DESI+PP+CC (0.0σ). If the claim is meant to apply only to a subset of combinations (e.g., only CMB+BAO+SN combinations), that qualification must be stated in the sentence; otherwise the conclusion is factually incorrect as written.","section":"Section V bullet list"}],"minor_comments":[{"comment":"Since Δχ² = min(χ²_CPL) − min(χ²_ΛCDM) is non-positive by construction (the CPL model has two extra parameters), the use of |Δχ²| in Eq. (8) should be motivated and the sign convention stated explicitly.","section":"Section III, Eq. (8)"},{"comment":"The convergence criterion R−1 < 0.02 is given, but no chain lengths, numbers of walkers, or thinning details are reported. Adding these would improve reproducibility.","section":"Section III, methodology paragraph"},{"comment":"For rows with upper limits, the table lists two numbers in parentheses (e.g., '< −1.05 (< −0.238)'); the caption says '68% CL (95% CL)' for parameters with errors, but the convention for upper-limit entries is not explicitly defined and should be clarified.","section":"Table II and Section IV.B.3"},{"comment":"The text says DESI+CC 'fails to constrain wa within the considered flat prior,' while Table II lists an upper limit for wa. This is contradictory; the intended meaning is presumably that wa is not constrained from below and the posterior is prior-dominated. Please rephrase.","section":"Section IV.A.1"},{"comment":"The whisker plots show only 68% CL intervals. For rows with upper limits or open contours, the plots should use arrows or a different symbol to indicate that the 68% interval is truncated by the prior boundary.","section":"Figure 2"}],"recommendation":"major_revision","confidential_remarks":"The paper is essentially a systematic reanalysis of public data; its novelty is modest, but the synthesis could be useful to the community. The main barrier to acceptance is the significance-mapping issue: several headline values rely on a χ² approximation that is invalid at prior boundaries, and the paper's central 'robust across most combinations' claim depends in part on those values. If the authors fix the boundary treatment, report raw Δχ² values, and soften the 'independent catalogs' language, I would support publication. I do not see a reason to reject based on the dataset-dependence conclusion itself, which is plausible and consistent with current literature."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: this is a useful, honestly written systematic map of where the DESI dynamical dark energy preference lives, but its headline significance ladder is unreliable in exactly the rows where the wa posterior hits the prior boundary, and the 'independent SN catalogs' framing is contradicted by the paper's own dataset description. The qualitative conclusion — dataset dependence, with SDSS+PantheonPlus weakening the hint — probably survives. The specific sigma map does not.\n\nWhat's actually new: Table II and the cross-combination synthesis. The individual constraints reproduce earlier work (refs. [191,192,203,219,220]), and the paper says so; the CMB-alone results are explicitly not new. That's an acceptable contribution for a focused review: the field needs a clear map of which data choices drive the hint, and this paper provides it using standard public likelihoods.\n\nI checked the statistical concern about prior boundaries and it's real. The paper converts Δχ² = min χ²_CPL − min χ²_ΛCDM to significance using a χ²_2 null. That holds only when the CPL best fit is in the interior of the parameter space. In several rows of Table II the wa limit is truncated by the prior edge: CMB+DESI gives wa < −1.05, CMB+DESI+CC gives wa < −0.991, DESI+CC gives wa < −0.197, and CMB alone and CMB+U3 are also upper limits. For boundary maxima, the null distribution is a mixture with heavier tails, so the quoted 2.3σ and 2.6σ for CMB+DESI and CMB+DESI+CC are inflated. The statement 'the preference for DDE remains robust across most dataset combinations' depends partly on those inflated entries. It's fixable with a proper boundary-robust calculation or a profile-likelihood treatment, and I'd want that before relying on the detailed sigma map.\n\nSecond soft spot: the paper calls the three SN catalogs 'independent' in the abstract and conclusions, but the dataset section openly states that Union3 shares 1363 of its 2087 SNe with PantheonPlus, and DESY5 contains 194 low-redshift SNe that overlap with PantheonPlus. So the cross-catalog agreement is not an independent check; shared objects can inflate consistency. They disclose this, so it's not hidden, but the framing overstates the robustness.\n\nThird, no chains or scripts are provided, so Table II cannot be independently verified. Minor, but worth noting.\n\nOverall: this is a serious, useful paper for someone who wants the bird's-eye view of the current DDE evidence. I'd send it to a competent referee with specific request to fix the boundary significance analysis and re-frame the SN independence claim. The qualitative message will likely stand.","headline":"Useful systematic map of where the DESI dark-energy hint lives, but the sigma ladder is unreliable where wa hits the prior boundary and the 'independent SN' framing overstates the cross-checks.","tokens_in":34929,"tokens_out":5051,"would_cite":true,"duration_ms":49226,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The evidence that dark energy's equation of state changes over time reaches 3.9σ in the most favorable dataset combination but nearly vanishes when SDSS BAO and PantheonPlus supernovae are used together, so the hint is dataset-dependent…","keywords":["dynamical dark energy","Chevallier-Polarski-Linder parametrization","baryon acoustic oscillations","Type Ia supernovae","cosmic microwave background","cosmological constant","dataset combination robustness","equation of state"],"falsifier":"Recompute the CMB$+$DESI$+$DESY5 significance using only the DESY5 supernovae that are not shared with PantheonPlus; if the $3.9\\sigma$ preference for evolving dark energy drops below $2\\sigma$ once the shared low-redshift objects are excluded, then the headline signal is carried by PantheonPlus overlap rather than by independent DESY5 data.","tokens_in":33776,"feed_emoji":"🔭","tokens_out":10108,"duration_ms":86967,"temperature":0.7,"pith_summary":"Recent DESI baryon acoustic oscillation data, combined with supernova distances, have hinted that dark energy's equation of state changes with time. This paper tests how robust that hint is by reanalyzing more than 35 combinations of Planck CMB, DESI and SDSS BAO, three Type Ia supernova catalogs (PantheonPlus, Union3, DESY5), and cosmic chronometer $H(z)$ measurements under the Chevallier-Polarski-Linder ansatz $w(a)=w_0+w_a(1-a)$. The authors find that the preference for evolving dark energy appears in most combinations and peaks at $3.9\\sigma$ for CMB$+$DESI$+$DESY5, always with the same qualitative pattern: a phantom-like past and a quintessence-like present. The single configuration that systematically erases the preference is SDSS BAO paired with PantheonPlus supernovae, which drops the significance to $1.6\\sigma$ or below. The paper's message is that the preference is consistent across most combinations yet dataset-dependent, not a universal feature of current data.","feed_headline":"Dark energy evolution hint peaks at 3.9 sigma—then drops","feed_subtitle":"Reanalysis of 35 dataset combos finds the signal is strong with DESI/DESY5 but nearly dies with SDSS/PantheonPlus.","key_machinery":"The engine of the analysis is the Chevallier-Polarski-Linder (CPL) parametrization, $w(a)=w_0+w_a(1-a)$, a two-parameter linear model of how the dark energy equation of state varies with scale factor, where $w_0$ is the present value and $w_a$ encodes its evolution, with $w_a=0$ reducing to the cosmological constant. The paper quantifies the evidence for a nonzero $w_a$ by taking the difference in minimum $\\chi^2$ between the CPL model and $\\Lambda$CDM on identical datasets, converting the resulting $p$-value into a $\\sigma$ scale. The second piece of machinery is the dataset grid itself: two BAO surveys (DESI and SDSS), three supernova catalogs (PantheonPlus, Union3, DESY5), Planck CMB, and cosmic chronometers, combined in over 35 configurations so that every probe's contribution to the signal can be isolated.","core_discovery":"On the authors' own terms, the central result is a map of where the dynamical dark energy signal does and does not appear. Across nearly all of the 35-plus combinations, the CPL fit drives $w_0$ toward values above $-1$ (quintessence today) and $w_a$ below $0$ (phantom in the past), and this trend alone is remarkably consistent. The statistical strength, however, splits along two axes: the choice of supernova catalog (DESY5 and Union3 give strong evidence; PantheonPlus weakens it) and the choice of BAO survey (DESI strengthens, SDSS softens). The strongest case is CMB$+$DESI$+$DESY5 at $3.9\\sigma$; CMB$+$DESI$+$Union3 gives $3.5\\sigma$; and CMB$+$SDSS$+$DESY5 gives $2.6\\sigma$. The exception that stands out is any combination containing both SDSS BAO and PantheonPlus supernovae, where the deviation from $\\Lambda$CDM falls to $1.6\\sigma$ and below, even to $1.0\\sigma$ with cosmic chronometers added. That single configuration is what prevents the paper from declaring the DDE preference robust across all current data.","pith_inferences":["The three supernova catalogs are not independent measurements of the same sky: Union3 shares roughly 1,363 of its 2,087 supernovae with PantheonPlus, and DESY5 includes 194 low-redshift supernovae that also live in PantheonPlus, so the cross-catalog agreement is partly the same objects appearing twice.","If the shared supernovae were removed and each catalog analyzed only on its unique objects, the spread in significance between PantheonPlus and the other two catalogs might shrink—or grow, depending on which objects actually drive the signal.","The paper's grid cannot distinguish a genuinely evolving dark energy from a low-redshift distance systematic that tilts supernova magnitudes in a redshift-dependent way, because both would produce the same $w_0$-$w_a$ pattern.","A natural extension would be to repeat the full 35-combination grid with a non-parametric reconstruction of $w(a)$ or a different two-parameter ansatz; if the dataset-dependence persists across parametrizations, it likely reflects real tension in the data rather than the shape of the CPL curve."],"forward_implications":["If the pattern holds, the DESI-reported DDE hint is not a single-survey fluke: it survives when DESI is replaced by SDSS if the supernova catalog is Union3 or DESY5.","Any claim that data favor evolving dark energy should carry a footnote about which combination is being quoted, since the same model ranges from $3.9\\sigma$ to $0.1\\sigma$ across the grid.","The consistent sign pattern ($w_0>-1$, $w_a<0$) means that even weak configurations move parameter estimates in the same direction, which is what one expects if a real effect is being diluted rather than manufactured.","Cosmic chronometer $H(z)$ data add little constraining power once CMB, BAO, and supernovae are included, so future gains in settling this question will come from new BAO or supernova data, not additional chronometers.","A maximum of $3.9\\sigma$ is suggestive but not discovery-level; the spread across combinations is itself the headline uncertainty."],"supporting_citations":[{"why":"Supplies the Planck 2018 CMB temperature, polarization, and lensing likelihoods used in every high-redshift combination.","marker":"[6]"},{"why":"Supplies SDSS/eBOSS BAO and RSD measurements; the combination SDSS+PantheonPlus is the one case that kills the DDE preference.","marker":"[9]"},{"why":"Introduces the CPL parametrization that the whole analysis fits.","marker":"[132, 133]"},{"why":"Provides the DESI BAO measurements in galaxies and quasars that anchor the DESI dataset.","marker":"[193]"},{"why":"The DESI 2024 BAO analysis whose reported 2.5–3.9 sigma DDE hint is the starting point tested here.","marker":"[195]"},{"why":"PantheonPlus supernova catalog and its cosmological analysis; this catalog systematically weakens the DDE evidence.","marker":"[241, 242]"},{"why":"DESY5 supernova dataset and analysis; gives the strongest evidence at 3.9 sigma with CMB+DESI.","marker":"[243–245]"},{"why":"Union3 supernova compilation; yields intermediate evidence between DESY5 and PantheonPlus.","marker":"[246]"},{"why":"Raises the possibility that low-redshift supernova systematics bias the DESY5-based DDE hint, a caveat the paper acknowledges.","marker":"[261]"}],"fun_headline_variants":["Dark energy hint peaks at 3.9σ, then fades with data choice","Data choices split dark energy evidence","3.9σ dark energy hint depends on dataset mix","Dark energy signal robust only with DESI and DESY5","Dark energy hint: 3.9σ with some data, 1σ with others"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The three supernova catalogs are treated as independent probes in the robustness argument, yet they substantially overlap—Union3 shares 1,363 of its 2,087 supernovae with PantheonPlus and DESY5 includes 194 low-redshift supernovae in common with PantheonPlus—so the agreement across catalogs is partly the same data counted multiple times.","fun_headline_variants_meta":{"raw":{"variants":["Dark energy hint peaks at 3.9σ, then fades with data choice","Data choices split dark energy evidence","3.9σ dark energy hint depends on dataset mix","Dark energy signal robust only with DESI and DESY5","Dark energy hint: 3.9σ with some data, 1σ with others"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000784,"raw_usage":{"total_tokens":3573,"prompt_tokens":1168,"completion_tokens":2405,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":784,"completion_tokens_details":{"reasoning_tokens":2315}},"tokens_in":784,"tokens_out":2405,"duration_ms":17312,"temperature":1.0,"reasoning_tokens":2315,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T18:44:18.232550+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the CMB$+$DESI$+$DESY5 significance using only the DESY5 supernovae that are not shared with PantheonPlus; if the $3.9\\sigma$ preference for evolving dark energy drops below $2\\sigma$ once the shared low-redshift objects are excluded, then the headline signal is carried by PantheonPlus overlap rather than by independent DESY5 data.","supporting_citations":[{"cited_title":"Structure Formation in Various Dynamical Dark Energy Scenarios","cited_arxiv_id":"2403.15202","evidence_quote":"The DESI 2024 BAO analysis whose reported 2.5–3.9 sigma DDE hint is the starting point tested here."}],"review_version":1}