{"id":"1e690c93-4109-406e-b73f-0282107b527a","arxiv_id":"2607.13282","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":15,"one_line_summary":"In a systematic head-to-head analysis, early dark energy and early modified gravity models reduce the Hubble tension to about 3σ and are favored over ΛCDM, while radiation and late-time alternatives are not.","lead":"This paper compares fourteen proposed fixes to the standard model of cosmology in one shared analysis of the latest CMB, galaxy clustering and supernova data. Early dark energy-type models come out ahead, reducing the Hubble tension from over 5σ to roughly 2.5–3.6σ, while radiation and late-time fixes fail.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"All-four-Group-E claim rests on Cold NEDE's ln BF=3.03, a marginal prior-dependent crossing with no reported numerical uncertainty; the Bayesian selection threshold could flip.","rationale":"In good faith, the paper is a carefully framed benchmark with standard tools, multiple statistics, and a consistent qualitative pattern. The strongest_claim, however, contains a precise quantitative assertion: all four Group E models cross both thresholds. The risk to that assertion is concentrated in Cold NEDE's log-Bayes factor, 3.03 vs. 3.00, with no error bar and with priors only specified in an unpublished companion. The paper itself flags prior sensitivity at the point where ln BF is defined, so the limitation is acknowledged; that does not remove the fragility of a 0.03 margin. The test I propose is one concrete re-analysis, not a request for new data. If the evidence recomputation confirms ln BF > 3 across reasonable prior variations and estimators, the claim stands and the conditional verdict can be upgraded. If it does not, the headline should be softened. The reader's weakest_assumption—priors inherited from Paper II—is exactly the same load-bearing point, so agreement is 'agree.' No change to the reader's CONDITIONAL verdict is needed now; the concern is a reason to keep the condition rather than to accept outright.","tokens_in":14255,"tokens_out":7957,"duration_ms":91309,"concrete_test":"Recompute ln BF for Cold NEDE and, as controls, EDE and EMG on the baseline A+B dataset using (i) an independent nested-sampling evidence estimator (e.g., PolyChord/CosmoChord), and (ii) the Paper II priors widened and narrowed by a factor of 2 on the model-specific physical parameters (e.g., the trigger/decay parameters and energy fraction), while keeping the same likelihoods. Report the numerical uncertainty on ln BF (e.g., from repeated runs or chain-splitting). If Cold NEDE's ln BF falls below 3 in any of these variations, the 'all four Group E models cross both thresholds' claim should be revised to 'three of four' and the Bayesian selection threshold treated as borderline.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that 'all four Group E models cross both selection thresholds' and that a localized non-radiative contribution is the most effective mechanism. The frequentist thresholds are comfortable (−ΔAIC 19.6–23.4 vs >10), but the Bayesian leg is not: Cold NEDE has ln BF = 3.03 against a threshold of 3.00, a margin of 0.03 in a quantity the paper itself says 'must be interpreted with the adopted priors in mind.' No uncertainty is reported on ln BF, and the priors are those of parent studies/Paper II, many of which were constructed while addressing the H0 tension, so they may already be centered near the H0-preferred region. Because evidence averages over prior volume, a modest widening/narrowing of the Cold NEDE priors, a switch to an independent nested-sampling evidence estimator, or a small numerical error could push ln BF below 3 and remove one of the four Group E qualifiers. The qualitative hierarchy would likely survive, but the 'all four' claim is more fragile than the table suggests.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This Letter presents a systematic comparison of fourteen cosmological models proposed to address the Hubble tension, all analyzed with a common CMB+BAO+SN dataset combination (A) and an optional local supernova-calibration prior B. The authors use two frequentist statistics, ΔDMAP and ΔAIC, and two Bayesian statistics, the log-Bayes factor and Δshift, to evaluate each model's ability to reduce the residual calibration tension and to improve the global fit relative to flat ΛCDM. The central claim is that all four Group E models (EDE, Cold NEDE, EMG, RnR) cross both selection thresholds, shift the inferred H0 to approximately 70 km/s/Mpc, and reduce the residual tension to about 2.5–3.6σ, making a localized non-radiative contribution near matter–radiation equality the most effective mechanism in the baseline comparison. Varying electron mass is the only non-Group-E model that qualifies; radiation and late-time models do not. Full model definitions, priors, and robustness tests are deferred to an in-preparation companion paper (Paper II).","tokens_in":14639,"tokens_out":3181,"duration_ms":39047,"significance":"If the results are correct, this is a valuable community benchmark: it places fourteen often-disparate models in a single likelihood framework with current CMB, BAO, and supernova data, and it cross-checks frequentist and Bayesian selection statistics. The paper is refreshingly explicit about its limitations: it notes that no model fully resolves the tension, that the Bayes factor depends on prior volume, and that conclusions change when ACT data are removed. The use of public pipelines and emulators, and the explicit comparison to the earlier H0 Olympics, make the work reproducible in principle. However, the central Bayesian claim rests on a single marginal crossing (Cold NEDE ln BF = 3.03 versus the 3.0 threshold) and on model priors that are not specified in this Letter. The qualitative hierarchy may well be robust, but the specific claim that 'all four Group E models cross both selection thresholds' is more fragile than the presentation suggests.","major_comments":[{"comment":"Cold NEDE has ln BF = 3.03 against the stated threshold ln BF > 3, a margin of 0.03. No numerical uncertainty or stability analysis is reported for any of the evidence values in Table I. The paper itself warns that ln BF 'must be interpreted with the adopted priors in mind,' but it does not report the numerical accuracy of the nested-sampling estimates, nor does it test the sensitivity of the Cold NEDE crossing to reasonable changes in the inherited priors or to the evidence estimator. Since the crossing of the Bayesian threshold is a load-bearing part of the claim that 'all four Group E models' qualify, this needs a quantitative robustness check or an explicit statement of the evidence uncertainty.","section":"Table I and Section 'We use parallel frequentist and Bayesian criteria'"},{"comment":"All model definitions, parameter priors, and implementation details are deferred to the companion Paper II [1], which is 'in preparation.' The Bayesian evidence integrates over the prior volume, so the ranking and qualification status depend on the priors. The reader cannot verify that the priors are representative or unbiased, especially for models whose parent studies were themselves motivated by the Hubble tension. This is a correctness-risk issue rather than a presentation issue. The Letter should either include the prior ranges in an appendix, or the companion paper must be available and referenced before the central Bayesian comparison can be fully assessed.","section":"Paragraph beginning 'In every case, the reference cosmology...'"},{"comment":"The paper claims that 'a localized non-radiative contribution near matter–radiation equality is the most effective mechanism' based on the Group E results. This is a model-selection statement, but all four Group E models leave a residual tension of 2.5–3.6σ, and the paper admits they are not favored over ΛCDM without the local H0 prior. The conclusion is therefore an empirical ranking of phenomenological proxies, not a detection of the physical mechanism. This is acknowledged in the text, but the abstract and title may overstate the case. I would recommend softening the causal language or making the proxy status more prominent.","section":"Section 'The baseline competition gives a concise empirical target...'"}],"minor_comments":[{"comment":"The caption says 'all 13 contenders' but the table lists 14 rows. The agnostic-reionization row has dashes for model-comparison metrics, so the caption should say '14 contenders, 13 of which have defined model-comparison statistics' or similar.","section":"Table I caption"},{"comment":"The left panel shows H0 intervals from A alone, while the middle and right panels show statistics from A+B. The caption is clear, but the reader must reconcile the different data treatments across the three panels. A sentence in the caption explicitly stating that only the left panel uses A alone would be helpful.","section":"Figure 1"},{"comment":"The definition of ΔDMAP as a square root of a chi-square difference is reasonable, but the label 'non-Gaussian tension' is not defined precisely. It would help to state explicitly that this reduces to the usual number of sigma in the Gaussian limit and to note the interpretation when the profile is non-Gaussian.","section":"Equation (2)"},{"comment":"The reference list is extensive, but Ref. [17] is cited in the introduction for the DESI dynamical dark energy preference and again in the text; consider consolidating. Also, several arXiv entries have incomplete bibliographic details (e.g., Refs. [14,15] with DOI placeholders); these should be completed during production.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The central question for the editor is whether a summary Letter whose main Bayesian claim depends on a marginal threshold crossing and on an unavailable companion paper is publishable in its current form. I think the comparison framework is valuable and the frequentist results are solid, but the Bayesian leg of the 'all four Group E' claim needs either numerical evidence uncertainties, prior-robustness tests, or the companion paper to be available. This is fixable within revision, so major revision seems appropriate rather than rejection. The paper is also honest about dataset sensitivity, which strengthens its credibility."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth reading, and worth refereeing. The paper updates the H0 Olympics framework to current data (Planck PR4/NPIPE, ACT DR6, SPT-3G, DESI DR2 BAO, Pantheon+) and compares 14 proposed solutions in one pipeline. The main result is a hierarchy: four early-energy models (EDE, Cold NEDE, EMG, RnR) shift H0 to roughly 70 km/s/Mpc and cut the residual tension to about 2.5–3.6σ; varying me is intermediate; radiation and late-time fixes do little. Frequentist and Bayesian statistics agree in broad shape, which is a genuine cross-check.\n\nThe design is clean: it keeps the uncalibrated dataset A separate from the local MB prior B, and reports both tension (ΔDMAP, Δshift) and model preference (ΔAIC, ln BF). That is the right way to separate 'does it reconcile the datasets' from 'is it worth the extra parameters.' The paper is also honest: it notes the evidence must be interpreted with priors in mind, states that removing ACT loosens radiation constraints back to 3–3.5σ, and flags the models as phenomenological proxies.\n\nSoft spots, in proportion. The most concrete one is the one the stress-test flagged: the statement that all four Group E models cross both thresholds rests on Cold NEDE's ln BF = 3.03 against a threshold of 3.00. No numerical uncertainty is reported on the evidence, and the priors come from parent studies that were often constructed with the tension in mind. A modest change in prior volume or a different evidence estimator could push that number below 3. That would not overturn the qualitative ranking — EDE, RnR, and EMG are comfortable — but the 'all four' phrasing is more fragile than the table suggests, and the paper should either quote a robustness check or soften the claim. A related limitation is that all model definitions, priors, and knock-out round robustness tests are deferred to the in-preparation Paper II. That is acceptable for a programmatic Letter only if Paper II actually appears and the numbers reproduce; for now a referee cannot fully verify the ranking from this text alone. The circularity concern raised in the reader's report I do not share: each model is fit to the data, no result is assumed, and the parent-study priors are a reasonable starting point.\n\nBottom line: a solid benchmark update and a useful common reference for the field, not a new physical mechanism. Anyone working on the Hubble tension or early-universe model comparison will get value from it. I would send it to peer review rather than desk reject, with a request that the Cold NEDE evidence sensitivity be addressed. I'd also bring it to reading group and cite it in my own work.","headline":"Solid, well-executed H0 Olympics update that deserves refereeing; the Cold NEDE Bayesian qualification is knife-edge and should be checked before the 'all four' claim is trusted.","tokens_in":15159,"tokens_out":3468,"would_cite":true,"duration_ms":35088,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["98.80.-k"],"model":"deepseek-v4-flash","headline":"A head-to-head comparison of 14 proposed Hubble-tension solutions finds that only transient early-universe energy injections both shift the inferred Hubble constant to about 70 km/s/Mpc and cut the residual discrepancy with local measuremen","keywords":["Hubble tension","early dark energy","H0","cosmological model comparison","Bayesian evidence","Akaike information criterion","sound horizon","recombination"],"falsifier":"Re-run the baseline comparison without the ACT high-multipole dataset (or with a different high-resolution CMB experiment in its place). If an extra-radiation or late-time model then crosses the AIC/Bayes thresholds and pushes its residual tension below 3σ, the claim that a localized non-radiative injection is the most effective mechanism would be refuted.","tokens_in":14165,"feed_emoji":"🌌","tokens_out":6060,"duration_ms":59920,"temperature":0.7,"pith_summary":"This paper asks whether any of the fourteen most studied fixes for the Hubble tension genuinely works when all are run through the same analysis pipeline with the same CMB, BAO, and supernova data. It claims that only models with a localized, non-radiative energy injection around matter–radiation equality – the early dark energy family – pass both a frequentist and a Bayesian bar: they shift the inferred expansion rate to about 70 km/s/Mpc and reduce the residual tension from more than 5 sigma to around 2.5–3.6 sigma, while improving the joint fit over the standard model. Varying the electron mass at recombination helps only halfway, and neither extra radiation nor late-time modifications improve over the standard model. The paper's punchline is that the mechanism matters: the fix has to shrink the sound horizon, not simply tweak what happens later.","feed_headline":"Early dark energy models cut Hubble tension to ~3 sigma","feed_subtitle":"In a 14-model showdown on identical data, only pre-recombination energy injections pass fit tests and push H0 near 70.","key_machinery":"The central mechanism is a transient, non-radiative energy component peaking near matter–radiation equality – a short-lived scalar field that briefly accelerates expansion before recombination and shrinks the sound horizon. This allows a higher H0 from early-universe data without disturbing the well-measured CMB damping tail and the BAO/SN distance relation. The comparison is carried by two complementary statistics: frequentist (AIC and ΔDMAP) and Bayesian (log-evidence and Δshift), with thresholds −ΔAIC > 10 and ln BF > 3.","core_discovery":"The central claim is that, in a common statistical framework applied to the baseline CMB+BAO+SN dataset, the four early-energy models (axion-like early dark energy, cold NEDE, early modified gravity, and its Rock'n'Roll limit) are the only contenders that both reduce the residual calibration tension to the 2.5–3.6σ level and achieve strong joint-fit support over ΛCDM (−ΔAIC > 10 and ln BF > 3). A localized non-radiative contribution near matter–radiation equality is thus identified as the most effective mechanism, because it shrinks the sound horizon while leaving the detailed high-multipole CMB spectra and the low-redshift distance relation intact. The paper does not claim any model fully r","pith_inferences":["The paper's group-stage framing suggests a competition with winners, but the honest reading is that a class of mechanisms (early, non-radiative injections) is favoured; which specific model wins depends on the statistic.","The stated sensitivity of radiation models to the ACT dataset implies that the eliminated or relegated status of the radiation group is not robust to future high-multipole CMB data; a rerun with different high-resolution data could reshuffle the lower ranks.","A natural testable extension is to feed the same baseline into multi-field or otherwise extended early-energy models, which the paper hints can relax the tight constraints on the single-field case.","If the early-energy mechanism is correct, its signatures should appear in complementary probes such as cosmic birefringence, fifth-force constraints, or accelerator searches of the associated particles – a thread the paper leaves for later work."],"forward_implications":["Successful models must reduce the sound horizon while preserving the high-multipole CMB spectra and the low-redshift distances fixed jointly by BAO and supernovae.","Radiation-only additions are limited by damping-tail and phase-shift signatures; post-recombination changes have too little freedom once BAO and SN are included.","Modified recombination (varying electron mass) can shift the relevant scales and qualifies, but with a larger residual tension (~4σ).","The neutrino-mass sum can be varied freely without significantly affecting the tension results, and its constraints are nearly as tight in most extended models as in ΛCDM.","Both statistical frameworks agree on the overall hierarchy, even though the internal ordering within the early-energy group depends on whether best-fit or posterior-volume measures are used."],"fun_headline_variants":["Early dark energy tops 14-model H0 field, tension cut to ~3 sigma","H0 World Cup: early dark energy wins group stage, easing tension to 3σ","Four early-energy models beat LambdaCDM, lower H0 tension to 2.5-3.6σ","In 14-model H0 bake-off, early dark energy is clear winner on fit","Baseline H0 group stage: only pre-recombination energy shifts survive"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The verdict depends on the prior ranges chosen for each model's extra parameters (taken from the original proposals, not derived here) and on the inclusion of the ACT high-multipole data; the paper itself notes that removing ACT relaxes radiation-model constraints back to 3–3.5σ.","fun_headline_variants_meta":{"raw":{"variants":["Early dark energy tops 14-model H0 field, tension cut to ~3 sigma","H0 World Cup: early dark energy wins group stage, easing tension to 3σ","Four early-energy models beat LambdaCDM, lower H0 tension to 2.5-3.6σ","In 14-model H0 bake-off, early dark energy is clear winner on fit","Baseline H0 group stage: only pre-recombination energy shifts survive"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00131,"raw_usage":{"total_tokens":5222,"prompt_tokens":837,"completion_tokens":4385,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":581,"completion_tokens_details":{"reasoning_tokens":4270}},"tokens_in":581,"tokens_out":4385,"duration_ms":28182,"temperature":1.0,"reasoning_tokens":4270,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T05:37:02.599822+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the baseline comparison without the ACT high-multipole dataset (or with a different high-resolution CMB experiment in its place). If an extra-radiation or late-time model then crosses the AIC/Bayes thresholds and pushes its residual tension below 3σ, the claim that a localized non-radiative injection is the most effective mechanism would be refuted.","supporting_citations":[],"review_version":1}