{"id":"97bbdccb-8cb0-4765-8e1c-725d492d7f82","arxiv_id":"2607.13283","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":10,"one_line_summary":"Against a common 2025-26 dataset (Planck PR4, ACT DR6, SPT-3G, DESI DR2, Pantheon+), early dark energy and early modified gravity models win the H0 competition (~3σ residual tension), while radiation and late-time solutions fail and nothing fully resolves the tension.","lead":"This paper ran a systematic head-to-head competition among 14 proposed solutions to the Hubble tension, analyzing every model with the same current CMB, galaxy and supernova data and the same statistical tools. It finds that early-dark-energy-style models reduce the tension most, to about 3σ, but that none fully reconciles the local and early-universe measurements.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Baseline hierarchy is ACT-dependent: the paper's own no-ACT runs (Sec. IV E) show groups R and M recover to ~3–3.5σ, so the 'clear hierarchy' claim rests on ACT+Planck cross-calibration that is not independently validated.","rationale":"The reader's weakest_assumption identifies the internal consistency of the combined CMB data vector as load-bearing, and the paper's own no-ACT runs demonstrate that the baseline ranking is materially ACT-dependent. My stress-test agrees with this assessment. The paper is transparent about the dependence and qualifies its headline with 'in the full baseline analysis,' but the actual conclusion still asserts a 'clear hierarchy among physical mechanisms.' That hierarchy is not robust to excluding ACT, so the strength of the claim should be conditional on ACT being correct. This does not change the reader's CONDITIONAL verdict; it reinforces it. I do not find a more fundamental flaw: the multi-metric framework is internally consistent, the model implementations are mostly public, and the robustness tests (extended Planck multipoles, alternative SN samples, S8 prior, BBN) show the group-E preference is stable to many other choices. The principal unresolved issue remains the ACT+Planck cross-calibration, which the paper itself flags but does not resolve. A concrete cross-check—using a different ACT likelihood or a free calibration parameter—would settle whether the hierarchy is physical or dataset-driven. Until then, CONDITIONAL is the appropriate verdict.","tokens_in":60751,"tokens_out":3219,"duration_ms":38709,"concrete_test":"Re-run the baseline CMB+BAO+SN analysis replacing the ACT DR6 'lite' likelihood with (a) the ACT DR4 likelihood, and separately (b) the ACT DR6 full likelihood with a free relative calibration amplitude between ACT and Planck, keeping all other likelihoods and priors identical. If in either case group R or M models achieve ΔDMAP < 3.5σ and ln BF > 3 (or if NEDE/EDE lose their separation), the baseline hierarchy is not robust to the ACT treatment and the 'clear hierarchy' claim should be softened to 'dataset-dependent.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that group E (early dark energy/modified gravity) clearly outperforms recombination (M) and dark-radiation (R) mechanisms. This ranking is driven by including ACT DR6 'lite' likelihoods in the baseline CMB combination. The paper's own no-ACT analysis (Sec. IV E, tables VIIIa–VIIIb) shows that without ACT, most group R and M models reduce their residual tension from ~5σ to ~3–3.5σ, with NEDE reaching 2.3σ, and Bayes factors shift by several units. The paper acknowledges that 'ACT data are preventing several models from efficiently reducing the tension' (Sec. IV E) and concludes in Sec. VI that ACT is 'central to the stronger baseline constraints on groups R and M and to the clearer preference for group E.' Yet no cross-calibration test is performed to verify that the ACT DR6 and Planck PR4 data vectors are mutually consistent under the chosen multipole cuts. The known 2–3σ ACT-vs-Planck Neff disagreement (Refs. [3,7]) and ACT's slight preference for extra high-ℓ power (Sec. IV E) mean the 'clear hierarchy' could be an artifact of using two CMB experiments with imperfectly understood relative calibration. This is not an internal inconsistency in the analysis, but it is a load-bearing assumption: if the ACT high-ℓ preference is a systematics artifact, group E's advantage over R and M would substantially weaken, and the headline conclusion—'clear hierarchy among physical mechanisms'—would not survive contact with a different but equally valid CMB combination.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a systematic, model-by-model reassessment of proposed Hubble tension solutions, updating the 'H0 Olympics' framework. Fourteen models (grouped into late-time, recombination, radiation, and early energy-injection mechanisms, plus an early+late hybrid) are fit to a common baseline of CMB data (Planck PR4 CamSpec with multipole cuts, ACT DR6 and SPT-3G 'lite' primary likelihoods, ACT/SPT lensing), DESI DR2 BAO, Pantheon+, and an SH0ES MB prior, with free neutrino mass. Performance is assessed with Frequentist (DMAP, AIC) and Bayesian (parameter shift, Bayes factor) metrics, with thresholds fixed before inspecting the results. The headline finding is that early dark energy-type models (EDE, NEDE, EMG, RnR) reduce the residual tension to ~2.5-3.1 sigma and are strongly preferred by AIC/BF, while recombination and dark-radiation models fare worse and late-time models fail; no model fully resolves the tension. The robustness section varies CMB likelihoods, SN samples, neutrino-mass priors, S8, baryonic feedback, and BBN, and the authors explicitly show that removing ACT changes the ranking substantially.","tokens_in":61097,"tokens_out":9120,"duration_ms":101312,"significance":"If correct, this is an important field-level result: it sharpens the case that early-time sound-horizon reduction is the only currently viable mechanism class, and it places quantitative constraints on the viability of other proposals. The analysis has real strengths: metrics and thresholds are pre-registered, both Bayesian and Frequentist rankings agree, the neutrino mass sum is varied, and the robustness tests include alternative CMB likelihoods, extended multipole cuts, DES Dovekie, BBN, and S8. The paper ships public code/chain repositories (with a placeholder for analysis notebooks) and is transparent about the ACT-dependence. The central quantitative exercise is careful and reproducible. However, the robustness of the headline hierarchy hinges on the internal consistency of the ACT+Planck data combination, which is not independently validated, and the knockout round contains one untested finalist (EMG) in the BBN analysis.","major_comments":[{"comment":"The headline 'clear hierarchy' is ACT-dependent. In the baseline, DeltaNeff, SIDR, and WZDR have ln BF = -0.09, -1.42, and 1.77, respectively; without ACT these become 9.33, 9.54, and 9.91, and the residual tensions drop to ~3-3.5 sigma (DRMD Delta_shift = 2.8 sigma, NEDE 2.3 sigma). The paper explicitly states that ACT is 'central to the stronger baseline constraints on groups R and M and to the clearer preference for group E' (Sec. VI). Yet no cross-calibration or consistency test between ACT DR6 and Planck PR4 is performed, despite the known 2-3 sigma ACT-vs-Planck Neff discrepancy cited in Sec. IV E. Because the headline ranking is precisely the difference between the with-ACT baseline and the no-ACT case, this is load-bearing; the claim of a 'clear hierarchy' should be made conditional on the relative calibration of the two experiments or supported by an explicit null test (e.g., AC","section":"Sec. IV E, VI; Tables VIIIa-VIIIb vs Table VI"},{"comment":"The EMG model is excluded from the BBN tests because 'specific modifications of BBN theory codes are required.' This matters because BBN constraints worsen all other tested finalists by +0.4-0.7 sigma in Delta_shift(MB) (for example, EDE from 3.0 to 3.5-3.7, NEDE from 2.7 to 3.2-3.5), and because EMG is a headline member of group E. The authors' statement that Ref. [252] bounds 'likely apply' is not a substitute for running the test. The conclusion that early dark energy injection 'minimally and non-minimally coupled to gravity' currently performs best is therefore not fully supported for EMG. Please either implement the BBN treatment for EMG or explicitly restrict the robustness claim to models for which BBN was computed.","section":"Sec. V G, Table X"},{"comment":"The 'Agnos. Reion.' case is not compared on the same data vector, because the SROLL2 low-ell EE likelihood is removed. Its Total chi2 = 5575.66 vs LambdaCDM's 5970.34, and Delta chi2 = -394.68, are not measures of model performance but of data removal. Reporting this row in the same model-comparison tables (DeltaAIC, ln BF) is inconsistent with the paper's stated 'common datasets, likelihoods' framework. Mark this row as not comparable, or re-run it against a LambdaCDM baseline that also drops SROLL2, or remove it from the common ranking tables. This does not change the Group E/R/M ordering, but it is a methodological flaw in a competition whose purpose is fair comparison.","section":"Sec. II L and Tables V-VI"}],"minor_comments":[{"comment":"The abstract says 'four broad mechanisms,' while the conclusions and figures distinguish five categories including 'Group E+L.' Please make the count consistent.","section":"Abstract vs Sec. VI"},{"comment":"The reproducibility statement contains a placeholder: 'A repository containing the analysis outputs and reproducibility notebooks is available at XXX.' This must be replaced with a working URL before publication.","section":"Acknowledgments"},{"comment":"In the +w0,wa columns, the entries for f_idm and log10(zstop) appear transposed: f_idm is bounded by physics to be < 1, yet is listed as '>2.4', while log10(zstop) is listed as '<0.023'. Check and correct the column ordering.","section":"Table XI, DRMD row"},{"comment":"The name 'Chevallier-Polarski-Lindner' should be 'Chevallier-Polarski-Linder'.","section":"Abstract"},{"comment":"The alpha_s/beta_s extension parameterization is discussed as a robustness test for group R, but the paper notes that full-shape galaxy clustering and Lyman-alpha constraints are not included. This caveat is reported, but it should be repeated in the conclusions where 'Deviations from a pure power-law ... substantially improve' group R models, so readers do not over-interpret the rescue.","section":"Sec. IV F"}],"recommendation":"major_revision","confidential_remarks":"The paper is a significant and careful community resource, and I would like to see it published after the robustness points are addressed. The main unresolved issue is not internal inconsistency but external validity of the headline ranking: the authors themselves demonstrate that the group E vs R/M hierarchy depends on including ACT, and they cite a known 2-3 sigma inter-experiment Neff tension, yet no cross-calibration test is provided. The BBN exclusion of EMG and the unequal-data row for 'Agnos. Reion.' should also be fixed or explicitly caveated. None of these issues strike me as fatal; they require additional work or a more conditional conclusion, hence major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is the update to the H0 Olympics that the field needed: fourteen mechanisms, a common pipeline, current data (Planck PR4, ACT DR6, SPT-3G, DESI DR2, Pantheon+, SH0ES), and both Bayesian and frequentist metrics. The competition format is inherited, but the execution is genuinely new — free neutrino mass, curvature/CPL/running extensions, and a knockout round with BBN, S8, and alternate SN samples. The rankings are consistent across metrics, thresholds were set in advance, and the robustness tests are unusually thorough. Credit where earned: this is a careful, self-aware measurement exercise, and the main message — no single mechanism fully resolves the tension, and early-time sound-horizon reduction currently does best — is credible for the baseline dataset.\n\nThe soft spot is the ACT-dependence. The paper's own no-ACT runs (Table VIII) show group R and M models recovering to roughly 3–3.5σ, with NEDE at 2.3σ, and Bayes factors for WZDR, SIDR, and ΔNeff becoming comparable to EDE. The authors do state that ACT is central to the clearer preference for group E, and the extended-multipole robustness tests help, but the headline conclusion in the abstract is not qualified accordingly. If ACT's high-ℓ preference is a systematics artifact, the hierarchy among mechanisms substantially weakens. That is a load-bearing caveat, not a fatal flaw — the paper is honest about it inside — but an abstract reader would miss it.\n\nTwo smaller issues: the acknowledgements point to a \"XXX\" repository placeholder, and two model codes are only \"planned\" releases. For a paper whose selling point is reproducibility, that is sloppy. And the reported Bayes-factor errors (±0.01) are acknowledged to understate the true uncertainty (±0.5); several threshold calls sit right at the edge, so some \"decisive\" vs \"strong\" labels are fragile. Minor-to-moderate, but both should be fixed before publication.\n\nBottom line: this deserves a serious referee. It is the best current benchmark for comparing Hubble-tension mechanisms, and even if you're skeptical of the ACT-dependent ranking, the multi-dataset, multi-metric framework is valuable. I would cite it, and it would stimulate a good reading-group discussion. My recommendation: send to peer review, but ask for a moderated abstract and a real reproducibility link.","headline":"The definitive H0-tension bake-off: careful and transparent, but the 'clear hierarchy' claim rests on ACT more than the abstract lets on.","tokens_in":61881,"tokens_out":3382,"would_cite":true,"duration_ms":37650,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["98.80.-k","98.80.Es"],"model":"deepseek-v4-flash","headline":"Fourteen proposed fixes for the Hubble tension, scored on shared CMB, BAO, and supernova data, leave a clear winner — early dark energy and early modified gravity — and no fully successful solution.","keywords":["Hubble tension","early dark energy","cosmic microwave background","sound horizon","model comparison","dark radiation","neutrino mass","DESI BAO"],"falsifier":"Run the baseline with ACT DR6 replaced by an independent high-multipole CMB dataset, or with ACT and SPT jointly recalibrated: if the radiation and recombination models then recover to 3–3.5σ while early dark energy stays near 2.5–3σ, the headline hierarchy is an ACT artifact; if they remain near 5σ, the hierarchy is physics. A second check: tighten BBN deuterium and helium measurements — the paper finds these worsen the early-dark-energy models by +0.4 to +0.7σ, enough to push them toward the ΛCDM tension level.","tokens_in":60414,"feed_emoji":"🏆","tokens_out":12075,"duration_ms":115517,"temperature":0.7,"pith_summary":"Cosmology's sharpest anomaly is the Hubble tension: local distance measurements put the expansion rate near 73 km/s/Mpc while the standard cosmological model calibrated on the early universe yields about 67–68, a disagreement now quoted above 7σ. This paper stages a competition among fourteen proposed fixes — late-time expansion changes, altered recombination, extra dark radiation, and early dark-energy injection — run through one shared analysis of current CMB, BAO, and supernova data and scored with both frequentist and Bayesian statistics. Its central claim is that a clear hierarchy emerges: only the early-universe mechanisms that shrink the sound horizon (early dark energy, and gravity modified before recombination) cut the residual tension to roughly 2.5–3σ and are strongly preferred over the standard model, while modified recombination and extra radiation leave 4–5σ tensions and late-time fixes barely help. No model fully resolves the discrepancy. The paper also shows the ranking leans heavily on one data set: remove the ACT measurements and most radiation and recombination models recover to 3–3.5σ.","feed_headline":"No Hubble-tension fix fully works; early dark energy comes closest","feed_subtitle":"Fourteen cures, one shared pipeline: only early-time sound-horizon fixes cut the gap to ~3σ.","key_machinery":"The load-bearing object is the comoving sound horizon at recombination: models that shrink it push the CMB-inferred expansion rate upward through the angular-diameter-distance degeneracy, and the whole contest is about which mechanism does this without wrecking the fit to the CMB damping tail and lensing. The scoring apparatus is a single shared pipeline — Planck PR4, ACT DR6, and SPT-3G CMB likelihoods with fixed multipole cuts, DESI DR2 BAO, Pantheon+ supernovae, and a Gaussian prior on the supernova absolute magnitude — evaluated with both a frequentist tension metric (ΔDMAP) and a Bayesian parameter-shift metric, plus two preference criteria (−ΔAIC and the log-Bayes factor). All fourteen","core_discovery":"On its own terms, the paper's finding is a hierarchy. In the baseline analysis, ΛCDM sits at a 5.4σ (ΔDMAP) tension with the locally calibrated supernova magnitude, and no contender closes the gap completely. The best performers are early-time mechanisms: axion-like early dark energy reaches a residual 2.5σ with −ΔAIC ≈ 23 and log-Bayes ≈ 10.5; early modified gravity, Rock'n'Roll, and NEDE cluster at 2.7–3.1σ. The varying-electron-mass model is the only non-early contender to pass the selection thresholds, at roughly 4σ; radiation and late-time groups stay near or above 4.5σ. Without ACT data, most radiation and recombination models recover to 3–3.5σ, so the preference for early dark energy","pith_inferences":["A near-term decisive test is an ACT–SPT cross-calibration at high multipoles: if ACT's preference for extra small-scale power turns out to be a calibration artifact, the paper's own no-ACT numbers suggest the early-dark-energy lead would collapse to a near-tie with radiation and recombination models at 3–3.5σ.","The running-spectrum rescue of the radiation models hints that one shared modification — a blue-tilted primordial spectrum — could mimic either an early-dark-energy or a dark-radiation solution depending on which model family is fitted; a joint fit of both classes together with running parameters would settle which mechanism is actually needed.","If early dark energy is real, it should leave correlated imprints beyond the CMB — a higher matter-clustering amplitude (S8) and signatures in future high-redshift probes such as 21-cm or CMB spectral-distortion measurements — which current data only weakly constrain.","Because the paper scores tension through the supernova magnitude and quotes Bayes factors that it flags as prior-dependent, the absolute significance levels are the least transferable numbers in it; a recalibrated distance ladder would rescale all σ values roughly uniformly, and different priors would move the evidence values by several units — the ordering of mechanisms is the more robust conclus"],"forward_implications":["If the hierarchy holds, the productive path to resolving the Hubble tension runs through pre-recombination physics that shrinks the sound horizon; late-time fixes and weakly interacting extra radiation are observably disfavored.","ACT data carry the ranking: the claim that radiation and recombination solutions fail is contingent on ACT's high-multipole measurements, and would weaken if those data were revised.","The leading early-dark-energy models are exactly the ones most squeezed by BBN light-element abundances and by the DES Y6 clustering amplitude, so better deuterium and helium measurements and large-scale-structure data are the sharpest independent tests.","Allowing a running of the primordial spectrum restores the dark-radiation models to 3–3.5σ tension, so conclusions about radiation mechanisms cannot be separated from assumptions about the initial power spectrum.","Models that ease the tension also dilute the DESI preference for dynamical dark energy, linking the two anomalies."],"fun_headline_variants":["Hubble-tension contest: early dark energy tops 14 rivals, still falls short","No Hubble fix clears 5σ; early dark energy gets closest at 2.5σ","Fourteen fixes for Hubble tension: only early dark energy makes a dent","Early dark energy wins $H_0$ world cup, but no cure is complete"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The ranking assumes the combined CMB dataset (Planck PR4, ACT DR6, SPT-3G, with the chosen multipole cuts) is internally consistent with no significant double-counting — and the paper's own no-ACT runs show this assumption is load-bearing, since dropping ACT lets most radiation and recombination models recover from ~5σ to 3–3.5σ.","fun_headline_variants_meta":{"raw":{"variants":["Hubble-tension contest: early dark energy tops 14 rivals, still falls short","No Hubble fix clears 5σ; early dark energy gets closest at 2.5σ","Fourteen fixes for Hubble tension: only early dark energy makes a dent","Early dark energy wins $H_0$ world cup, but no cure is complete"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000251,"raw_usage":{"total_tokens":1460,"prompt_tokens":874,"completion_tokens":586,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":618,"completion_tokens_details":{"reasoning_tokens":497}},"tokens_in":618,"tokens_out":586,"duration_ms":20331,"temperature":1.0,"reasoning_tokens":497,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T05:39:52.762299+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the baseline with ACT DR6 replaced by an independent high-multipole CMB dataset, or with ACT and SPT jointly recalibrated: if the radiation and recombination models then recover to 3–3.5σ while early dark energy stays near 2.5–3σ, the headline hierarchy is an ACT artifact; if they remain near 5σ, the hierarchy is physics. A second check: tighten BBN deuterium and helium measurements — the paper finds these worsen the early-dark-energy models by +0.4 to +0.7σ, enough to push them toward the ΛCDM tension level.","supporting_citations":[],"review_version":1}