{"id":"083004b7-0950-46f6-a902-c0d9e1b89ecf","arxiv_id":"2411.13621","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"A historical meta-analysis of 163 Hubble constant measurements argues published error bars were often underestimated and recalibrates the 4.4-sigma Hubble tension down to an equivalent 2.1-sigma effect.","lead":"This paper analyzes 163 published Hubble constant measurements and finds their scatter is far larger than the reported error bars, concluding that error estimates were often too small. It then applies a historical error calibration to argue the much-debated 4.4-sigma Hubble tension is equivalent to only a 2.1-sigma fluctuation.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eq. (3) is calibrated on per-point residuals x=|H_i−Hbar|/σ_i but is applied to the pairwise Riess–Planck 4.4σ tension; no distribution for pairwise differences is derived, so the 4.4→2.1σ conversion is unsupported.","rationale":"The paper's descriptive historical finding—163 published H0 values scatter far more than their quoted errors (χ²≈580)—is well supported and is a useful cautionary observation. However, the central quantitative claim depends on a calibration function whose input variable does not match the target quantity. The paper fits Eqs. (2)–(3) to per-measurement standardized residuals from a weighted mean, then applies the result to a pairwise tension between Riess et al. and Planck. These are different statistics: the tail behavior of individual residuals does not, by itself, give the tail probability of a pairwise difference, and the paper provides no convolution or pairwise survival calculation. The internal inconsistency is visible in Table 1, where the 2019 Riess entry appears at 4.1σ, not the 4.4σ quoted in the abstract. The reader's in-sample concern is valid and related, but the variable mismatch is the more decisive flaw; the two together remove the statistical basis for the 2.1σ recalibration. The verdict should remain REJECT for the central claim, while preserving the value of the historical dataset as a caution about overconfident error bars.","tokens_in":8374,"tokens_out":11397,"duration_ms":128627,"concrete_test":"Using the published 163-value compilation (Faerber & López-Corredoira 2020), compute the empirical distribution of pairwise standardized differences D_ij=|H_i−H_j|/√(σ_i²+σ_j²) for a random subset of non-overlapping pairs, and estimate P(D_ij>4.4). If this probability differs substantially from the paper's 0.036, Eq. (3) was applied to the wrong statistic; as a secondary check, refit Eq. (3) with the 2019 Riess entry excluded and see whether x_eq(4.4) moves by more than 0.3σ.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3 defines x as |H_i−Hbar|/σ_i, a single measurement's deviation from the sample weighted average, and fits Eqs. (2)–(3) to the empirical survival function of these residuals. The abstract then applies the resulting 4.4σ→2.1σ conversion to the Riess-versus-Planck tension. But that tension is a pairwise standardized difference T=(74.0−67.4)/√(1.4²+0.5²)≈4.4, not a single residual; indeed Table 1 lists the 2019 Riess point as only 4.1σ from the weighted average. A heavy-tailed distribution for individual residuals does not determine the tail of pairwise differences, which involve a convolution of two error distributions. No distribution of pairwise tensions is derived, so P=0.036 is not a valid probability for the 4.4σ tension. This is compounded by the in-sample nature of the fit: the 2019 Riess point is part of the 163-point sample used to derive Eq. (3).","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compiles 163 published Hubble–Lemaître constant measurements from 1976 to 2019 and reports that their scatter around the weighted average H0 = 68.26 ± 0.40 km s−1 Mpc−1 gives χ2 = 575.7, far larger than the number of points. Interpreting this as evidence of systematically underestimated error bars, the author fits a heavy-tailed survival function P(|H0,i − H0,avg| > xσ) = 0.93 exp(−0.72x) and an equivalent-Gaussian conversion x_eq = 0.83 x^0.62. The conversion is then applied to the Riess et al. versus Planck 4.4σ tension to claim an equivalent 2.1σ tension with probability P = 0.036, and the paper concludes that Hubble tensions were common before 2019 and that the present-day attention to the Hubble tension reflects sociological groupthink.","tokens_in":8634,"tokens_out":2207,"duration_ms":24046,"significance":"The historical observation that published H0 error bars frequently underestimate the true dispersion is interesting and is supported by the simple chi-squared calculation in Section 2; this part of the paper is a useful reminder of the checkered history of H0 measurements. The paper therefore deserves credit for quantifying, in one transparent statistic, that the historical record is far more scattered than the quoted errors imply. However, the central quantitative claim — that a 4.4σ tension should be read as a 2.1σ tension — rests on two unsupported steps: the equivalent-Gaussian curve is fitted in-sample on the same data point it is used to reinterpret, and it is fitted to single-measurement residuals while being applied to a pairwise comparison. If the historical dispersion result is retained but the recalibration is removed, the paper becomes a valid historical critique rather than a statistical tool for downgrading modern tensions.","major_comments":[{"comment":"The conversion x_eq = 0.83 x^0.62 is fitted to the full 163-point sample, which includes the 2019 Riess et al. measurement listed as the last row of Table 1 (4.1σ from the weighted average). Applying this same fitted curve to reinterpret that very point is in-sample prediction, not independent evidence. The claim that the 4.4σ tension becomes 2.1σ therefore has no out-of-sample support. A leave-one-out or holdout validation, or at least a fit from which the 2019 point is excluded, is needed before this curve can be used to recalibrate that measurement.","section":"Section 3, Eq. (3)"},{"comment":"The probability distribution in Eq. (2) is defined for x = |H0,i − H0,avg|/σ_i, i.e. the deviation of a single measurement from the sample weighted average. The 4.4σ tension that the paper reinterprets is, however, a pairwise standardized difference T = (74.0 − 67.4)/sqrt(1.4^2 + 0.5^2) ≈ 4.4 between the Riess et al. local measurement and the Planck value. The tail probability of individual residuals does not determine the tail probability of a pairwise difference, because the latter requires the convolution of two error distributions; no such pairwise-tension distribution is derived. Consequently P = 0.036 is not a valid probability for the 4.4σ tension as stated in the abstract and Section 3.","section":"Section 3 versus Sections 1 and 4"},{"comment":"The chi-squared analysis treats the 163 measurements as independent, but the paper itself acknowledges that many measurements are incremental updates sharing common foundations. This is not merely a caveat: if the effective number of independent measurements is substantially smaller than 163, then both the quoted Q ≈ 10^−47 and the fitted tail of Eq. (2) are distorted. Since Eq. (2) is the basis of the recalibration, the dependence of the fit on the assumed independence needs to be quantified, for example by removing duplicate chains or by using an effective number of degrees of freedom.","section":"Section 2, χ2 calculation"}],"minor_comments":[{"comment":"The manuscript contains several typographical errors, including 'compililation', 'posslibility', 'understimated', 'Feynmann', and the missing space in 'constantH0' on page 1.","section":"General"},{"comment":"Ref. [19] is a blog post rather than a peer-reviewed source, and large parts of Sections 1, 2, 3, and 4 are excerpts from Ref. [1], the author's own previous paper; this should be stated more clearly in the text so that the reader can distinguish the new material from the reused material.","section":"References"},{"comment":"The fitted exponential amplitude 0.93 ± 0.06 is close to but not equal to 1.0 at x = 0; if the survival function is to be interpreted as a properly normalized probability, the behavior near x = 0 should be commented on or the amplitude should be fitted with the normalization constraint.","section":"Section 3, Eq. (2)"}],"recommendation":"reject","confidential_remarks":"The historical chi-squared result is worth preserving as a cautionary note, but the central recalibration is not statistically defensible. The in-sample application of Eq. (3) to the Riess 2019 point and the mismatch between single-residual statistics and pairwise-tension statistics are both load-bearing. These are not presentation issues that a minor revision could fix within the current scope. The paper might be suitable as a commentary if the quantitative recalibration were removed, but as submitted it does not meet the standard for a research article."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is mostly a proceedings condensation of the author's earlier work (MNRAS 2022 and the 2020 compilation). The one genuinely new section is the 'groupthink' sociology, which is an opinion piece. The historical observation is real: 163 published H0 values between 1976 and 2019 give chi^2 around 576 against the weighted average, which is overwhelming evidence that quoted errors were too small. That point is well made and worth taking seriously.\n\nThe problem is the paper's central quantitative claim. Equation (3), x_eq = 0.83 x^0.62, is fit to the tail of per-point residuals x = |H_i - Hbar|/sigma_i, i.e. how far an individual measurement sits from the sample's weighted average. The 4.4 sigma tension it is then applied to is a pairwise comparison: (74.0 - 67.4)/sqrt(1.4^2 + 0.5^2). Those are different random variables. A heavy-tailed distribution for a single measurement's residual does not determine the tail of a pairwise difference, which involves both errors. You would need to derive the distribution of pairwise tensions from the historical sample, or fit an error distribution and convolve it. The paper does neither, so P = 0.036 is not a valid probability for the 4.4 sigma discrepancy.\n\nThere is also an in-sample issue. The 2019 Riess point with 74.0 +/- 1.4 is row 27 of Table 1, part of the same 163-point sample used to fit Eq. (3). Using a curve fitted to that sample to reinterpret that same point is circular, not independent evidence. The 2.8 sigma cutoff used to get Q >= 0.05 is also post hoc, though that is a secondary issue.\n\nThe groupthink section is speculation, not sociology: it cites Sunstein and Janis but no data on cosmologists' behavior. The author presents it as a suggestion, so I don't hold it to a high evidentiary bar, but it does not belong in the quantitative argument.\n\nWhere does that leave the paper? The historical caution is useful and probably correct: many past H0 error bars were underestimated. But the paper's contribution is mostly a restatement, and its new quantitative claim fails on the residual-vs-pairwise confusion. If this were submitted as a research paper, I would not send it to a referee for the quantitative claim as it stands. I would encourage the author to either present it explicitly as a proceedings summary and opinion, or redo the recalibration using actual pairwise differences. The Hubble tension is important enough that the 4.4 -> 2.1 sigma downgrade, if it were valid, would matter. But it isn't supported by the evidence here.","headline":"The historical point is sound, but the statistical bridge from single-residual scatter to the pairwise 4.4σ tension doesn't exist, and most of the content was already published.","tokens_in":9207,"tokens_out":3398,"would_cite":false,"duration_ms":35118,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"After recalibrating with 163 historical measurements, the 4.4-sigma Hubble tension is a 2.1-sigma fluctuation.","keywords":["Hubble constant","Hubble tension","error-bar underestimation","cosmological parameter estimation","statistical significance","meta-analysis","groupthink in science","standard cosmological model"],"falsifier":"Compile all independent $H_0$ measurements published from 2020 onward (which were not used to fit the calibration curve), compute each one's deviation from the current inverse-variance weighted mean in units of its quoted error, and count the fraction that exceeds $3\\sigma$. A Gaussian error distribution predicts about 0.27% of points beyond $3\\sigma$; the historical sample gives 11.7%. If the modern fraction is close to 11.7%, the recalibration is confirmed and the 4.4-$\\sigma$ tension should indeed be read as roughly a 2.1-$\\sigma$ event; if the modern fraction is close to 0.27%, the historical curve does not transfer to current-era measurements and the tension stands. The companion extension of the sample to 2012-2022 already provides a partial version of this test.","tokens_in":8153,"feed_emoji":"🔭","tokens_out":19448,"duration_ms":158988,"temperature":0.7,"pith_summary":"This paper asks whether the much-discussed “Hubble tension” between the local Cepheid-supernova distance ladder and the cosmic microwave background is really as significant as claimed. Working with a historical compilation of 163 published values of the Hubble constant $H_0$ from 1976 to 2019, it finds that the measurements scatter far more than their quoted error bars would allow: the $\\chi^2$ around the weighted mean is 580 for 163 points. The author recalibrates the meaning of an “x-$\\sigma$” deviation using the empirical distribution of the historical sample, obtaining an equivalent Gaussian significance $x_{\\rm eq.}=0.83 x^{0.62}$. On this scale the 4.4-$\\sigma$ tension becomes 2.1 $\\sigma$ (a one-in-28 fluctuation), and even a claimed 6-$\\sigma$ tension becomes 2.5 $\\sigma$. The paper's message is that underestimated error bars have been the norm in $H_0$ measurements, so present-day tensions are not historically unusual and may not require new physics; it adds that the post-2019 uproar likely reflects a sociological phenomenon of groupthink.","feed_headline":"Historical error bars shrink the Hubble tension from 4.4 to 2.1 sigma","feed_subtitle":"A 163-measurement history shows quoted error bars were too small; the claimed 4.4-sigma anomaly is a 1-in-28 fluctuation, not a crisis.","key_machinery":"The load-bearing object is the historical sample of 163 published $H_0$ measurements (1976-2019) with their quoted errors, assembled by an automated search of paper abstracts. For each point the paper computes the deviation $x=|H_0-\\bar{H}_0|/\\sigma$ relative to the inverse-variance weighted mean $\\bar{H}_0=68.26$ km/s/Mpc, then measures the frequency of large deviations. The empirical survival probability is fitted by an exponential, $P(>x)=(0.93\\pm0.06)\\exp[-(0.720\\pm0.013)x]$, and re-expressed as an equivalent number of sigmas in a normal distribution, $x_{\\rm eq.}=(0.830\\pm0.004)x^{0.621\\pm0.003}$. This equivalence curve is the mechanism that converts any claimed tension, stated in units of the quoted error, into the significance it would carry if error bars had been honest; the $\\chi^2=575.7$ for 163 points is the supporting diagnostic showing that the quoted errors are collectively far too small.","core_discovery":"Using 163 $H_0$ measurements published between 1976 and 2019, the paper shows that the dispersion around the inverse-variance weighted average ($\\bar{H}_0 = 68.26 \\pm 0.40$ km/s/Mpc) gives $\\chi^2 = 575.7$ for 163 data points, a chance probability of $Q = 1.0\\times 10^{-47}$; 27 points deviate by more than $2.8\\sigma$ from the mean. The fraction of measurements lying more than $3\\sigma$ away is 11.7%, whereas a Gaussian distribution predicts 0.27%. The tail of the empirical distribution is fitted by $P(>x) = (0.93\\pm0.06)\\exp[-(0.720\\pm0.013)x]$, which is equivalent to a Gaussian significance of $x_{\\rm eq.}=(0.830\\pm0.004)\\,x^{0.621\\pm0.003}$ for $1\\le x\\le 12$. Applying this calibration, the 4.4-$\\sigma$ discrepancy between the 2019 local ladder measurement and the CMB value is a $2.1\\sigma$ effect with probability $P=0.036$ (1 in 28), and a claimed $6\\sigma$ tension would be $2.5\\sigma$ ($P=0.012$, 1 in 83). The paper concludes that $H_0$ tensions of this size have always been present in the literature and are best explained by underestimated statistical error bars or unaccounted systematic errors, not by new physics; the recent attention, it argues, is amplified by conformity within the cosmology community.","pith_inferences":["The same calibration could be carried over to other cosmological parameter tensions (for instance the growth or $S_8$ tension); if the historical error-underestimation statistics are universal, several reported ‘crises’ may weaken to 1–2 sigma effects.","A prospective test would compile independent $H_0$ measurements from 2020 onward that were not used in the fit and check whether their deviations from the current weighted mean are Gaussian; if they are, the calibration curve would not transfer to the modern era, while a heavy tail would confirm that systematic error budgets remain incomplete.","The groupthink explanation could be checked bibliometrically: the paper would predict that the volume and rhetoric of the Hubble-tension literature track the prominence of the teams making the claim, not the objective z-score of the discrepancy.","Because the calibration is fit to the sample containing the very 2019 measurement it later reinterprets, its application to that point is partly self-confirming; fully independent support would come from a curve fitted only to pre-2019 data and then applied to the 2019 and later measurements."],"forward_implications":["If the calibration is correct, the 4.4-sigma Hubble tension is really a 2.1-sigma fluctuation with probability 0.036 (one in 28), well within the range of ordinary statistical noise.","Even the largest widely quoted tensions, around 6 sigma, reduce to about 2.5 sigma (one in 83), so no current claim demands new physics at the traditional discovery threshold.","The historical record implies that error bars on $H_0$ measurements have been systematically underestimated for decades, so the present tension is not a historical anomaly.","The comparison with recent JWST standard-candle measurements, which agree with the CMB value, reinforces the reading that the earlier local-ladder analysis underestimated its errors.","Consequently, a recalibrated significance should be used when judging future $H_0$ comparisons, rather than taking quoted $1\\sigma$ error bars at face value."],"supporting_citations":[{"why":"Supplies the 163-point historical dataset of H0 measurements and quoted errors used for the chi-squared dispersion and tail fits.","marker":"[13]"},{"why":"The earlier analysis whose chi-squared statistic, tail fit, and calibration equations are reproduced in Sections 2-3 of this paper.","marker":"[1]"},{"why":"The 2019 local distance-ladder measurement that defines the 4.4-sigma discrepancy with the CMB value.","marker":"[3]"},{"why":"The cosmic microwave background measurement of H0 under the standard cosmological model that is the comparison baseline for the tension.","marker":"[4]"},{"why":"Reference for the claimed 6-sigma tension obtained with some different datasets, used to show that even the worst-case tension reduces to about 2.5 sigma.","marker":"[5]"},{"why":"Recent JWST standard-candle measurements reported as compatible with the CMB value, cited as evidence that the earlier ladder analysis underestimated its errors.","marker":"[16]"},{"why":"Extends the historical analysis to 2012-2022 data, supporting the claim that underestimated error bars persist into the modern era.","marker":"[2]"}],"fun_headline_variants":["Hubble tension shrinks from 4.4 to 2.1 sigma in historical recalibration","163 H0 measurements expose underestimated errors, tension drops to 2.1 sigma","Hubble tension is not a crisis: history shows it's just underestimated error bars","4.4 sigma Hubble tension recalibrated to 2.1 sigma with 163 data points","Historical H0 error bars inflate tensions; true discrepancy is only 2.1 sigma"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The calibration curve is fitted to the historical sample that includes the very 2019 measurement it is later used to reinterpret, and it assumes the 2019 measurement's error bar suffers the same underestimation statistics as past ones; if modern high-precision measurements are genuinely better calibrated, the recalibrated 2.1-sigma significance would not apply.","fun_headline_variants_meta":{"raw":{"variants":["Hubble tension shrinks from 4.4 to 2.1 sigma in historical recalibration","163 H0 measurements expose underestimated errors, tension drops to 2.1 sigma","Hubble tension is not a crisis: history shows it's just underestimated error bars","4.4 sigma Hubble tension recalibrated to 2.1 sigma with 163 data points","Historical H0 error bars inflate tensions; true discrepancy is only 2.1 sigma"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000563,"raw_usage":{"total_tokens":2842,"prompt_tokens":1289,"completion_tokens":1553,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":905,"completion_tokens_details":{"reasoning_tokens":1439}},"tokens_in":905,"tokens_out":1553,"duration_ms":10892,"temperature":1.0,"reasoning_tokens":1439,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T16:40:21.539545+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compile all independent $H_0$ measurements published from 2020 onward (which were not used to fit the calibration curve), compute each one's deviation from the current inverse-variance weighted mean in units of its quoted error, and count the fraction that exceeds $3\\sigma$. A Gaussian error distribution predicts about 0.27% of points beyond $3\\sigma$; the historical sample gives 11.7%. If the modern fraction is close to 11.7%, the recalibration is confirmed and the 4.4-$\\sigma$ tension should indeed be read as roughly a 2.1-$\\sigma$ event; if the modern fraction is close to 0.27%, the historical curve does not transfer to current-era measurements and the tension stands. The companion extension of the sample to 2012-2022 already provides a partial version of this test.","supporting_citations":[{"cited_title":"Faerber and M","cited_arxiv_id":null,"evidence_quote":"Supplies the 163-point historical dataset of H0 measurements and quoted errors used for the chi-squared dispersion and tail fits."},{"cited_title":"L´ opez-Corredoira, Mon","cited_arxiv_id":null,"evidence_quote":"The earlier analysis whose chi-squared statistic, tail fit, and calibration equations are reproduced in Sections 2-3 of this paper."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The 2019 local distance-ladder measurement that defines the 4.4-sigma discrepancy with the CMB value."},{"cited_title":"Astrophys","cited_arxiv_id":null,"evidence_quote":"The cosmic microwave background measurement of H0 under the standard cosmological model that is the comparison baseline for the tension."},{"cited_title":"Di Valentino, O","cited_arxiv_id":null,"evidence_quote":"Reference for the claimed 6-sigma tension obtained with some different datasets, used to show that even the worst-case tension reduces to about 2.5 sigma."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Extends the historical analysis to 2012-2022 data, supporting the claim that underestimated error bars persist into the modern era."}],"review_version":1}