Pith. sign in

REVIEW 3 major objections 3 minor 21 references

Underestimation of Hubble constant error bars: a historical analysis

T0 review · 3 major / 3 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read After recalibrating with 163 historical measurements, the 4.4-sigma Hubble tension is a 2.1-sigma fluctuation.

desk verdict The historical point is sound, but the statistical bridge from single-residual scatter to the pairwise 4.4σ tension doesn't exist, and most of the content was already published. read the letter →

arxiv 2411.13621 v1 pith:MZIDCTAW submitted 2024-11-20 physics.hist-ph astro-ph.CO

classification physics.hist-phastro-ph.CO
keywords Hubbleconstanttensionerror-barunderestimationcosmologicalparameterestimationstatisticalsignificancemeta-analysisgroupthinkinsciencestandardmodel
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether the much-discussed “Hubble tension” between the local Cepheid-supernova distance ladder and the cosmic microwave background is really as significant as claimed. Working with a historical compilation of 163 published values of the Hubble constant $H_0$ from 1976 to 2019, it finds that the measurements scatter far more than their quoted error bars would allow: the $\chi^2$ around the weighted mean is 580 for 163 points. The author recalibrates the meaning of an “x-$\sigma$” deviation using the empirical distribution of the historical sample, obtaining an equivalent Gaussian significance $x_{\rm eq.}=0.83 x^{0.62}$. On this scale the 4.4-$\sigma$ tension becomes 2.1 $\sigma$ (a one-in-28 fluctuation), and even a claimed 6-$\sigma$ tension becomes 2.5 $\sigma$. The paper's message is that underestimated error bars have been the norm in $H_0$ measurements, so present-day tensions are not historically unusual and may not require new physics; it adds that the post-2019 uproar likely reflects a sociological phenomenon of groupthink.

What carries the argument

The load-bearing object is the historical sample of 163 published $H_0$ measurements (1976-2019) with their quoted errors, assembled by an automated search of paper abstracts. For each point the paper computes the deviation $x=|H_0-\bar{H}_0|/\sigma$ relative to the inverse-variance weighted mean $\bar{H}_0=68.26$ km/s/Mpc, then measures the frequency of large deviations. The empirical survival probability is fitted by an exponential, $P(>x)=(0.93\pm0.06)\exp[-(0.720\pm0.013)x]$, and re-expressed as an equivalent number of sigmas in a normal distribution, $x_{\rm eq.}=(0.830\pm0.004)x^{0.621\pm0.003}$. This equivalence curve is the mechanism that converts any claimed tension, stated in units of the quoted error, into the significance it would carry if error bars had been honest; the $\chi^2=575.7$ for 163 points is the supporting diagnostic showing that the quoted errors are collectively far too small.

What would settle it

Compile all independent $H_0$ measurements published from 2020 onward (which were not used to fit the calibration curve), compute each one's deviation from the current inverse-variance weighted mean in units of its quoted error, and count the fraction that exceeds $3\sigma$. A Gaussian error distribution predicts about 0.27% of points beyond $3\sigma$; the historical sample gives 11.7%. If the modern fraction is close to 11.7%, the recalibration is confirmed and the 4.4-$\sigma$ tension should indeed be read as roughly a 2.1-$\sigma$ event; if the modern fraction is close to 0.27%, the historical curve does not transfer to current-era measurements and the tension stands. The companion extension of the sample to 2012-2022 already provides a partial version of this test.

Watch

Extended reading notes

Core claim

Using 163 $H_0$ measurements published between 1976 and 2019, the paper shows that the dispersion around the inverse-variance weighted average ($\bar{H}_0 = 68.26 \pm 0.40$ km/s/Mpc) gives $\chi^2 = 575.7$ for 163 data points, a chance probability of $Q = 1.0\times 10^{-47}$; 27 points deviate by more than $2.8\sigma$ from the mean. The fraction of measurements lying more than $3\sigma$ away is 11.7%, whereas a Gaussian distribution predicts 0.27%. The tail of the empirical distribution is fitted by $P(>x) = (0.93\pm0.06)\exp[-(0.720\pm0.013)x]$, which is equivalent to a Gaussian significance of $x_{\rm eq.}=(0.830\pm0.004)\,x^{0.621\pm0.003}$ for $1\le x\le 12$. Applying this calibration, the 4.4-$\sigma$ discrepancy between the 2019 local ladder measurement and the CMB value is a $2.1\sigma$ effect with probability $P=0.036$ (1 in 28), and a claimed $6\sigma$ tension would be $2.5\sigma$ ($P=0.012$, 1 in 83). The paper concludes that $H_0$ tensions of this size have always been present in the literature and are best explained by underestimated statistical error bars or unaccounted systematic errors, not by new physics; the recent attention, it argues, is amplified by conformity within the cosmology community.

Load-bearing premise

The calibration curve is fitted to the historical sample that includes the very 2019 measurement it is later used to reinterpret, and it assumes the 2019 measurement's error bar suffers the same underestimation statistics as past ones; if modern high-precision measurements are genuinely better calibrated, the recalibrated 2.1-sigma significance would not apply.

Editorial extensions

If this is right

  • If the calibration is correct, the 4.4-sigma Hubble tension is really a 2.1-sigma fluctuation with probability 0.036 (one in 28), well within the range of ordinary statistical noise.
  • Even the largest widely quoted tensions, around 6 sigma, reduce to about 2.5 sigma (one in 83), so no current claim demands new physics at the traditional discovery threshold.
  • The historical record implies that error bars on $H_0$ measurements have been systematically underestimated for decades, so the present tension is not a historical anomaly.
  • The comparison with recent JWST standard-candle measurements, which agree with the CMB value, reinforces the reading that the earlier local-ladder analysis underestimated its errors.
  • Consequently, a recalibrated significance should be used when judging future $H_0$ comparisons, rather than taking quoted $1\sigma$ error bars at face value.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same calibration could be carried over to other cosmological parameter tensions (for instance the growth or $S_8$ tension); if the historical error-underestimation statistics are universal, several reported ‘crises’ may weaken to 1–2 sigma effects.
  • A prospective test would compile independent $H_0$ measurements from 2020 onward that were not used in the fit and check whether their deviations from the current weighted mean are Gaussian; if they are, the calibration curve would not transfer to the modern era, while a heavy tail would confirm that systematic error budgets remain incomplete.
  • The groupthink explanation could be checked bibliometrically: the paper would predict that the volume and rhetoric of the Hubble-tension literature track the prominence of the teams making the claim, not the objective z-score of the discrepancy.
  • Because the calibration is fit to the sample containing the very 2019 measurement it later reinterprets, its application to that point is partly self-confirming; fully independent support would come from a curve fitted only to pre-2019 data and then applied to the 2019 and later measurements.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper compiles 163 published Hubble–Lemaître constant measurements from 1976 to 2019 and reports that their scatter around the weighted average H0 = 68.26 ± 0.40 km s−1 Mpc−1 gives χ2 = 575.7, far larger than the number of points. Interpreting this as evidence of systematically underestimated error bars, the author fits a heavy-tailed survival function P(|H0,i − H0,avg| > xσ) = 0.93 exp(−0.72x) and an equivalent-Gaussian conversion x_eq = 0.83 x^0.62. The conversion is then applied to the Riess et al. versus Planck 4.4σ tension to claim an equivalent 2.1σ tension with probability P = 0.036, and the paper concludes that Hubble tensions were common before 2019 and that the present-day attention to the Hubble tension reflects sociological groupthink.

Significance. The historical observation that published H0 error bars frequently underestimate the true dispersion is interesting and is supported by the simple chi-squared calculation in Section 2; this part of the paper is a useful reminder of the checkered history of H0 measurements. The paper therefore deserves credit for quantifying, in one transparent statistic, that the historical record is far more scattered than the quoted errors imply. However, the central quantitative claim — that a 4.4σ tension should be read as a 2.1σ tension — rests on two unsupported steps: the equivalent-Gaussian curve is fitted in-sample on the same data point it is used to reinterpret, and it is fitted to single-measurement residuals while being applied to a pairwise comparison. If the historical dispersion result is retained but the recalibration is removed, the paper becomes a valid historical critique rather than a statistical tool for downgrading modern tensions.

major comments (3)
  1. [Section 3, Eq. (3)] The conversion x_eq = 0.83 x^0.62 is fitted to the full 163-point sample, which includes the 2019 Riess et al. measurement listed as the last row of Table 1 (4.1σ from the weighted average). Applying this same fitted curve to reinterpret that very point is in-sample prediction, not independent evidence. The claim that the 4.4σ tension becomes 2.1σ therefore has no out-of-sample support. A leave-one-out or holdout validation, or at least a fit from which the 2019 point is excluded, is needed before this curve can be used to recalibrate that measurement.
  2. [Section 3 versus Sections 1 and 4] The probability distribution in Eq. (2) is defined for x = |H0,i − H0,avg|/σ_i, i.e. the deviation of a single measurement from the sample weighted average. The 4.4σ tension that the paper reinterprets is, however, a pairwise standardized difference T = (74.0 − 67.4)/sqrt(1.4^2 + 0.5^2) ≈ 4.4 between the Riess et al. local measurement and the Planck value. The tail probability of individual residuals does not determine the tail probability of a pairwise difference, because the latter requires the convolution of two error distributions; no such pairwise-tension distribution is derived. Consequently P = 0.036 is not a valid probability for the 4.4σ tension as stated in the abstract and Section 3.
  3. [Section 2, χ2 calculation] The chi-squared analysis treats the 163 measurements as independent, but the paper itself acknowledges that many measurements are incremental updates sharing common foundations. This is not merely a caveat: if the effective number of independent measurements is substantially smaller than 163, then both the quoted Q ≈ 10^−47 and the fitted tail of Eq. (2) are distorted. Since Eq. (2) is the basis of the recalibration, the dependence of the fit on the assumed independence needs to be quantified, for example by removing duplicate chains or by using an effective number of degrees of freedom.
minor comments (3)
  1. [General] The manuscript contains several typographical errors, including 'compililation', 'posslibility', 'understimated', 'Feynmann', and the missing space in 'constantH0' on page 1.
  2. [References] Ref. [19] is a blog post rather than a peer-reviewed source, and large parts of Sections 1, 2, 3, and 4 are excerpts from Ref. [1], the author's own previous paper; this should be stated more clearly in the text so that the reader can distinguish the new material from the reused material.
  3. [Section 3, Eq. (2)] The fitted exponential amplitude 0.93 ± 0.06 is close to but not equal to 1.0 at x = 0; if the survival function is to be interpreted as a properly normalized probability, the behavior near x = 0 should be commented on or the amplitude should be fitted with the normalization constraint.

Circularity Check

1 steps flagged · score 6.0 of 10

The 4.4σ→2.1σ recalibration reduces to an in-sample fit: Eq. (3) is calibrated on the historical sample containing the 2019 Riess point, then applied to that same tension.

  1. fitted input called prediction [Abstract; Section 3, Eqs. (2)–(3); Section 2, Table 1 (last row).]
    "Here we have carried out a recalibration of the probabilities with the present sample of measurements and we find that x-σs deviation is indeed equivalent in a normal distribution to the xeq.σs deviation, where xeq. = 0.83x0.62. Hence, the tension of 4.4-σ, estimated between the local Cepheid-supernova distance ladder and cosmic microwave background (CMB) data, is indeed a 2.1-σ tension in equivalent terms of a normal distribution, with an associated probability P = 0.036 (1 in 28)."

    Eqs. (2)–(3) are fits to the empirical distribution of x = |H0,i − H0|/σi over all 163 historical measurements, including the 2019 Riess point, which Table 1 lists with |H0−H0|/σ = 4.1. The abstract then evaluates the fitted power law at 4.4σ (the Riess–Planck pairwise tension) to obtain 2.1σ. This is in-sample prediction: the recalibration curve is constrained by the very 2019 point whose tension it is used to downgrade, so P = 0.036 is not independent evidence. Moreover, the fitted x is defined as a single measurement's deviation from the sample weighted average, while the 4.4σ target is a pairwise standardized difference between two published values; no distribution for such pairwise differences is derived.

full rationale

The paper's historical observation that H0 measurements are overdispersed relative to quoted error bars is a legitimate empirical finding and is not circular in itself. However, the central quantitative claim—that the 4.4σ Riess–Planck tension is really 2.1σ with P = 0.036—does reduce by construction. The recalibration curve in Eqs. (2)–(3) is fitted to the same 163-point compilation that contains the 2019 Riess measurement; applying that fitted curve to the tension built around that measurement (or to its in-sample 4.1σ residual) is an in-sample prediction rather than an independent test. The calibration statistic is also a residual from the weighted average, whereas the target is a pairwise tension, and the paper supplies no derivation connecting the two distributions. The extensive reuse of the author's own previous compilations and analyses (Refs. [1,13]) is transparent but does not add independent confirmation. These issues warrant a score of 6 (partial circularity), not higher, because the heavy-tailed historical distribution would likely remain even if the single 2019 point were excluded, and the tail statistics are reported rather than hidden.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The central claim rests on the 163-point historical compilation and on a fitted tail distribution. The fitting constants are free parameters, the extrapolation from past to present is a domain assumption, and the 2.8-sigma outlier threshold is chosen post hoc. No new physical entities are introduced; groupthink is a sociological label rather than a modeled entity.

free parameters (5)
  • Exponential amplitude A in Eq. (2) = 0.93 +/- 0.06
    Fitted to observed frequency of deviations in the historical sample; used in the recalibration.
  • Exponential decay constant lambda in Eq. (2) = 0.720 +/- 0.013
    Fitted to the same historical sample; defines the tail distribution.
  • Conversion scale c in x_eq = c x^beta = 0.830 +/- 0.004
    Derived fit to the same sample; maps reported sigma to equivalent Gaussian sigma.
  • Conversion exponent beta = 0.621 +/- 0.003
    Derived fit to the same sample; controls the compression of large sigma values.
  • Outlier threshold = > 2.8 sigma
    Chosen by hand so that removing outliers yields Q>=0.05; post hoc selection criterion.
assumptions (5)
  • standard math Chi-squared goodness-of-fit with N-1 degrees of freedom and Gaussian errors are used to quantify dispersion.
    Invoked in Section 2, Eq. (1), to compute the probability Q of the observed chi-squared.
  • domain assumption The ADS abstract-based 163-point compilation is representative and unbiased.
    Stated in Section 2: 'this is representative and should not produce any statistical bias'.
  • domain assumption Historical error-underestimation statistics can be transferred to modern measurements, including Riess et al. 2019.
    Implicit in Sections 3-4, where Eq. (3) is applied to the 4.4-sigma tension.
  • domain assumption Covariance between measurements is neglected.
    Acknowledged in Section 2: 'neglect of the covariance of each observed measurement' is a simplifying assumption.
  • standard math The weighted average H0 = 68.26 is treated as the reference value for computing deviations.
    Used throughout Section 2 to define deviations and outliers.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Underestimation of Hubble constant error bars: a historical analysis." pith.science (2026). https://pith.science/paper/MZIDCTAW

@misc{pith2026241113621,
  author       = {Pith},
  title        = {Pith review of: Underestimation of Hubble constant error bars: a historical analysis},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MZIDCTAW}},
  note         = {Machine review of arXiv:2411.13621}
}
abstract

An analysis of a historical compilation of Hubble-Lema\^itre constant values ($H_0$: 163 data points measured between 1976 and 2019) assuming the standard cosmological model gives a $\chi^2$ value of the dispersion with respect to the weighted average of 580, much larger than the number of points, which has an associated probability that is very low. This means that Hubble tensions were always present in the literature, due either to the underestimation of statistical error bars associated with the observed parameter measurements, or to the fact that systematic errors were not properly taken into account in many of the measurements. The fact that the underestimation of error bars for $H_0$ is so common might explain the apparent 4.4-sigma discrepancy by Riess et al. As a matter of fact, more recent precise $H_0$ measurements with JWST data by Freedman et al. using standard candles in galaxies find there is no tension with CMBR data, possibly indicating that previously Riess et al. had underestimated their errors. Here we have carried out a recalibration of the probabilities. The tension of 4.4-$\sigma $, estimated between the local Cepheid-supernova distance ladder and cosmic microwave background (CMB) data, is indeed a 2.1-$\sigma $ tension in equivalent terms of a normal distribution, with an associated probability $P$ = 0.036 (1 in 28). This can be increased to an equivalent tension of 2.5-$\sigma $ in the worst cases of claimed 6-$\sigma $ tension, which may in any case happen as a random statistical fluctuation. If Hubble tensions were always present in the literature, and present day tensions are not more important than previous ones, why, then, is there so much noise and commotion surrounding Hubble tension after 2019? It is suggested here that this obeys a sociological phenomenon of ``groupthink''.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

21 extracted references · 19 canonical work pages

  1. [1]

    L´ opez-Corredoira, Mon

    M. L´ opez-Corredoira, Mon. Not. R. Astron. Soc. 517, 5805 (2022)

  2. [2]

    B. Wang, M. L´ opez-Corredoira and J.-J. Wei, Mon. Not. R. Astron. Soc. 527, 7692 (2024)

  3. [3]

    A. G. Riess, S. Casertano, W. Yuan, L. M. Macri and D. Scolni c, Astrophys. J. 876, 85 (2019)

  4. [4]

    Astrophys

    Planck Collaboration et al., 2020, Astron. Astrophys. 641, A6 (2020)

  5. [5]

    Di Valentino, O

    E. Di Valentino, O. Mena and S. Pan, Class. Quantum Grav. 38, 153001 (2021)

  6. [6]

    G. Chen, J. R. I. Gott and B. Ratra B., Publ. Astron. Soc. Pacific , 115, 1269 (2003)

  7. [7]

    J. Gott, I. Richard, M. S. Vogeley, S. Podariu and B. Ratra, Astrophys. J. 549, 1 (2001)

  8. [8]

    G. A. Tammann, A. Sandage and B. Reindl Astron. Astrophys. 404, 423 (2003)

Show all 21 references
  1. [9]

    D. R. Matravers, G. F. R. Ellis and W. R. Stoeger, Quarterly J. R. Astron. Soc. 36, 29 (1995)

  2. [10]

    L´ opez-Corredoira, Int

    M. L´ opez-Corredoira, Int. J. Mod. Phys. D 22, 1350032 (2013)

  3. [11]

    Axelsson, H

    M. Axelsson, H. T. Ihle, S. Scodeller and F. K. Hansen, Astron. Astrophys. 578, A44 (2015)

  4. [12]

    Creswell and P

    J. Creswell and P. Naselsky, J. Cosmol. Astropart. Phys. 2021(3), 103 (2021)

  5. [13]

    Faerber and M

    T. Faerber and M. L´ opez-Corredoira M., Universe 6, 114 (2020)

  6. [14]

    W. L. Freedman, B. F. Madore, B. K. Gibson, et al., Astrophys. J. 553, 47 (2001)

  7. [15]

    B. M. Leith, S. C. C. Ng, D. L. Wiltshire, Astrophys. J. Lett. 672, L91 (2008)

  8. [16]

    W. L. Freedman, B. F. Madore, I. S. Jang, T. J. Hoyt, A. J. Le e and K. A. Owens, arXiv.org, 2408.06153 (2024)

  9. [17]

    A. G. Riess, D. Scolnic, G. S. Anand, et al. arXiv.org, 2408.11770 (2024)

  10. [18]

    R. P. Feynman, Engineering and Science 37, 10 (1974)

  11. [19]

    L´ opez-Corredoira, Science 2.0 , December 22nd (2023)

    M. L´ opez-Corredoira, Science 2.0 , December 22nd (2023)

  12. [20]

    L´ opez-Corredoira, Fundamental Ideas in Cosmology

    M. L´ opez-Corredoira, Fundamental Ideas in Cosmology. Scientific, philosophical and sociological critical perspectives (IoP Science, 2022)

  13. [21]

    C. R. Sunstein, Conformity. The Power of Social Influences (NYU Press, 2021)

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.