Pith. sign in

REVIEW 4 major objections 4 minor 33 references

When correcting for regression to the mean is worse than no correction at all

T0 review · 4 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read This paper shows that the most widely used correction for regression to the mean—the Berry et al. slope adjustment—is itself biased and can create false positives and false negatives, and argues that the uncorrected crude slope tested again

desk verdict The repeatability-null recipe is worth stealing, but the paper's algebra on the Berry correction is wrong in ways that undermine its central theoretical claims. read the letter →

arxiv 2509.04718 v2 pith:VYWI7GOQ submitted 2025-09-05 stat.ME physics.data-anq-bio.QM

classification stat.MEphysics.data-anq-bio.QM MSC 62F0362J0562F40
keywords regressiontothemeanmeasurementerrorhypothesistestingbootstraprepeatabilitydifferentialtreatmenteffectchangescoresstructurallinearmodel
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the most popular correction for regression to the mean (RTM), the Berry et al. slope adjustment popularized in ecology by Kelly and Price, does more harm than good: its corrected slope is systematically biased and can produce both false positives and false negatives, because it only recovers the true slope in an unrealistic special case. The authors build a simple structural model in which the true post-test value is a linear function of the pre-test value plus a treatment effect and noise, and measured values add independent error. Within that model they show the uncorrected crude slope has a clean bias expressible through repeatability, the proportion of total variance due to true individual differences, and they propose testing the crude slope against that structural null expectation using bootstrap. They reanalyze published lizard and bird datasets to show that conclusions change depending on the assumed repeatability. The punchline is that without knowing, or at least bounding, an experiment's repeatability, claims about differential treatment effects are statistically unfounded.

What carries the argument

The structural change model X2 = X1 + (α+βX1) + ξ, with independent measurement noise ϵ₁,ϵ₂ ~ N(0,δ²), is the engine of the paper. It yields the exact identity δ²/σ²₁ = 1−R, where R is repeatability, and the crude-slope formula βc = β − (1+β)δ²/(σ²+δ²). This identity does the work: it turns the unobservable measurement-error variance into a quantity researchers can estimate or bound, and converts the null hypothesis β=0 into the testable claim βc = −δ²/σ²₁ = R−1.

What would settle it

Simulate two-time-point data with known true slope β, known between-subject treatment variance ν²>0, and known measurement-error variance δ², then estimate the Berry et al. slope βB on many samples; if the average βB equals β for any parameter combination with ν²>0, equation (18)—the claim that the correction is biased except at β=ν²=0—is false. Equivalently, a real dataset with an external gold-standard measure of the trait could be used to compare corrected vs. true slopes across repeatability values.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is a closed-form comparison of three estimates of the slope of true change on true initial value: the crude slope βc, the Berry et al. corrected slope βB, and the Blomqvist slope βe. For the crude slope, βc = β − (1+β)δ²/(σ²+δ²), so the RTM bias is entirely driven by within-subject variance δ² and is independent of the stochastic variance ν² of the treatment effect. For the Berry et al. correction, βB = −ρ + (1+β)σ²/(σ²+δ²) = βc + (1−ρ), which is always greater than or equal to the crude slope and equals the true slope only when β = ν² = 0. The Blomqvist correction is algebraically unbiased but has roughly twice the sampling variance of the cru

Load-bearing premise

The load-bearing premise is that the true post-test value is a linear function of the true pre-test value plus treatment effect and independent noise, and that measured pre- and post-values add independent, equal-variance normal error; if change is nonlinear in the initial value, or the errors are correlated or heteroscedastic across time, the paper's formulas for the crude, corrected, and null slopes no longer hold.

Editorial extensions

If this is right

  • If the paper is right, every conclusion drawn from the Kelly–Price/Berry et al. adjustment about differential treatment effects is unsupported, because the corrected slope is biased whenever there is any between-subject variation in treatment effect (ν²>0).
  • The Blomqvist method cannot be recommended as a routine fix: although unbiased in principle, its high sampling variance makes it less reliable than the biased crude slope in sample sizes typical of ecology and physiology.
  • A practical testing recipe emerges: compute the crude slope, bootstrap its confidence interval, and reject a differential treatment effect only if the interval excludes R−1 for the experiment's expected repeatability.
  • Re-analysis of existing datasets (lizards, birds) may overturn published conclusions: evidence for or against trade-offs and biomarkers depends on repeatability values that are often not reported.
  • Reporting repeatability becomes a prerequisite for valid inference about change; readers should demand it.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the structural model is approximately right, the same logic extends to designs with more than two time points: repeatability estimated from repeated measures could replace the qualitative bound with a direct null test, a step the paper leaves implicit.
  • The paper's framework suggests a simple diagnostic for any published RTM-corrected result: check whether a plausible repeatability range places the reported corrected slope's null value inside the bootstrap interval; if not, the inference is fragile.
  • By reframing RTM as a problem of knowing R, the paper connects to the measurement-error and latent-variable literature, where the same identifiability problem (σ² vs δ² from two observations) is well known; a Bayesian or mixed-model approach with informative priors on R might be a natural next step.
  • A quantitative prediction follows: meta-analyses of traits with low repeatability (R<0.5) should show apparent negative slopes between initial value and change even when no differential treatment effect exists, and published positive results should cluster in low-repeatability studies.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes a structural two-level model in which true pre- and post-test values follow X2 = X1 + (α + βX1) + ξ and measurements add independent normal noise with variance δ² at both time points. Within this model it derives population expressions for the crude slope βc, the Berry et al. corrected slope βB, and the Blomqvist estimator βe, and then uses simulations and two empirical case studies to argue that (i) the Berry correction is systematically biased and unreliable for hypothesis testing, (ii) the Blomqvist estimator, though unbiased, has high sampling variance and is often worse than the crude slope, and (iii) the most practical approach is to test the crude slope against a repeatability-based null expectation. The paper claims that βB equals the true slope only in the highly restrictive case β = ν² = 0, and uses Eq. (23) and Eq. (25) to support the bias and false-positive narrative.

Significance. If the derivations were correct, the paper would be a useful cautionary contribution: it connects the RTM bias directly to repeatability, gives a simple bootstrap recipe for the null β = 0, and applies it to published datasets where earlier Kelly-Price-style corrections were used. The algebraic structure is transparent, no parameters are fitted to force the conclusions, and the empirical examples are concrete. However, the central mathematical characterization of the Berry correction contains load-bearing errors: the 'only at β = ν² = 0' claim is false, and the limit formulas in Eqs. (23) and (25) are misstated. These issues need to be corrected and the corresponding narrative revised before the paper's conclusions can be accepted.

major comments (4)
  1. [Correcting for RTM using the Berry et al. method, Eq. (18)] The statement that βB = β only when β = ν² = 0 is false. Setting βB = β in Eq. (18) gives ρ = (σ² − βδ²)/(σ² + δ²). Combining this with Eq. (17) yields a one-dimensional family of non-null parameters, e.g. σ² = δ² = 1, β = 0.1, ν² ≈ 0.778 gives βB = β = 0.1. The paper's own δ² = 0 limit also admits βB = β for all β > −1 when ν² = 0. The unbiasedness condition is therefore a continuum, not a point. This does not necessarily rescue the Berry correction as a general method, but the current claim and its downstream wording are incorrect and must be revised.
  2. [Testing for a Differential Treatment Effect, Eq. (25)] Under the null β = 0, the correct expression is βB = σ²/σ₁² [1 − 1/sqrt(1 + ν²/σ₁²)], which is positive for ν² > 0. Equation (25) as printed has δ² in place of σ² and the opposite sign, and the accompanying statement 'This expression is always negative for ν² > 0' is wrong. The qualitative warning that the Berry slope is nonzero under the null survives, but the sign and the false-positive direction are misdescribed. The text should be corrected and the implications re-stated.
  3. [Analysis of the population regression slopes, Eq. (23)] For δ² = 0, the correct limit of Eq. (18) is βB = 1 + β − sgn(1 + β)/sqrt(1 + ν²/[(1 + β)²σ²]). The displayed Eq. (23) lacks the reciprocal and gives β + 1 − sgn(1 + β) * sqrt(1 + ν²/[(1 + β)²σ²]), which is not the limit. For example, for small ν² and β > −1 the true limiting slope exceeds β by O(ν²), whereas the printed expression behaves differently. This is not a typo in an ancillary remark; it is the equation used to describe the method's behavior in the no-measurement-error regime.
  4. [Sample Size Effects on Regression Slopes, Figures 2–3] The general claim that the Blomqvist estimator is practically inferior to the crude slope rests on a single simulation configuration: N = 100, the systolic blood pressure parameters, δ²/σ² = 0.45, and one value of ν/σ. The variance comparison and the crossing probabilities in Figure 3 may depend on these choices. To support a general recommendation against Blomqvist, the authors should either derive an analytic expression for the sampling variance or provide a systematic exploration over N, repeatability, and ν²/σ². At minimum, the abstract's blanket statement should be tempered.
minor comments (4)
  1. [Telomere case study] In the paragraph after Figure 6, 'The Barry et al. method' should read 'The Berry et al. method'.
  2. [Throughout] The repeatability R is introduced implicitly as R = 1/(1 + δ²/σ²). An explicit equation in the model section would help readers see why the null value R − 1 appears.
  3. [Bootstrap procedure] The proposed null test requires a chosen value of R, and the conclusion is conditional on that choice. The paper acknowledges this, but it would be clearer to state explicitly that the resulting p-value/decision is not a test of β = 0 without an externally justified repeatability.
  4. [Model assumptions, Eqs. (1)–(3)] The derivations assume equal error variance at pre and post and independent errors. The authors should add a sentence noting that correlated or heteroscedastic measurement error changes the formulas, and that the empirical recommendations are scoped to the stated model.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper's bias derivations are algebraic consequences of an explicitly stated model; self-citations are peripheral, not load-bearing.

full rationale

The paper's derivation chain is self-contained. Starting from the explicit generative model (Eqs. 1–3), it computes the crude slope (Eq. 13), the Berry et al. corrected slope (Eqs. 15–18), and the Blomqvist estimator (Eqs. 20–22) by direct covariance algebra. No parameter is fitted to force the conclusion; the repeatability R enters as an external input from prior empirical studies or from stated hypothetical ranges, not as a fitted quantity. The citations to the authors' own prior work (Santos & Fontanari 2025) appear only as examples of the ρ=0/permutation null that the paper criticizes, and as background; they are not premises of the central algebra. The central claim—that Berry et al.'s corrected slope is biased—is derived from the model rather than assumed. The skeptical objection about the truth of Eq. (18)'s consequence ('only at β=ν²=0') is a mathematical/correctness challenge, not a circularity: it does not show that a prediction reduces to an input by construction. Therefore no circular step is exhibited.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The paper introduces no new entities or forces. It relies on a standard measurement-error model, a linear true-change model, and external repeatability estimates. The central derivations are parameter-free algebra, while the simulations and empirical examples require hand-set or literature-supplied variance components.

free parameters (3)
  • repeatability R (or noise ratio δ²/σ²) = 0.69 (blood pressure), 0.479/0.398 (telomere), hypothetical ranges in lizard example
    The proposed null value for the crude slope is R − 1; R is not estimable from the two time points in the data and is taken from external literature or assumed.
  • between-subject treatment variation ν² = ν = 10 mmHg in simulations
    Set arbitrarily in the simulations; it affects the Berry bias magnitude and the Blomqvist variance but is not identifiable from two time points.
  • additive treatment offset α = -20 mmHg
    Follows Hayes (1988); slopes are independent of α, so it is illustrative only.
assumptions (5)
  • domain assumption True pre-test values are normal, X1 ~ N(μ, σ²)
    Used for Eqs. (4)-(5); required for the variance decomposition and the crude slope formula.
  • domain assumption True post-test value is linear in true pre-test with additive noise, X2 = X1 + (α + βX1) + ξ
    Eq. (1); defines β as the slope of true change on true initial value and underpins all derived bias formulas.
  • domain assumption Measurement errors are independent, normal, equal variance δ², independent of X1 and ξ
    Eqs. (2)-(3); the crude slope bias, the repeatability null R − 1, and the Berry/Blomqvist corrections all rely on this.
  • standard math Bootstrap resampling of the empirical sample gives valid confidence intervals for the crude slope
    Used in Figs. 4, 5, and 7 to test hypotheses; assumes i.i.d. sampling and large enough N.
  • domain assumption External repeatability estimates (e.g., from Wolak 2012, Kärkkäinen 2022) transfer to the target population
    The testing recipe substitutes literature repeatability for the unidentifiable within-study R; this is acknowledged in the text but is load-bearing for empirical conclusions.

how reviews work

0 comments
Cite this review

Pith. "Pith review of When correcting for regression to the mean is worse than no correction at all." pith.science (2026). https://pith.science/paper/VYWI7GOQ

@misc{pith2026250904718,
  author       = {Pith},
  title        = {Pith review of: When correcting for regression to the mean is worse than no correction at all},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VYWI7GOQ}},
  note         = {Machine review of arXiv:2509.04718}
}
read the original abstract

The ubiquitous regression to the mean (RTM) effect complicates statistical inference regarding the relationship between baseline levels of a biological variable and its subsequent change. We demonstrate that common RTM correction methods are problematic: the Berry et al. method, popularized by Kelly & Price in The American Naturalist, is unreliable for hypothesis testing or effect-size estimation, leading to systematic bias and inflated error rates. Conversely, while the Blomqvist method is theoretically unbiased, its high sampling variance limits its practical utility in small-to-moderate datasets. Using a structural linear model, we show that the most robust approach to navigating RTM is not to correct the data, but to evaluate the uncorrected crude slope against a structural null expectation derived from measurement repeatability-the proportion of total variance attributable to true individual differences. We illustrate this approach using empirical data from studies on lizard thermal physiology and bird telomere dynamics. Ultimately, we argue that any conclusion regarding a differential treatment effect is statistically unfounded without a clear understanding of the experiment's repeatability.

Figures

Figures reproduced from arXiv: 2509.04718 by the authors.

Figure 1
Figure 1. Crude (βc) and Berry et al. (βB) estimates of the true slope as function of the ratio of within-subject to between-subject variance. The left panel shows β = 0, the middle panel shows β = −0.5, and the right panel shows β = −1.5. The true slopes are shown as horizontal lines. The other parameters are µ = 141, σ = 13.6, α = −20, and ν = 10 [PITH_FULL_IMAGE:figures/full_fig_p013_1.png] view at source ↗
Figure 2
Figure 2. Distribution of the estimates of the regression slopes for [PITH_FULL_IMAGE:figures/full_fig_p015_2.png] view at source ↗
Figure 3
Figure 3. Probability that the crude slope (βc) yields a better estimate of the true slope β than the Berry et al. (βB) correction (left panel) and than the Blomqvist (βe) correction (right panel) as function of β. The other parameters are µ = 141, σ = 13.6, δ = 9.1, α = −20, and ν = 10. 16 [PITH_FULL_IMAGE:figures/full_fig_p016_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: The empirical regression slope is βc = −0.423. From our population 18 [PITH_FULL_IMAGE:figures/full_fig_p018_4.png]
Figure 4
Figure 4. Figure 4: Procedure to test the null hypothesis β = 0 using the crude slope. The left panel shows the scatter plot of change d = x2 − x1 against initial value x1 in a computer-simulated sample of size N = 100 for systolic blood pressure. The empirical slope of the regression lin…
Figure 5
Figure 5. Figure 5: Analysis of heat tolerance plasticity for [PITH_FULL_IMAGE:figures/full_fig_p021_5.png]
Figure 6
Figure 6. Figure 6: Telomere attrition rates adjusted for the RTM effect. The left panel shows the scatter [PITH_FULL_IMAGE:figures/full_fig_p023_6.png]
Figure 7
Figure 7. Figure 7: Bootstrap histograms for the crude and Blomqvist regression slopes. Histograms [PITH_FULL_IMAGE:figures/full_fig_p024_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

33 extracted references · 33 canonical work pages

  1. [1]

    Archie, J. P. 1981. Mathematic coupling of data: a common source of error. Annals of Surgery 193:296--303

  2. [2]

    Berry, D. A., M. L. Eaton, B. P. Ekholm, and T. L. Fox. 1984. Assessing differential drug effect. Biometrics 40:1109--1115

  3. [3]

    Blomqvist, N. 1977. On the relation between change and initial value. Journal of the American Statistical Association 72:746--749

  4. [4]

    Castaneda, L. E., G. Calabria, L. A. Betancourt, E. L. Rezende, and M. Santos. 2012. Measurement error in heat tolerance assays. Journal of Thermal Biology 37:432--437

  5. [5]

    Paradis, B

    Chiolero, A., G. Paradis, B. Rich, and J. A. Hanley. 2013. Assessing the relationship between the baseline value of a continuous variable and subsequent change over time. Frontiers in Public Health 1:29

  6. [6]

    a, L. Hillstr\

    Cicho\'n, M., J. Meril\"a, L. Hillstr\"om, and D. Wiggins. 1999. Mass-dependent mass loss in breeding birds: getting the null hypothesis right. Oikos 87:191--194

  7. [7]

    Chuang-Stein, C. 1993. The regression fallacy. Drug Information Journal 27:1213--1220

  8. [8]

    Deery, S. A., J. E. Rej, D. Haro, and A. R. Gunderson. 2021. Heat hardening in a pair of Anolis lizards: constraints, dynamics and ecological consequences. Journal of Experimental Biology 224:jeb240994

Show all 33 references
  1. [9]

    Efron, B., and R. J. Tibshirani. 1993. An Introduction to the Bootstrap. Chapman & Hall, New York

  2. [10]

    Ferrer, E., and J. J. McArdle. 2010. Longitudinal modeling of developmental changes in psychological research. Current Directions in Psychological Science, 19:149--154

  3. [11]

    Wagenmakers, and T

    Forstmeier W., E.-J. Wagenmakers, and T. H. Parker. 2017. Detecting and avoiding likely false-positive findings - a practical guide. Biological Reviews 92:1941--1968

  4. [12]

    Galton, F. 1886. Regression towards mediocrity in hereditary stature. The Journal of the Anthropological Institute of Great Britain and Ireland 15:246--263

  5. [13]

    Gardner, M. J. , and J. A. Heady. 1973. Some effects of within-person variability in epidemiological studies. Journal of Chronic Diseases 26:781--795

  6. [14]

    Geenen, R., and F. J. R. van de Vijver. 1993. A simple test of the law of initial values. Psychophysiology 30:525--530

  7. [15]

    Gunderson, A. R. 2023. Trade-offs between baseline thermal tolerance and thermal tolerance plasticity are much less common than it appears. Global Change Biology, 29:3519--3524

  8. [16]

    Hanushek, E. A., L. Kinne, F. Witthöft, and L. Woessmann. 2025. Age and cognitive skills: Use it or lose it. Sci. Adv. 11:eads1560

  9. [17]

    Hayes, R. J. 1988. Methods for assessing whether change depends on initial value. Statistics in Medicine 7:915--927

  10. [18]

    A., and K

    Jackson, D. A., and K. M. Somers. 1991. The spectre of 'spurious' correlations. Oecologia 86:147--151

  11. [19]

    Briga, T

    Kärkkäinen, T., M. Briga, T. Laaksonen, and A. Stier. 2022. Within-individual repeatability in telomere length: A meta-analysis in nonmammalian vertebrates. Molecular Ecology 31:6339--6359

  12. [20]

    Kelly, C., and T. D. Price. 2005. Correcting for regression to the mean in behavior and ecology. American Naturalist 166:700--707

  13. [21]

    Kronmal, R. A. 1993. Spurious correlation and the fallacy of the ratio standard revisited. Journal of the Royal Statistical Society A 156:379--392

  14. [22]

    Diekmann

    Mazalla, L., and M. Diekmann. 2022. Regression to the mean in vegetation science. Journal of Vegetation Science 33:e13117

  15. [23]

    Nesselroade, J. R., S. M. Stigler, and P. B. Baltes. 1980. Regression toward the mean and the study of change. Psychological Bulletin 88:622--637

  16. [24]

    Pearson, K. 1897. Mathematical contributions to the theory of evolution. -On a form of spurious correlation which may arise when indices are used in the measurement of organs. Proceedings of the Royal Society of London 60:489--497

  17. [25]

    Pitman, E. J. G. 1939. A note on normal correlation. Biometrika 31:9--12

  18. [26]

    Brandt, and M

    Rogosa, D., D. Brandt, and M. Zimowski. 1982. A growth curve approach to the measurement of change. Psychological Bulletin, 92:726--748

  19. [27]

    Santos, M., J. F. Fontanari. 2025. On testing the tolerance-plasticity trade-off hypothesis as the change of thermal tolerance across two environments. Journal of Thermal Biology, 132:104248

  20. [28]

    Slessarev E. W., A. Mayer, C. Kelly, K. Georgiou, J. Pett-Ridge, and E. E. Nuccio. 2023. Initial soil organic carbon stocks govern changes in soil carbon: reality or artifact? Global Change Biology 29:1239--1247

  21. [29]

    Ingre, G

    Sorjonen, K., M. Ingre, G. Nilsonne, and B. Melin. 2023. Dangers of including outcome at baseline as a covariate in latent change score models: results from simulations and empirical re-analyses. Heliyon 9:e15746

  22. [30]

    M., Gustafsson, L

    Sudyka, J., Arct, A., Drobniak, S. M., Gustafsson, L. , and Cicho\'n, M. 2019. Birds with high lifetime reproductive success experience increased telomere loss. Biol. Lett. 15:20180637

  23. [31]

    Wasserman, L. 2004. All of Statistics: A Concise Course in Statistical Inference. Springer, New York

  24. [32]

    Wilder, J. 1967. Stimulus and response: the law of initial value. John Wright & Sons LTD, Bristol

  25. [33]

    Wolak, M. E., D. J. Fairbairn, and Y. R. Paulsen. 2012. Guidelines for estimating repeatability. Methods in Ecology and Evolution 3:129–137

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.