REVIEW 4 major objections 4 minor 33 references
When correcting for regression to the mean is worse than no correction at all
T0 review · 4 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read This paper shows that the most widely used correction for regression to the mean—the Berry et al. slope adjustment—is itself biased and can create false positives and false negatives, and argues that the uncorrected crude slope tested again
desk verdict The repeatability-null recipe is worth stealing, but the paper's algebra on the Berry correction is wrong in ways that undermine its central theoretical claims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The structural change model X2 = X1 + (α+βX1) + ξ, with independent measurement noise ϵ₁,ϵ₂ ~ N(0,δ²), is the engine of the paper. It yields the exact identity δ²/σ²₁ = 1−R, where R is repeatability, and the crude-slope formula βc = β − (1+β)δ²/(σ²+δ²). This identity does the work: it turns the unobservable measurement-error variance into a quantity researchers can estimate or bound, and converts the null hypothesis β=0 into the testable claim βc = −δ²/σ²₁ = R−1.
What would settle it
Simulate two-time-point data with known true slope β, known between-subject treatment variance ν²>0, and known measurement-error variance δ², then estimate the Berry et al. slope βB on many samples; if the average βB equals β for any parameter combination with ν²>0, equation (18)—the claim that the correction is biased except at β=ν²=0—is false. Equivalently, a real dataset with an external gold-standard measure of the trait could be used to compare corrected vs. true slopes across repeatability values.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is a closed-form comparison of three estimates of the slope of true change on true initial value: the crude slope βc, the Berry et al. corrected slope βB, and the Blomqvist slope βe. For the crude slope, βc = β − (1+β)δ²/(σ²+δ²), so the RTM bias is entirely driven by within-subject variance δ² and is independent of the stochastic variance ν² of the treatment effect. For the Berry et al. correction, βB = −ρ + (1+β)σ²/(σ²+δ²) = βc + (1−ρ), which is always greater than or equal to the crude slope and equals the true slope only when β = ν² = 0. The Blomqvist correction is algebraically unbiased but has roughly twice the sampling variance of the cru
Load-bearing premise
The load-bearing premise is that the true post-test value is a linear function of the true pre-test value plus treatment effect and independent noise, and that measured pre- and post-values add independent, equal-variance normal error; if change is nonlinear in the initial value, or the errors are correlated or heteroscedastic across time, the paper's formulas for the crude, corrected, and null slopes no longer hold.
Editorial extensions
If this is right
- If the paper is right, every conclusion drawn from the Kelly–Price/Berry et al. adjustment about differential treatment effects is unsupported, because the corrected slope is biased whenever there is any between-subject variation in treatment effect (ν²>0).
- The Blomqvist method cannot be recommended as a routine fix: although unbiased in principle, its high sampling variance makes it less reliable than the biased crude slope in sample sizes typical of ecology and physiology.
- A practical testing recipe emerges: compute the crude slope, bootstrap its confidence interval, and reject a differential treatment effect only if the interval excludes R−1 for the experiment's expected repeatability.
- Re-analysis of existing datasets (lizards, birds) may overturn published conclusions: evidence for or against trade-offs and biomarkers depends on repeatability values that are often not reported.
- Reporting repeatability becomes a prerequisite for valid inference about change; readers should demand it.
Reading between the lines
- If the structural model is approximately right, the same logic extends to designs with more than two time points: repeatability estimated from repeated measures could replace the qualitative bound with a direct null test, a step the paper leaves implicit.
- The paper's framework suggests a simple diagnostic for any published RTM-corrected result: check whether a plausible repeatability range places the reported corrected slope's null value inside the bootstrap interval; if not, the inference is fragile.
- By reframing RTM as a problem of knowing R, the paper connects to the measurement-error and latent-variable literature, where the same identifiability problem (σ² vs δ² from two observations) is well known; a Bayesian or mixed-model approach with informative priors on R might be a natural next step.
- A quantitative prediction follows: meta-analyses of traits with low repeatability (R<0.5) should show apparent negative slopes between initial value and change even when no differential treatment effect exists, and published positive results should cluster in low-repeatability studies.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a structural two-level model in which true pre- and post-test values follow X2 = X1 + (α + βX1) + ξ and measurements add independent normal noise with variance δ² at both time points. Within this model it derives population expressions for the crude slope βc, the Berry et al. corrected slope βB, and the Blomqvist estimator βe, and then uses simulations and two empirical case studies to argue that (i) the Berry correction is systematically biased and unreliable for hypothesis testing, (ii) the Blomqvist estimator, though unbiased, has high sampling variance and is often worse than the crude slope, and (iii) the most practical approach is to test the crude slope against a repeatability-based null expectation. The paper claims that βB equals the true slope only in the highly restrictive case β = ν² = 0, and uses Eq. (23) and Eq. (25) to support the bias and false-positive narrative.
Significance. If the derivations were correct, the paper would be a useful cautionary contribution: it connects the RTM bias directly to repeatability, gives a simple bootstrap recipe for the null β = 0, and applies it to published datasets where earlier Kelly-Price-style corrections were used. The algebraic structure is transparent, no parameters are fitted to force the conclusions, and the empirical examples are concrete. However, the central mathematical characterization of the Berry correction contains load-bearing errors: the 'only at β = ν² = 0' claim is false, and the limit formulas in Eqs. (23) and (25) are misstated. These issues need to be corrected and the corresponding narrative revised before the paper's conclusions can be accepted.
major comments (4)
- [Correcting for RTM using the Berry et al. method, Eq. (18)] The statement that βB = β only when β = ν² = 0 is false. Setting βB = β in Eq. (18) gives ρ = (σ² − βδ²)/(σ² + δ²). Combining this with Eq. (17) yields a one-dimensional family of non-null parameters, e.g. σ² = δ² = 1, β = 0.1, ν² ≈ 0.778 gives βB = β = 0.1. The paper's own δ² = 0 limit also admits βB = β for all β > −1 when ν² = 0. The unbiasedness condition is therefore a continuum, not a point. This does not necessarily rescue the Berry correction as a general method, but the current claim and its downstream wording are incorrect and must be revised.
- [Testing for a Differential Treatment Effect, Eq. (25)] Under the null β = 0, the correct expression is βB = σ²/σ₁² [1 − 1/sqrt(1 + ν²/σ₁²)], which is positive for ν² > 0. Equation (25) as printed has δ² in place of σ² and the opposite sign, and the accompanying statement 'This expression is always negative for ν² > 0' is wrong. The qualitative warning that the Berry slope is nonzero under the null survives, but the sign and the false-positive direction are misdescribed. The text should be corrected and the implications re-stated.
- [Analysis of the population regression slopes, Eq. (23)] For δ² = 0, the correct limit of Eq. (18) is βB = 1 + β − sgn(1 + β)/sqrt(1 + ν²/[(1 + β)²σ²]). The displayed Eq. (23) lacks the reciprocal and gives β + 1 − sgn(1 + β) * sqrt(1 + ν²/[(1 + β)²σ²]), which is not the limit. For example, for small ν² and β > −1 the true limiting slope exceeds β by O(ν²), whereas the printed expression behaves differently. This is not a typo in an ancillary remark; it is the equation used to describe the method's behavior in the no-measurement-error regime.
- [Sample Size Effects on Regression Slopes, Figures 2–3] The general claim that the Blomqvist estimator is practically inferior to the crude slope rests on a single simulation configuration: N = 100, the systolic blood pressure parameters, δ²/σ² = 0.45, and one value of ν/σ. The variance comparison and the crossing probabilities in Figure 3 may depend on these choices. To support a general recommendation against Blomqvist, the authors should either derive an analytic expression for the sampling variance or provide a systematic exploration over N, repeatability, and ν²/σ². At minimum, the abstract's blanket statement should be tempered.
minor comments (4)
- [Telomere case study] In the paragraph after Figure 6, 'The Barry et al. method' should read 'The Berry et al. method'.
- [Throughout] The repeatability R is introduced implicitly as R = 1/(1 + δ²/σ²). An explicit equation in the model section would help readers see why the null value R − 1 appears.
- [Bootstrap procedure] The proposed null test requires a chosen value of R, and the conclusion is conditional on that choice. The paper acknowledges this, but it would be clearer to state explicitly that the resulting p-value/decision is not a test of β = 0 without an externally justified repeatability.
- [Model assumptions, Eqs. (1)–(3)] The derivations assume equal error variance at pre and post and independent errors. The authors should add a sentence noting that correlated or heteroscedastic measurement error changes the formulas, and that the empirical recommendations are scoped to the stated model.
Circularity Check
No circularity: the paper's bias derivations are algebraic consequences of an explicitly stated model; self-citations are peripheral, not load-bearing.
full rationale
The paper's derivation chain is self-contained. Starting from the explicit generative model (Eqs. 1–3), it computes the crude slope (Eq. 13), the Berry et al. corrected slope (Eqs. 15–18), and the Blomqvist estimator (Eqs. 20–22) by direct covariance algebra. No parameter is fitted to force the conclusion; the repeatability R enters as an external input from prior empirical studies or from stated hypothetical ranges, not as a fitted quantity. The citations to the authors' own prior work (Santos & Fontanari 2025) appear only as examples of the ρ=0/permutation null that the paper criticizes, and as background; they are not premises of the central algebra. The central claim—that Berry et al.'s corrected slope is biased—is derived from the model rather than assumed. The skeptical objection about the truth of Eq. (18)'s consequence ('only at β=ν²=0') is a mathematical/correctness challenge, not a circularity: it does not show that a prediction reduces to an input by construction. Therefore no circular step is exhibited.
Assumptions & free parameters
free parameters (3)
- repeatability R (or noise ratio δ²/σ²) =
0.69 (blood pressure), 0.479/0.398 (telomere), hypothetical ranges in lizard example
- between-subject treatment variation ν² =
ν = 10 mmHg in simulations
- additive treatment offset α =
-20 mmHg
assumptions (5)
- domain assumption True pre-test values are normal, X1 ~ N(μ, σ²)
- domain assumption True post-test value is linear in true pre-test with additive noise, X2 = X1 + (α + βX1) + ξ
- domain assumption Measurement errors are independent, normal, equal variance δ², independent of X1 and ξ
- standard math Bootstrap resampling of the empirical sample gives valid confidence intervals for the crude slope
- domain assumption External repeatability estimates (e.g., from Wolak 2012, Kärkkäinen 2022) transfer to the target population
Cite this review
Pith. "Pith review of When correcting for regression to the mean is worse than no correction at all." pith.science (2026). https://pith.science/paper/VYWI7GOQ
@misc{pith2026250904718,
author = {Pith},
title = {Pith review of: When correcting for regression to the mean is worse than no correction at all},
year = {2026},
howpublished = {\url{https://pith.science/paper/VYWI7GOQ}},
note = {Machine review of arXiv:2509.04718}
}
read the original abstract
The ubiquitous regression to the mean (RTM) effect complicates statistical inference regarding the relationship between baseline levels of a biological variable and its subsequent change. We demonstrate that common RTM correction methods are problematic: the Berry et al. method, popularized by Kelly & Price in The American Naturalist, is unreliable for hypothesis testing or effect-size estimation, leading to systematic bias and inflated error rates. Conversely, while the Blomqvist method is theoretically unbiased, its high sampling variance limits its practical utility in small-to-moderate datasets. Using a structural linear model, we show that the most robust approach to navigating RTM is not to correct the data, but to evaluate the uncorrected crude slope against a structural null expectation derived from measurement repeatability-the proportion of total variance attributable to true individual differences. We illustrate this approach using empirical data from studies on lizard thermal physiology and bird telomere dynamics. Ultimately, we argue that any conclusion regarding a differential treatment effect is statistically unfounded without a clear understanding of the experiment's repeatability.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Archie, J. P. 1981. Mathematic coupling of data: a common source of error. Annals of Surgery 193:296--303
work page 1981
-
[2]
Berry, D. A., M. L. Eaton, B. P. Ekholm, and T. L. Fox. 1984. Assessing differential drug effect. Biometrics 40:1109--1115
work page 1984
-
[3]
Blomqvist, N. 1977. On the relation between change and initial value. Journal of the American Statistical Association 72:746--749
work page 1977
-
[4]
Castaneda, L. E., G. Calabria, L. A. Betancourt, E. L. Rezende, and M. Santos. 2012. Measurement error in heat tolerance assays. Journal of Thermal Biology 37:432--437
work page 2012
-
[5]
Chiolero, A., G. Paradis, B. Rich, and J. A. Hanley. 2013. Assessing the relationship between the baseline value of a continuous variable and subsequent change over time. Frontiers in Public Health 1:29
work page 2013
-
[6]
Cicho\'n, M., J. Meril\"a, L. Hillstr\"om, and D. Wiggins. 1999. Mass-dependent mass loss in breeding birds: getting the null hypothesis right. Oikos 87:191--194
work page 1999
-
[7]
Chuang-Stein, C. 1993. The regression fallacy. Drug Information Journal 27:1213--1220
work page 1993
-
[8]
Deery, S. A., J. E. Rej, D. Haro, and A. R. Gunderson. 2021. Heat hardening in a pair of Anolis lizards: constraints, dynamics and ecological consequences. Journal of Experimental Biology 224:jeb240994
work page 2021
Show all 33 references
-
[9]
Efron, B., and R. J. Tibshirani. 1993. An Introduction to the Bootstrap. Chapman & Hall, New York
1993
-
[10]
Ferrer, E., and J. J. McArdle. 2010. Longitudinal modeling of developmental changes in psychological research. Current Directions in Psychological Science, 19:149--154
2010
-
[11]
Wagenmakers, and T
Forstmeier W., E.-J. Wagenmakers, and T. H. Parker. 2017. Detecting and avoiding likely false-positive findings - a practical guide. Biological Reviews 92:1941--1968
2017
-
[12]
Galton, F. 1886. Regression towards mediocrity in hereditary stature. The Journal of the Anthropological Institute of Great Britain and Ireland 15:246--263
-
[13]
Gardner, M. J. , and J. A. Heady. 1973. Some effects of within-person variability in epidemiological studies. Journal of Chronic Diseases 26:781--795
1973
-
[14]
Geenen, R., and F. J. R. van de Vijver. 1993. A simple test of the law of initial values. Psychophysiology 30:525--530
1993
-
[15]
Gunderson, A. R. 2023. Trade-offs between baseline thermal tolerance and thermal tolerance plasticity are much less common than it appears. Global Change Biology, 29:3519--3524
2023
-
[16]
Hanushek, E. A., L. Kinne, F. Witthöft, and L. Woessmann. 2025. Age and cognitive skills: Use it or lose it. Sci. Adv. 11:eads1560
2025
-
[17]
Hayes, R. J. 1988. Methods for assessing whether change depends on initial value. Statistics in Medicine 7:915--927
1988
-
[18]
A., and K
Jackson, D. A., and K. M. Somers. 1991. The spectre of 'spurious' correlations. Oecologia 86:147--151
1991
-
[19]
Briga, T
Kärkkäinen, T., M. Briga, T. Laaksonen, and A. Stier. 2022. Within-individual repeatability in telomere length: A meta-analysis in nonmammalian vertebrates. Molecular Ecology 31:6339--6359
2022
-
[20]
Kelly, C., and T. D. Price. 2005. Correcting for regression to the mean in behavior and ecology. American Naturalist 166:700--707
2005
-
[21]
Kronmal, R. A. 1993. Spurious correlation and the fallacy of the ratio standard revisited. Journal of the Royal Statistical Society A 156:379--392
1993
-
[22]
Diekmann
Mazalla, L., and M. Diekmann. 2022. Regression to the mean in vegetation science. Journal of Vegetation Science 33:e13117
2022
-
[23]
Nesselroade, J. R., S. M. Stigler, and P. B. Baltes. 1980. Regression toward the mean and the study of change. Psychological Bulletin 88:622--637
1980
-
[24]
Pearson, K. 1897. Mathematical contributions to the theory of evolution. -On a form of spurious correlation which may arise when indices are used in the measurement of organs. Proceedings of the Royal Society of London 60:489--497
-
[25]
Pitman, E. J. G. 1939. A note on normal correlation. Biometrika 31:9--12
1939
-
[26]
Brandt, and M
Rogosa, D., D. Brandt, and M. Zimowski. 1982. A growth curve approach to the measurement of change. Psychological Bulletin, 92:726--748
1982
-
[27]
Santos, M., J. F. Fontanari. 2025. On testing the tolerance-plasticity trade-off hypothesis as the change of thermal tolerance across two environments. Journal of Thermal Biology, 132:104248
2025
-
[28]
Slessarev E. W., A. Mayer, C. Kelly, K. Georgiou, J. Pett-Ridge, and E. E. Nuccio. 2023. Initial soil organic carbon stocks govern changes in soil carbon: reality or artifact? Global Change Biology 29:1239--1247
2023
-
[29]
Ingre, G
Sorjonen, K., M. Ingre, G. Nilsonne, and B. Melin. 2023. Dangers of including outcome at baseline as a covariate in latent change score models: results from simulations and empirical re-analyses. Heliyon 9:e15746
2023
-
[30]
M., Gustafsson, L
Sudyka, J., Arct, A., Drobniak, S. M., Gustafsson, L. , and Cicho\'n, M. 2019. Birds with high lifetime reproductive success experience increased telomere loss. Biol. Lett. 15:20180637
2019
-
[31]
Wasserman, L. 2004. All of Statistics: A Concise Course in Statistical Inference. Springer, New York
2004
-
[32]
Wilder, J. 1967. Stimulus and response: the law of initial value. John Wright & Sons LTD, Bristol
1967
-
[33]
Wolak, M. E., D. J. Fairbairn, and Y. R. Paulsen. 2012. Guidelines for estimating repeatability. Methods in Ecology and Evolution 3:129–137
2012
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.