{"id":"ce0e1ad6-27e4-47d3-9622-9b684ddcbb7f","arxiv_id":"2509.04718","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Two common regression-to-the-mean corrections are biased or too variable; the paper recommends testing the uncorrected slope against a null based on trait repeatability.","lead":"A statistics paper shows that two common methods for correcting regression to the mean are unreliable, and that the raw crude slope, tested against a repeatability-based null, can be the safer choice. The result matters because many ecology and evolution papers use the Berry correction to infer trade-offs between baseline traits and plasticity.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Berry-correction unbiasedness claim is false: βB=β can occur for β≠0, ν²>0, contradicting the paper's 'only at β=ν²=0' assertion.","rationale":"The reader's verdict identifies display errors in Eqs. (23) and (25) and valid concerns about finite-sample variance, but it does not catch a deeper algebraic issue: the paper's own equations admit a continuum of parameter values where the Berry correction is exactly unbiased. This directly contradicts the strongest claim as summarized: 'the method only yields the true slope β in the highly restrictive case where β = ν² = 0.' The counterexample is concrete and easy to verify, so the paper must be corrected before its theoretical characterization of the Berry method is accepted. However, the paper's practical conclusion—that the Berry correction can produce false positives and negatives and is unreliable for hypothesis testing—still appears to hold for generic parameter values, especially because ν² is unknown and unidentifiable from two time points. Thus the appropriate verdict remains CONDITIONAL rather than REJECT: the authors need to fix the unbiasedness condition, correct Eqs. (23) and (25), and rephrase the scope of the bias claims. The finite-sample variance claims for Blomqvist also need broader simulation support, as the reader noted.","tokens_in":14136,"tokens_out":14360,"duration_ms":122256,"concrete_test":"Symbolically solve βB=β using Eqs. (17) and (18) for ν² in terms of σ², δ², and β. Then instantiate σ²=1, δ²=1, β=0.1, ν²≈0.778 and simulate N=10^5 independent samples from the model (Eqs. 1-3). Compute the sample Berry slope from Eq. (16) and check whether the regression slope of dB on x1 equals 0.1 within sampling error. If it does, the claim that βB=β only when β=ν²=0 is refuted.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central theoretical claim is that the Berry et al. corrected slope βB equals the true slope β only when β=ν²=0. This is mathematically false within the paper's own model. Setting βB=β in Eq. (18) yields ρ = (σ² - βδ²)/(σ²+δ²). Substituting the population ρ from Eq. (17) and solving for ν² gives a continuum of non-null parameter values where βB=β exactly. For example, with σ²=1, δ²=1, β=0.1, the required ν² is (1.1)²·2/(0.9)² - 1.21 - 1 ≈ 0.778, giving ρ=0.45 and βB=0.1=β. Thus the 'highly restrictive case' is not a point but a one-dimensional manifold. This also affects Eq. (25): under β=0 and ν²>0, the correct βB is σ²/σ1²[1 - 1/√(1+ν²/σ1²)] > 0, not the negative expression displayed. The paper's described false-positive direction is therefore wrong, though false positives can still occur. The practical recommendation may survive, but the central theoretical characterization of the Berry correction's bias is incorrect and must be revised.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a structural two-level model in which true pre- and post-test values follow X2 = X1 + (α + βX1) + ξ and measurements add independent normal noise with variance δ² at both time points. Within this model it derives population expressions for the crude slope βc, the Berry et al. corrected slope βB, and the Blomqvist estimator βe, and then uses simulations and two empirical case studies to argue that (i) the Berry correction is systematically biased and unreliable for hypothesis testing, (ii) the Blomqvist estimator, though unbiased, has high sampling variance and is often worse than the crude slope, and (iii) the most practical approach is to test the crude slope against a repeatability-based null expectation. The paper claims that βB equals the true slope only in the highly restrictive case β = ν² = 0, and uses Eq. (23) and Eq. (25) to support the bias and false-positive narrative.","tokens_in":14459,"tokens_out":13382,"duration_ms":115509,"significance":"If the derivations were correct, the paper would be a useful cautionary contribution: it connects the RTM bias directly to repeatability, gives a simple bootstrap recipe for the null β = 0, and applies it to published datasets where earlier Kelly-Price-style corrections were used. The algebraic structure is transparent, no parameters are fitted to force the conclusions, and the empirical examples are concrete. However, the central mathematical characterization of the Berry correction contains load-bearing errors: the 'only at β = ν² = 0' claim is false, and the limit formulas in Eqs. (23) and (25) are misstated. These issues need to be corrected and the corresponding narrative revised before the paper's conclusions can be accepted.","major_comments":[{"comment":"The statement that βB = β only when β = ν² = 0 is false. Setting βB = β in Eq. (18) gives ρ = (σ² − βδ²)/(σ² + δ²). Combining this with Eq. (17) yields a one-dimensional family of non-null parameters, e.g. σ² = δ² = 1, β = 0.1, ν² ≈ 0.778 gives βB = β = 0.1. The paper's own δ² = 0 limit also admits βB = β for all β > −1 when ν² = 0. The unbiasedness condition is therefore a continuum, not a point. This does not necessarily rescue the Berry correction as a general method, but the current claim and its downstream wording are incorrect and must be revised.","section":"Correcting for RTM using the Berry et al. method, Eq. (18)"},{"comment":"Under the null β = 0, the correct expression is βB = σ²/σ₁² [1 − 1/sqrt(1 + ν²/σ₁²)], which is positive for ν² > 0. Equation (25) as printed has δ² in place of σ² and the opposite sign, and the accompanying statement 'This expression is always negative for ν² > 0' is wrong. The qualitative warning that the Berry slope is nonzero under the null survives, but the sign and the false-positive direction are misdescribed. The text should be corrected and the implications re-stated.","section":"Testing for a Differential Treatment Effect, Eq. (25)"},{"comment":"For δ² = 0, the correct limit of Eq. (18) is βB = 1 + β − sgn(1 + β)/sqrt(1 + ν²/[(1 + β)²σ²]). The displayed Eq. (23) lacks the reciprocal and gives β + 1 − sgn(1 + β) * sqrt(1 + ν²/[(1 + β)²σ²]), which is not the limit. For example, for small ν² and β > −1 the true limiting slope exceeds β by O(ν²), whereas the printed expression behaves differently. This is not a typo in an ancillary remark; it is the equation used to describe the method's behavior in the no-measurement-error regime.","section":"Analysis of the population regression slopes, Eq. (23)"},{"comment":"The general claim that the Blomqvist estimator is practically inferior to the crude slope rests on a single simulation configuration: N = 100, the systolic blood pressure parameters, δ²/σ² = 0.45, and one value of ν/σ. The variance comparison and the crossing probabilities in Figure 3 may depend on these choices. To support a general recommendation against Blomqvist, the authors should either derive an analytic expression for the sampling variance or provide a systematic exploration over N, repeatability, and ν²/σ². At minimum, the abstract's blanket statement should be tempered.","section":"Sample Size Effects on Regression Slopes, Figures 2–3"}],"minor_comments":[{"comment":"In the paragraph after Figure 6, 'The Barry et al. method' should read 'The Berry et al. method'.","section":"Telomere case study"},{"comment":"The repeatability R is introduced implicitly as R = 1/(1 + δ²/σ²). An explicit equation in the model section would help readers see why the null value R − 1 appears.","section":"Throughout"},{"comment":"The proposed null test requires a chosen value of R, and the conclusion is conditional on that choice. The paper acknowledges this, but it would be clearer to state explicitly that the resulting p-value/decision is not a test of β = 0 without an externally justified repeatability.","section":"Bootstrap procedure"},{"comment":"The derivations assume equal error variance at pre and post and independent errors. The authors should add a sentence noting that correlated or heteroscedastic measurement error changes the formulas, and that the empirical recommendations are scoped to the stated model.","section":"Model assumptions, Eqs. (1)–(3)"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to be of interest to the statistical-methods audience of this journal, and the applied examples are relevant. The main barrier is not scope but correctness of the core equations. I would ask the authors to re-derive Eqs. (23) and (25), revise the unbiasedness characterization of βB, and either broaden or soften the finite-sample claims about Blomqvist. The paper cites its own prior work in a few places, but the central algebra is independent, so I do not see a novelty-disclosure concern."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The practical recommendation here is better than the theory that surrounds it. Testing the crude slope against a repeatability-based null (R−1) is a coherent, useful idea, and the two empirical re-analyses show how published conclusions can flip. The identity βB = βc + (1−ρ) is clean and worth having. The paper also correctly identifies the pitfalls of using ρ=0 or variance equality as null hypotheses, and it honestly states that δ² is not identifiable from two time points. That is real value.\n\nBut the central theoretical claim about the Berry correction is not correct. The statement that βB equals β only when β=ν²=0 is false within the paper's own model. Setting βB=β in Eq. (18) gives a continuum of solutions: for σ²=δ²=1 and β=0.1, ν²≈0.78 makes βB=β exactly. So the paper overstates the restrictiveness of the unbiasedness condition. Eq. (25) also has a sign error: under β=0 the population βB is positive for ν²>0, not negative. This reverses the claimed direction of the false-positive bias. Eq. (23) misstates the δ²=0 limit; the square root should be in the denominator, and the formula is only accidentally correct at ν²=0. These are not typos in peripheral text; they are displayed equations used to quantify the Berry bias.\n\nWhat survives? The qualitative message that the Berry method is biased, so conclusions drawn from it are not automatically trustworthy, still holds—just not for the reasons the paper gives. The repeatability-based null and the bootstrap recipe are practical and independent of those errors. The finite-sample variance comparison for Blomqvist, however, is limited to one parameter set (N=100, δ²/σ²≈0.45), so the claim that it is worse in practice is not yet generalized. The population formulas mostly restate Hayes (1988), which the paper acknowledges, but the new identity and the recipe are the contribution.\n\nWho is this for? Ecologists and evolutionary biologists who use Kelly–Price style corrections, and applied statisticians who referee those papers. It deserves a serious referee, but not in its current form. I would send it to peer review with a clear request to fix the algebra and re-examine every conclusion that depends on the sign or magnitude of βB. Once the theory is corrected, the practical message could be an important one.","headline":"The repeatability-null recipe is worth stealing, but the paper's algebra on the Berry correction is wrong in ways that undermine its central theoretical claims.","tokens_in":14981,"tokens_out":5815,"would_cite":false,"duration_ms":48451,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F03","62J05","62F40"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that the most widely used correction for regression to the mean—the Berry et al. slope adjustment—is itself biased and can create false positives and false negatives, and argues that the uncorrected crude slope tested again","keywords":["regression to the mean","measurement error","hypothesis testing","bootstrap","repeatability","differential treatment effect","change scores","structural linear model"],"falsifier":"Simulate two-time-point data with known true slope β, known between-subject treatment variance ν²>0, and known measurement-error variance δ², then estimate the Berry et al. slope βB on many samples; if the average βB equals β for any parameter combination with ν²>0, equation (18)—the claim that the correction is biased except at β=ν²=0—is false. Equivalently, a real dataset with an external gold-standard measure of the trait could be used to compare corrected vs. true slopes across repeatability values.","tokens_in":1915,"feed_emoji":"📉","tokens_out":2154,"duration_ms":85519,"temperature":0.7,"pith_summary":"This paper argues that the most popular correction for regression to the mean (RTM), the Berry et al. slope adjustment popularized in ecology by Kelly and Price, does more harm than good: its corrected slope is systematically biased and can produce both false positives and false negatives, because it only recovers the true slope in an unrealistic special case. The authors build a simple structural model in which the true post-test value is a linear function of the pre-test value plus a treatment effect and noise, and measured values add independent error. Within that model they show the uncorrected crude slope has a clean bias expressible through repeatability, the proportion of total variance due to true individual differences, and they propose testing the crude slope against that structural null expectation using bootstrap. They reanalyze published lizard and bird datasets to show that conclusions change depending on the assumed repeatability. The punchline is that without knowing, or at least bounding, an experiment's repeatability, claims about differential treatment effects are statistically unfounded.","feed_headline":"The standard fix for regression to the mean is itself biased","feed_subtitle":"Popular regression-to-the-mean corrections are biased; the unadjusted slope plus repeatability is the safer test.","key_machinery":"The structural change model X2 = X1 + (α+βX1) + ξ, with independent measurement noise ϵ₁,ϵ₂ ~ N(0,δ²), is the engine of the paper. It yields the exact identity δ²/σ²₁ = 1−R, where R is repeatability, and the crude-slope formula βc = β − (1+β)δ²/(σ²+δ²). This identity does the work: it turns the unobservable measurement-error variance into a quantity researchers can estimate or bound, and converts the null hypothesis β=0 into the testable claim βc = −δ²/σ²₁ = R−1.","core_discovery":"On the paper's own terms, the central discovery is a closed-form comparison of three estimates of the slope of true change on true initial value: the crude slope βc, the Berry et al. corrected slope βB, and the Blomqvist slope βe. For the crude slope, βc = β − (1+β)δ²/(σ²+δ²), so the RTM bias is entirely driven by within-subject variance δ² and is independent of the stochastic variance ν² of the treatment effect. For the Berry et al. correction, βB = −ρ + (1+β)σ²/(σ²+δ²) = βc + (1−ρ), which is always greater than or equal to the crude slope and equals the true slope only when β = ν² = 0. The Blomqvist correction is algebraically unbiased but has roughly twice the sampling variance of the cru","pith_inferences":["If the structural model is approximately right, the same logic extends to designs with more than two time points: repeatability estimated from repeated measures could replace the qualitative bound with a direct null test, a step the paper leaves implicit.","The paper's framework suggests a simple diagnostic for any published RTM-corrected result: check whether a plausible repeatability range places the reported corrected slope's null value inside the bootstrap interval; if not, the inference is fragile.","By reframing RTM as a problem of knowing R, the paper connects to the measurement-error and latent-variable literature, where the same identifiability problem (σ² vs δ² from two observations) is well known; a Bayesian or mixed-model approach with informative priors on R might be a natural next step.","A quantitative prediction follows: meta-analyses of traits with low repeatability (R<0.5) should show apparent negative slopes between initial value and change even when no differential treatment effect exists, and published positive results should cluster in low-repeatability studies."],"forward_implications":["If the paper is right, every conclusion drawn from the Kelly–Price/Berry et al. adjustment about differential treatment effects is unsupported, because the corrected slope is biased whenever there is any between-subject variation in treatment effect (ν²>0).","The Blomqvist method cannot be recommended as a routine fix: although unbiased in principle, its high sampling variance makes it less reliable than the biased crude slope in sample sizes typical of ecology and physiology.","A practical testing recipe emerges: compute the crude slope, bootstrap its confidence interval, and reject a differential treatment effect only if the interval excludes R−1 for the experiment's expected repeatability.","Re-analysis of existing datasets (lizards, birds) may overturn published conclusions: evidence for or against trade-offs and biomarkers depends on repeatability values that are often not reported.","Reporting repeatability becomes a prerequisite for valid inference about change; readers should demand it."],"supporting_citations":[{"why":"Supplies the structural linear model (Eq. 1) and the derivation of the crude slope βc that the paper's comparisons build on.","marker":"Hayes (1988)"},{"why":"The correction method under critique; its adjusted-change definition and bivariate-normal framing are the target of the bias analysis.","marker":"Berry et al. (1984)"},{"why":"Popularized the Berry method in ecology and added the mean-adjustment term dB; the paper re-examines conclusions drawn from it.","marker":"Kelly & Price (2005)"},{"why":"The theoretically unbiased correction whose finite-sample variance the paper shows to be too large for practical use.","marker":"Blomqvist (1977)"},{"why":"Provides the systolic blood pressure parameters (µ, σ, δ) used in the simulation studies.","marker":"Gardner & Heady (1973)"},{"why":"Lizard heat-tolerance dataset reanalyzed to show that the no-differential-effect conclusion depends on repeatability.","marker":"Deery et al. (2021)"},{"why":"Telomere dataset reanalyzed; originally analyzed with the Berry et al. adjustment, here used to compare crude, Berry, and Blomqvist slopes.","marker":"Sudyka et al. (2019)"},{"why":"Supplies repeatability estimates for telomere length needed to implement the Blomqvist adjustment.","marker":"Kärkkäinen et al. (2022)"},{"why":"Reports typical median repeatability values for physiological and behavioral traits, connecting the method to practice.","marker":"Wolak et al. (2012)"},{"why":"Provides the bootstrap procedure used to build confidence intervals for the crude and corrected slopes.","marker":"Efron & Tibshirani (1993)"}],"fun_headline_variants":["RTM fixes can be worse than no fix","Why correcting regression to mean backfires","Skip the correction: raw slope plus repeatability","Standard RTM corrections biased, raw slope better","Don't correct for regression to the mean"],"cache_read_input_tokens":16640,"weakest_assumption_plain":"The load-bearing premise is that the true post-test value is a linear function of the true pre-test value plus treatment effect and independent noise, and that measured pre- and post-values add independent, equal-variance normal error; if change is nonlinear in the initial value, or the errors are correlated or heteroscedastic across time, the paper's formulas for the crude, corrected, and null slopes no longer hold.","fun_headline_variants_meta":{"raw":{"variants":["RTM fixes can be worse than no fix","Why correcting regression to mean backfires","Skip the correction: raw slope plus repeatability","Standard RTM corrections biased, raw slope better","Don't correct for regression to the mean"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000153,"raw_usage":{"total_tokens":1048,"prompt_tokens":755,"completion_tokens":293,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":499,"completion_tokens_details":{"reasoning_tokens":225}},"tokens_in":499,"tokens_out":293,"duration_ms":3319,"temperature":1.0,"reasoning_tokens":225,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T05:56:13.183107+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate two-time-point data with known true slope β, known between-subject treatment variance ν²>0, and known measurement-error variance δ², then estimate the Berry et al. slope βB on many samples; if the average βB equals β for any parameter combination with ν²>0, equation (18)—the claim that the correction is biased except at β=ν²=0—is false. Equivalently, a real dataset with an external gold-standard measure of the trait could be used to compare corrected vs. true slopes across repeatability values.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the structural linear model (Eq. 1) and the derivation of the crude slope βc that the paper's comparisons build on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The correction method under critique; its adjusted-change definition and bivariate-normal framing are the target of the bias analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Popularized the Berry method in ecology and added the mean-adjustment term dB; the paper re-examines conclusions drawn from it."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The theoretically unbiased correction whose finite-sample variance the paper shows to be too large for practical use."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the systolic blood pressure parameters (µ, σ, δ) used in the simulation studies."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Lizard heat-tolerance dataset reanalyzed to show that the no-differential-effect conclusion depends on repeatability."},{"cited_title":"M., Gustafsson, L","cited_arxiv_id":null,"evidence_quote":"Telomere dataset reanalyzed; originally analyzed with the Berry et al. adjustment, here used to compare crude, Berry, and Blomqvist slopes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Reports typical median repeatability values for physiological and behavioral traits, connecting the method to practice."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the bootstrap procedure used to build confidence intervals for the crude and corrected slopes."}],"review_version":1}