REVIEW 3 major objections 3 minor 16 references
Internal replication as a tool for evaluating reproducibility in preclinical experiments
T0 review · 3 major / 3 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper shows that many preclinical experiments already contain the replication needed to assess reproducibility, and that testing the treatment-by-batch interaction reveals whether an effect is stable across batches.
desk verdict A genuinely useful taxonomy for internal replication and a clean illustrative ANOVA, but the paper overreaches when it reads a significant batch-by-treatment interaction as 'not reproducible'. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the generalised randomised block design, in which batches such as sites, days, litters, or runs serve as blocks and each treatment appears in every block with genuine replicates. The treatment-by-block interaction term $\tau\beta$ carries the argument: testing $H_0: (\tau\beta)_{ij}=0$ for all $i,j$ directly assesses internal replication, while a second F-test can use the interaction mean square as the error term for the treatment effect, effectively treating the genuine replicates as pseudoreplicates. The design is specified as $y_{ijk} = \mu + \tau_i + \beta_j + (\tau\beta)_{ij} + \varepsilon_{ijk}$ with $\varepsilon_{ijk} \sim \mathrm{Normal}(0,\sigma)$, and the interaction test is what turns internal replication into a reproducibility assessment.
What would settle it
Compare interaction p-values from many multi-batch experiments with their later external replication outcomes; if significant treatment-by-batch interactions are not followed by replication failure more often than non-significant ones, the central claim would be refuted.
Extended reading notes
Core claim
The paper's central claim is that the stability of an experimental effect can be estimated from a single study whenever the experimental units are grouped into two or more batches, and that the right statistic is the treatment-by-batch interaction. Each batch is treated as a block in a generalised randomised block design with genuine replication, and a significant interaction means the effect varies across batches more than sampling variability alone would predict, which the paper interprets as non-reproducibility. In the illustrative reanalysis of Harrison and colleagues' mouse lifespan data, the site-by-treatment interaction was significant (p = 0.024), while the average 17aE2 effect was significant only when tested against the residual error (p = 0.002) and not when tested against the site-to-site variation (p = 0.254). The paper argues that reporting only the main effect in such cases misleads readers about how much the result can be trusted.
Load-bearing premise
The claim rests on treating a significant treatment-by-batch interaction as evidence that the effect is not reproducible, rather than as benign variation in scale or conditions that would still allow the effect to generalise.
Editorial extensions
If this is right
- A significant treatment-by-batch interaction means the main effect is not stable across batches, so reporting only the average effect would be misleading.
- Researchers can report an internal-reproducibility statistic from existing data at no experimental cost by modelling batches as blocks in a generalised randomised block design.
- The treatment effect can be tested against the batch-interaction variance; a significant result is strong evidence, though the test is usually underpowered because the effective sample size is the number of batches.
- Designing experiments with treatments crossed with batches, balanced samples, and as much batch independence as possible makes the interaction test informative.
- Internal-reproducibility metrics should become a standard reporting item in preclinical studies.
Reading between the lines
- By extension, the same interaction test could be run on published multi-batch datasets, such as multi-site or multi-litter studies, to flag effects whose stability is questionable without running new experiments.
- A natural next step would be to calibrate the interaction p-value against external replication outcomes; if the two do not correlate, the link between internal instability and replication failure would be weakened.
- The approach implicitly assumes batch differences that produce a significant interaction matter for generalisation, so a reader applying it to a new field should decide whether the batch factor is one over which reproducibility is actually desired.
- Standardising a minimal set of design features, namely crossing, balance, and independence of batches, could turn the interaction F-test into a routinely reported reproducibility statistic.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces the concept of 'internal replication' for preclinical experiments, defined as repetition across batches, runs, days, litters, or sites already present within a single study, and argues that such data can be used to assess reproducibility without additional experiments. Six types of internal replication are classified by independence and timing. The authors propose analyzing such designs as generalized randomized block designs (GRBDs), with the treatment-by-batch interaction test used to assess whether the treatment effect is stable across batches. The method is illustrated on a mouse lifespan dataset from a three-site study of 17α-Estradiol, where the site-by-treatment interaction is significant (p = 0.024) and the authors conclude the treatment effect was not reproducible across sites.
Significance. If the interpretive framework is accepted, the paper offers a practical, low-cost way to extract reproducibility information from existing experimental designs, which is valuable given the expense of full replication studies. The six-type taxonomy is clear and useful for practitioners. The illustrative reanalysis is transparent and reproducible from the reported ANOVA table, and the paper explicitly acknowledges the low power of the stringent treatment test. However, the central statistical interpretation—that a significant treatment-by-batch interaction equates to non-reproducibility—needs refinement, because it conflates observed heterogeneity among a fixed set of batches with generalizable instability across future replications.
major comments (3)
- [Section 2.2, Eq. (1)-(2) and Results] The claim that a significant treatment-by-batch interaction indicates the experiment is 'not reproducible' is a load-bearing interpretive leap. In a GRBD with fixed blocks, the interaction estimates systematic differences in treatment effects among the observed batches; these differences could be stable, reproducible properties of the system (e.g., site-specific genetics, diet, or husbandry) rather than stochastic instability. The manuscript's own Introduction limits the method to estimating 'the stability of results within an experiment,' which is weaker than 'not reproducible across sites' as stated in the Results. Please reframe the conclusion as evidence of heterogeneity across observed sites, explicitly acknowledge that context-dependence is a possible explanation, and add a discussion of what additional evidence (e.g., external replication or measurement of moderators) would be needed to distinguish stochastic instability from reproducible context-dependence.
- [Section 2.2, fixed vs. random block treatment] The paper argues that blocks should be treated as fixed effects because no inference about hypothetical future blocks is intended, yet the abstract and Discussion frame the method as estimating 'reproducibility.' With fixed blocks, the analysis only compares the batches actually observed and provides no estimate of between-batch variance that could be used to predict the behavior of a new batch or site. To support the reproducibility claim, the manuscript should either (a) limit the stated scope to consistency among the observed batches, or (b) supplement the fixed-effects analysis with a random-effects or variance-component analysis (e.g., mixed-effects model with a random site-by-treatment interaction) that quantifies the between-site variance component, along with a discussion of the small number of sites and the accompanying estimation uncertainty.
- [Results, Table 1 and Figure 2B] The binary conclusion 'not reproducible across sites' rests on a single p-value (p = 0.024) from three sites, and the visual pattern in Figure 2B suggests the interaction is driven primarily by one site (UT) with a large positive effect and another (TJL) with near-zero effect. The paper does not report effect-size measures for the interaction (e.g., variance component, range of site-specific effects, or a leave-one-site-out analysis). Because the conclusion is central to the paper's message, please add a sensitivity analysis (e.g., re-fitting the model after omitting each site) and report site-specific effect estimates with confidence intervals, so readers can judge how stable the interaction conclusion is.
minor comments (3)
- [Section 2.2] The normality assumption for the lifespan outcome is asserted without diagnostics. Given that lifespan is a bounded, positive outcome, please provide a Q-Q plot or a brief justification (e.g., citing central-limit-theorem robustness for the F-test), or run a sensitivity check using a survival analysis or nonparametric approach.
- [Section 3] The text refers to 'Figure 3A' and 'Figure 3B', but the corresponding figure is numbered 'Figure 2' in the caption. Please correct the cross-references.
- [Section 1.1] In Figure 1, the distinction between 'fully independent' and 'partially independent' for the staggered case could be made more explicit in the caption; currently the visual difference relies entirely on the text about shared materials or researchers.
Circularity Check
No significant circularity: the analysis is a standard GRBD ANOVA on published data; internal-reproducibility claims are operational definitions, not fitted predictions.
full rationale
The paper's central proposal is to use existing batch/site structure in designed experiments as 'internal replication' and to test the treatment-by-batch interaction in a generalised randomised block design. The illustrative analysis is a standard two-way ANOVA F-test on the published Harrison et al. (2009) lifespan data; no model parameters are fitted to a subset and then 'predicted' on a closely related quantity, and no distributional or uniqueness theorem is imported from the author's own prior work. The quantities reported—F statistics, p-values, degrees of freedom—are direct test statistics, not derived from fitted values. The only potentially definitional move is equating a significant treatment-by-batch interaction with 'not reproducible'; but this is presented explicitly as an operational definition of internal replication in Section 2.2 rather than as a result derived from that definition, so it does not constitute a circular derivation. Self-citations (Lazic and Essioux 2013; Lazic 2016; Lazic et al. 2018) appear only as design recommendations and a standard pseudoreplication point, and they do not carry the empirical conclusion. The conclusion that the 17aE2 effect differed by site is an interpretation of the interaction test in the data, not a fitted input renamed as a prediction. Therefore no circular step is present.
Assumptions & free parameters
assumptions (4)
- domain assumption The outcome (lifespan) is well-approximated by a normal distribution, justifying ANOVA.
- domain assumption A generalized randomized block design with genuine replication is an appropriate model for all six types of internal replication.
- domain assumption Blocks should be treated as fixed effects because inference is about the specific blocks in the experiment.
- domain assumption A significant treatment-by-block interaction means the experiment is not reproducible.
Cite this review
Pith. "Pith review of Internal replication as a tool for evaluating reproducibility in preclinical experiments." pith.science (2026). https://pith.science/paper/DKHU5DH5
@misc{pith2026250603468,
author = {Pith},
title = {Pith review of: Internal replication as a tool for evaluating reproducibility in preclinical experiments},
year = {2026},
howpublished = {\url{https://pith.science/paper/DKHU5DH5}},
note = {Machine review of arXiv:2506.03468}
}
read the original abstract
Reproducibility is central to the credibility of scientific findings, yet complete replication studies are costly and infrequent. However, many biological experiments contain internal replication, which is defined as repetition across batches, runs, days, litters, or sites that can be used to estimate reproducibility without requiring additional experiments. This internal replication is analogous to internal validation in prediction or machine learning models, but is often treated as a nuisance and removed by normalisation, missing an opportunity to assess the stability of results. Here, six types of internal replication are defined based on independence and timing. Using mice data from an experiment conducted at three independent sites, we demonstrate how to quantify and test for internal reproducibility. This approach provides a framework for quantifying reproducibility from existing data and reporting more robust statistical inferences in preclinical research.
Figures
Reference graph
Works this paper leans on
-
[8]
doi: 10.1038/nature11556. S. E. Lazic. Experimental Design for Laboratory Biologists: Maximising Information and Improving Reproducibility. Cambridge University Press, Cambridge, UK,
-
[13]
doi: 10.1038/nmeth.1312. S. H. Richter, J. P. Garner, C. Auer, J. Kunert, and H. Würbel. Systematic variation improves reproducibility of animal experiments. Nat Methods, 7(3):167–168,
-
[16]
doi: 10.1038/s41598-020-73503-4
ISSN 2045-2322. doi: 10.1038/s41598-020-73503-4. 8
-
[1936]
ISSN 1466-6162. doi: 10.2307/2983667. D. E. Harrison, R. Strong, Z. D. Sharp, J. F. Nelson, C. M. Astle, K. Flurkey, N. L. Nadon, J. E. Wilkinson, K. Frenkel, C. S. Carter, M. Pahor, M. A. Javors, E. Fernandez, and R. A. Miller. Rapamycin fed late in life extends lifespan in genetically heterogeneous mice. Nature, 460(7253):392–395, July
-
[2007]
doi: 10.1038/sj.bjp.0707372. Open Science Collaboration. Psychology. estimating the reproducibility of psychological science. Science, 349(6251): aac4716,
-
[2008]
doi: 10.1080/17482960701856300. E. W. Steyerberg. Clinical Prediction Models: A Practical Approach to Development, Validation, and Updating . Springer, Cham, Switzerland, 2nd edition,
-
[2009]
ISSN 1476-4687. doi: 10.1038/nature08221. K. Hinkelmann and O. Kempthorne. Design and Analysis of Experiments, Volume 1: Introduction to Experimental Design. Wiley, Hoboken, NJ, 2nd edition,
-
[2010]
doi: 10.1038/nmeth0310-167. S. Scott, J. E. Kranz, J. Cole, J. M. Lincecum, K. Thompson, N. Kelly, A. Bostrom, J. Theodoss, B. M. Al-Nakhala, F. G. Vieira, J. Ramasubbu, and J. A. Heywood. Design, power, and interpretation of studies in the standard murine model of ALS. Amyotroph Lateral Scler, 9(1):4–15,
Show all 16 references
-
[2011]
doi: 10.1038/nrd3439-c1. S. H. Richter, J. P. Garner, and H. Würbel. Environmental standardization: cure or cause of poor reproducibility in animal experiments? Nat Methods, 6(4):257–261,
-
[2012]
doi: 10.1038/483531a. C. Bodden, V . T. von Kortzfleisch, F. Karwinkel, S. Kaiser, N. Sachser, and S. H. Richter. Heterogenising study samples across testing time improves reproducibility of behavioural data. Scientific Reports, 9(1),
-
[2013]
doi: 10.1186/1471-2202-14-37. S. E. Lazic, C. J. Clarke-Williams, and M. R. Munafo. What exactly is ’N’ in cell culture and animal experiments? PLoS Biol., 16(4):e2005282,
-
[2015]
doi: 10.1126/science.aac4716. N. Percie du Sert, A. Ahluwalia, S. Alam, M. T. Avey, M. Baker, W. J. Browne, A. Clark, I. C. Cuthill, U. Dirnagl, M. Emerson, P. Garner, S. T. Holgate, D. W. Howells, V . Hurst, N. A. Karp, S. E. Lazic, K. Lidster, C. J. MacCallum, M. Macleod, E....
-
[2018]
doi: 10.1038/s41562-018-0399-z. G. Casella. Statistical Design. Springer, New York, NY ,
-
[2019]
doi: 10.1038/s41598-019-44705-2. C. F. Camerer, A. Dreber, F. Holzmeister, T.-H. Ho, J. Huber, M. Johannesson, M. Kirchler, G. Nave, B. A. Nosek, T. Pfeiffer, A. Altmejd, N. Buttrick, T. Chan, Y . Chen, E. Forsell, A. Gampa, E. Heikensten, L. Hummer, T. Imai, S. Isaksson, D. M...
-
[2020]
doi: 10.1038/s41598-020-62509-7. S. C. Landis, S. G. Amara, K. Asadullah, C. P. Austin, R. Blumenstein, E. W. Bradley, R. G. Crystal, R. B. Darnell, R. J. Ferrante, H. Fillit, R. Finkelstein, M. Fisher, H. E. Gendelman, R. M. Golub, J. L. Goudreau, R. A. Gross, A. K. Gubitz, S...
-
[2021]
doi: 10.7554/elife.71601. W. S. Gosset. Co-operation in large-scale experiments. Supplement to the Journal of the Royal Statistical Society, 3(2): 115,
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.