{"id":"500d3a20-0e79-4e97-83cc-8bef489bfc6b","arxiv_id":"2506.03468","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Internal replication, the repetition of experimental units across batches or sites, can be used to test reproducibility within a single study via the treatment-by-batch interaction.","lead":"This paper proposes a framework for assessing how reproducible a preclinical result is by looking at variation across batches or sites inside a single experiment. It shows on a three-site mouse lifespan study that the treatment effect was inconsistent across sites.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Treating a significant treatment-by-batch interaction as 'not reproducible' conflates stochastic instability with reproducible context-dependence, and the fixed-block analysis does not estimate cross-batch reproducibility.","rationale":"The reader's CONDITIONAL verdict is appropriate, and the reader's weakest_assumption identifies the same load-bearing concern. I sharpen it: the interaction test in a fixed-effects GRBD equates any systematic batch-to-batch difference in the effect with non-reproducibility, but such differences can be reproducible effect moderation. The illustrative analysis has only three sites, and no sensitivity analysis is given; the fixed-effects choice means no variance component is estimated, so the method cannot predict a new site's effect. A random-effects meta-analysis with a prediction interval would directly test the reproducibility interpretation. This does not invalidate the useful message about reporting within-study stability, but it requires the authors to temper the claim that internal replication 'estimates reproducibility' and to report heterogeneity metrics. A secondary issue is that the cited data source (Harrison et al. 2009) is a rapamycin study, not the 17aE2 study described; this undermines independent verification but is secondary to the conceptual concern.","tokens_in":7856,"tokens_out":16590,"duration_ms":181767,"concrete_test":"Re-analyze the three site-specific 17aE2 treatment-effect estimates from the Harrison et al. data with a random-effects meta-analysis (e.g., using the R package metafor), estimating the between-site variance tau-squared, a 95% prediction interval for the effect in a new site, and leave-one-site-out influence diagnostics. If the prediction interval lies mostly above zero or tau-squared is not significantly different from zero, the p=0.024 interaction is better interpreted as context-dependent variation rather than evidence of non-reproducibility; if the interval robustly spans zero, the paper's 'not reproducible' conclusion is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim (abstract, Section 2.2) rests on equating a significant treatment-by-batch interaction with non-reproducibility. This is an interpretive leap, not a statistical necessity. In the GRBD model, the interaction captures systematic differences in treatment effects across batches; those differences can be stable, reproducible properties of the system (e.g., site-specific genetics or diet), in which case the effect is context-dependent rather than irreproducible. The illustrative conclusion that 17aE2 was 'not reproducible across sites' (p=0.024, three sites) is a binary judgment from a test with no effect-size measure, no variance-component estimate, and no sensitivity analysis. Because blocks are treated as fixed effects, the test only compares the observed batches; it does not quantify the between-batch variance needed to predict a new batch's effect. Thus, the framework does not actually estimate generalizable reproducibility from internal replication, and its headline interpretation can mislead when the interaction is driven by one site or by a reproducible moderator.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the concept of 'internal replication' for preclinical experiments, defined as repetition across batches, runs, days, litters, or sites already present within a single study, and argues that such data can be used to assess reproducibility without additional experiments. Six types of internal replication are classified by independence and timing. The authors propose analyzing such designs as generalized randomized block designs (GRBDs), with the treatment-by-batch interaction test used to assess whether the treatment effect is stable across batches. The method is illustrated on a mouse lifespan dataset from a three-site study of 17α-Estradiol, where the site-by-treatment interaction is significant (p = 0.024) and the authors conclude the treatment effect was not reproducible across sites.","tokens_in":8000,"tokens_out":2627,"duration_ms":32217,"significance":"If the interpretive framework is accepted, the paper offers a practical, low-cost way to extract reproducibility information from existing experimental designs, which is valuable given the expense of full replication studies. The six-type taxonomy is clear and useful for practitioners. The illustrative reanalysis is transparent and reproducible from the reported ANOVA table, and the paper explicitly acknowledges the low power of the stringent treatment test. However, the central statistical interpretation—that a significant treatment-by-batch interaction equates to non-reproducibility—needs refinement, because it conflates observed heterogeneity among a fixed set of batches with generalizable instability across future replications.","major_comments":[{"comment":"The claim that a significant treatment-by-batch interaction indicates the experiment is 'not reproducible' is a load-bearing interpretive leap. In a GRBD with fixed blocks, the interaction estimates systematic differences in treatment effects among the observed batches; these differences could be stable, reproducible properties of the system (e.g., site-specific genetics, diet, or husbandry) rather than stochastic instability. The manuscript's own Introduction limits the method to estimating 'the stability of results within an experiment,' which is weaker than 'not reproducible across sites' as stated in the Results. Please reframe the conclusion as evidence of heterogeneity across observed sites, explicitly acknowledge that context-dependence is a possible explanation, and add a discussion of what additional evidence (e.g., external replication or measurement of moderators) would be needed to distinguish stochastic instability from reproducible context-dependence.","section":"Section 2.2, Eq. (1)-(2) and Results"},{"comment":"The paper argues that blocks should be treated as fixed effects because no inference about hypothetical future blocks is intended, yet the abstract and Discussion frame the method as estimating 'reproducibility.' With fixed blocks, the analysis only compares the batches actually observed and provides no estimate of between-batch variance that could be used to predict the behavior of a new batch or site. To support the reproducibility claim, the manuscript should either (a) limit the stated scope to consistency among the observed batches, or (b) supplement the fixed-effects analysis with a random-effects or variance-component analysis (e.g., mixed-effects model with a random site-by-treatment interaction) that quantifies the between-site variance component, along with a discussion of the small number of sites and the accompanying estimation uncertainty.","section":"Section 2.2, fixed vs. random block treatment"},{"comment":"The binary conclusion 'not reproducible across sites' rests on a single p-value (p = 0.024) from three sites, and the visual pattern in Figure 2B suggests the interaction is driven primarily by one site (UT) with a large positive effect and another (TJL) with near-zero effect. The paper does not report effect-size measures for the interaction (e.g., variance component, range of site-specific effects, or a leave-one-site-out analysis). Because the conclusion is central to the paper's message, please add a sensitivity analysis (e.g., re-fitting the model after omitting each site) and report site-specific effect estimates with confidence intervals, so readers can judge how stable the interaction conclusion is.","section":"Results, Table 1 and Figure 2B"}],"minor_comments":[{"comment":"The normality assumption for the lifespan outcome is asserted without diagnostics. Given that lifespan is a bounded, positive outcome, please provide a Q-Q plot or a brief justification (e.g., citing central-limit-theorem robustness for the F-test), or run a sensitivity check using a survival analysis or nonparametric approach.","section":"Section 2.2"},{"comment":"The text refers to 'Figure 3A' and 'Figure 3B', but the corresponding figure is numbered 'Figure 2' in the caption. Please correct the cross-references.","section":"Section 3"},{"comment":"In Figure 1, the distinction between 'fully independent' and 'partially independent' for the staggered case could be made more explicit in the caption; currently the visual difference relies entirely on the text about shared materials or researchers.","section":"Section 1.1"}],"recommendation":"major_revision","confidential_remarks":"The paper's main contribution is conceptual and its statistical execution is clean, but the central interpretive claim needs substantial revision. The fixed-effects GRBD analysis is appropriate for describing the observed data, but the manuscript currently overstates what can be concluded about 'reproducibility' in a generalizable sense. A revision that reframes the claims, adds effect-size or variance-component information, and acknowledges context-dependence as an alternative explanation would make the paper publishable. I do not see evidence of a fundamental statistical error in the ANOVA calculations themselves."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The six-type taxonomy of internal replication is the real contribution here. It gives preclinical researchers a practical vocabulary for spotting replication opportunities they already have, and it is framed well with figure 1. The application to the Harrison mouse data is also done cleanly: the ANOVA table is reported with degrees of freedom and sums of squares, the effect-size plot shows exactly what drives the interaction, and the secondary test using MS(Treatment x Batch) as the error term is a sensible illustration. The paper should get credit for making a textbook technique actionable for a non-specialist audience.\n\nThe soft spots are real but not fatal. The biggest one is exactly what the stress-test note flags: equating a significant treatment-by-batch interaction with 'not reproducible' is an interpretive leap. A stable, systematic site-specific moderator (e.g., genetics or diet) can produce a large interaction, and that is context-dependence, not stochastic instability. The paper half acknowledges this by defining internal replication as stability across the observed blocks, but the abstract and the conclusions use language like 'not reproducible across sites,' which generalizes beyond what a fixed-effects test can support. The test only compares the three sites in hand; it does not estimate the between-batch variance needed to say anything about future batches. I would have liked an effect-size measure for the interaction or a variance-component analysis, even a simple one, to give the binary p-value some context.\n\nTwo smaller items. Normality of the lifespan outcome is asserted without diagnostics; the paper says 'well-approximated' but shows nothing. That is a minor gap, easily fixed. Also, the paper does not provide the exact data subset or analysis code. The ANOVA table is reproducible in principle from the published data, but the subsetting rules are not specified precisely enough to check without guesswork. Minor because the table is there and the original data are public.\n\nWho is this for? Preclinical researchers who design multi-batch or multi-site experiments. They will get a clear framework for reporting stability metrics, and the paper will likely influence reporting standards if it lands in a venue they read. The statistical machinery is standard, but the packaging is new and the example is honest, including the low-power caveat on the second test.\n\nVerdict: send it to peer review. The central claim needs softening and a richer analysis, but the paper is serious, coherent, and useful. I would not desk-reject it.","headline":"A genuinely useful taxonomy for internal replication and a clean illustrative ANOVA, but the paper overreaches when it reads a significant batch-by-treatment interaction as 'not reproducible'.","tokens_in":8475,"tokens_out":2015,"would_cite":true,"duration_ms":25842,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62J10","62K10","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper shows that many preclinical experiments already contain the replication needed to assess reproducibility, and that testing the treatment-by-batch interaction reveals whether an effect is stable across batches.","keywords":["internal replication","reproducibility","treatment-by-batch interaction","generalised randomised block design","pseudoreplication","preclinical experiments","multi-site experiments","17aE2 lifespan"],"falsifier":"Compare interaction p-values from many multi-batch experiments with their later external replication outcomes; if significant treatment-by-batch interactions are not followed by replication failure more often than non-significant ones, the central claim would be refuted.","tokens_in":7656,"feed_emoji":"🧪","tokens_out":5492,"duration_ms":56398,"temperature":0.7,"pith_summary":"Many preclinical experiments are run in batches, across days, litters, runs, or sites, and these batches act as mini-experiments that can be used to judge whether an effect is stable without paying for a separate replication study. The paper formalises this as internal replication, defines six types distinguished by independence and timing, and shows that a treatment-by-batch interaction in a generalised randomised block design is the statistic that measures reproducibility. Reanalysing a three-site mouse lifespan study of 17aE2, the paper finds a significant site-by-treatment interaction (p = 0.024), meaning the drug's effect on lifespan was not consistent across sites even though the average treatment effect was significant (p = 0.002). The point is that reproducibility information is often normalised away or treated as nuisance, when it is already available in the data.","feed_headline":"Reproducibility can be tested with data you already have","feed_subtitle":"A three-site mouse study shows the 17aE2 lifespan effect was not stable across sites.","key_machinery":"The central mechanism is the generalised randomised block design, in which batches such as sites, days, litters, or runs serve as blocks and each treatment appears in every block with genuine replicates. The treatment-by-block interaction term $\\tau\\beta$ carries the argument: testing $H_0: (\\tau\\beta)_{ij}=0$ for all $i,j$ directly assesses internal replication, while a second F-test can use the interaction mean square as the error term for the treatment effect, effectively treating the genuine replicates as pseudoreplicates. The design is specified as $y_{ijk} = \\mu + \\tau_i + \\beta_j + (\\tau\\beta)_{ij} + \\varepsilon_{ijk}$ with $\\varepsilon_{ijk} \\sim \\mathrm{Normal}(0,\\sigma)$, and the interaction test is what turns internal replication into a reproducibility assessment.","core_discovery":"The paper's central claim is that the stability of an experimental effect can be estimated from a single study whenever the experimental units are grouped into two or more batches, and that the right statistic is the treatment-by-batch interaction. Each batch is treated as a block in a generalised randomised block design with genuine replication, and a significant interaction means the effect varies across batches more than sampling variability alone would predict, which the paper interprets as non-reproducibility. In the illustrative reanalysis of Harrison and colleagues' mouse lifespan data, the site-by-treatment interaction was significant (p = 0.024), while the average 17aE2 effect was significant only when tested against the residual error (p = 0.002) and not when tested against the site-to-site variation (p = 0.254). The paper argues that reporting only the main effect in such cases misleads readers about how much the result can be trusted.","pith_inferences":["By extension, the same interaction test could be run on published multi-batch datasets, such as multi-site or multi-litter studies, to flag effects whose stability is questionable without running new experiments.","A natural next step would be to calibrate the interaction p-value against external replication outcomes; if the two do not correlate, the link between internal instability and replication failure would be weakened.","The approach implicitly assumes batch differences that produce a significant interaction matter for generalisation, so a reader applying it to a new field should decide whether the batch factor is one over which reproducibility is actually desired.","Standardising a minimal set of design features, namely crossing, balance, and independence of batches, could turn the interaction F-test into a routinely reported reproducibility statistic."],"forward_implications":["A significant treatment-by-batch interaction means the main effect is not stable across batches, so reporting only the average effect would be misleading.","Researchers can report an internal-reproducibility statistic from existing data at no experimental cost by modelling batches as blocks in a generalised randomised block design.","The treatment effect can be tested against the batch-interaction variance; a significant result is strong evidence, though the test is usually underpowered because the effective sample size is the number of batches.","Designing experiments with treatments crossed with batches, balanced samples, and as much batch independence as possible makes the interaction test informative.","Internal-reproducibility metrics should become a standard reporting item in preclinical studies."],"supporting_citations":[{"why":"Supplies the three-site mouse lifespan dataset used to illustrate the internal-replication analysis.","marker":"Harrison et al. (2009)"},{"why":"Defines the generalised randomised block design and the genuine-replication requirement.","marker":"Hinkelmann and Kempthorne (2008)"},{"why":"Provides the F-test framework and error-term choice for testing the treatment-by-block interaction.","marker":"Casella (2008)"},{"why":"Clarifies what counts as N and supports treating genuine replicates as pseudoreplicates when testing against the interaction.","marker":"Lazic et al. (2018)"},{"why":"Supplies the internal-versus-external validation analogy that motivates internal replication.","marker":"Steyerberg (2019)"},{"why":"Supports the design recommendation that litters or batches be balanced across treatments.","marker":"Lazic (2016)"}],"fun_headline_variants":["Test reproducibility using batches you already have","Mouse lifespan effect unstable across three sites","Internal replication: gauge stability without new experiments","Reproducibility from existing data: a three-site demo","Batch interaction statistic exposes effect instability"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim rests on treating a significant treatment-by-batch interaction as evidence that the effect is not reproducible, rather than as benign variation in scale or conditions that would still allow the effect to generalise.","fun_headline_variants_meta":{"raw":{"variants":["Test reproducibility using batches you already have","Mouse lifespan effect unstable across three sites","Internal replication: gauge stability without new experiments","Reproducibility from existing data: a three-site demo","Batch interaction statistic exposes effect instability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000179,"raw_usage":{"total_tokens":1255,"prompt_tokens":855,"completion_tokens":400,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":471,"completion_tokens_details":{"reasoning_tokens":334}},"tokens_in":471,"tokens_out":400,"duration_ms":4873,"temperature":1.0,"reasoning_tokens":334,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:01:43.363104+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare interaction p-values from many multi-batch experiments with their later external replication outcomes; if significant treatment-by-batch interactions are not followed by replication failure more often than non-significant ones, the central claim would be refuted.","supporting_citations":[],"review_version":1}