REVIEW 2 major objections 5 minor 9 references
ALL-IN meta-analysis for flexibility and validity in prospective and retrospective evidence synthesis
T0 review · 2 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Widened confidence intervals can keep a repeatedly updated meta-analysis at 95% coverage with no maximum sample size.
desk verdict A clear, honest commentary on the authors' own anytime-valid meta-analysis method, worth publishing but with one unstated approximation about estimated standard errors. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine of the argument is the anytime-valid confidence interval, a confidence sequence: a nested family of intervals constructed so that, no matter when the analyst stops and looks, the true effect lies inside with probability at least 95%. In ALL-IN meta-analysis this appears as a single widened multiplier replacing the familiar $1.96$ times the standard error — about $3.037$ in the BCG example at a reasonable sample size — and the construction needs no maximum sample size or stopping rule. That multiplier is what converts an ordinary cumulative meta-analysis into one whose updates are all simultaneously valid.
What would settle it
Simulate a meta-analysis in which trials accrue one by one and their sizes, stopping times, or inclusion are triggered by the current interim estimate; construct ALL-IN intervals from aggregate log hazard ratios with estimated standard errors exactly as in Figure 1; if, under any such decision rule, the interval covers the true effect in fewer than 95% of repetitions, the central coverage claim is refuted.
Extended reading notes
Core claim
The paper's central claim is that ALL-IN meta-analysis is anytime-valid: type-I error for tests and coverage for confidence intervals are guaranteed simultaneously at every update, for unlimited updating, with no maximum sample size. Conventional meta-analysis loses this guarantee because each update is another chance to be wrong, and result-dependent decisions about which trials exist and when they stop — accumulation bias and unplanned early stopping — make the problem worse. ALL-IN absorbs these scenarios by widening intervals; in the paper's example a fixed-effects estimate with standard error 0.084 is reported with an anytime-valid interval roughly three standard errors wide, $[0.76;\,1.27]$, instead of the usual $1.96$-standard-error interval. The same forest plot can be read as before, and the wider interval automatically discourages strong conclusions when data are still thin.
Load-bearing premise
The anytime-valid coverage guarantee is imported from abstract anytime-valid inference and is assumed to remain exact when the method is applied to aggregate meta-analysis summaries in which log hazard ratios are treated as normally distributed with known standard errors, even though those standard errors are estimated from the data.
Editorial extensions
If this is right
- Living and prospective meta-analyses can be updated indefinitely, and each new forest plot still carries valid 95% intervals.
- Collaborative pandemic evidence synthesis can conclude from interim trials in real time, as the BCG example concluded early from two unfinished trials.
- Retrospective meta-analyses gain validity under accumulation bias and uncontrolled early stopping because coverage holds under any result-dependent continuation rule.
- Existing alpha-spending approaches are covered as a special case, but ALL-IN removes the need to pre-specify a maximum sample size.
- When the trial set is closed and no update-driven decisions shaped it, conventional intervals are shorter and should be preferred.
Reading between the lines
- One could extend the same widening construction to effect measures other than the log hazard ratio, such as risk differences or odds ratios, whenever a normal-with-known-standard-error approximation is defensible; the paper itself only demonstrates the log-hazard-ratio case.
- The width multiplier depends on a chosen "reasonable" sample size, so a natural next step is to make that tuning parameter explicit and study how early-interval width changes with it; the paper does not provide this sensitivity analysis.
- Because coverage holds under arbitrary decision rules, a practical guidance rule could tell analysts when to switch from conventional to ALL-IN intervals; the paper leaves that operational choice open.
- The same logic should protect meta-analyses used by regulators or guideline panels that combine trials stopped for futility, without needing to reconstruct the actual stopping rules; the paper states the principle but does not apply it to that governance setting.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript is a commentary advocating ALL-IN meta-analysis, an anytime-valid approach that widens meta-analytic confidence intervals so that the meta-analysis can be updated indefinitely while preserving error control. It presents the BCG-COVID case study, argues for applications in both pandemic and non-pandemic settings, and discusses a key limitation: the wider intervals are wasteful when there is a fixed, pre-planned set of trials with no updating. The paper does not derive new statistical theory; it imports anytime-valid confidence sequences from Ramdas et al. (2023) and ter Schure & Grünwald (2025) and illustrates their use in a familiar forest-plot workflow.
Significance. If the central claim holds as stated, ALL-IN meta-analysis would be a practically valuable tool for living and prospective systematic reviews, because it preserves coverage without requiring a pre-specified stopping rule or maximum sample size. The paper's strengths are that it connects established anytime-valid theory to a concrete applied workflow, provides R code and a replication package for the figures, and is explicit about the efficiency cost when no updating occurs. The main caveat is that the unconditional guarantee as stated in the Introduction is stronger than what the cited theory delivers in the exact aggregate-data workflow shown in Figure 1.
major comments (2)
- [Introduction and Figure 1 caption] The Introduction states that 'Type-I error for tests and coverage for confidence intervals are guaranteed for unlimited updating, with no maximum sample size.' The cited anytime-valid theory (Ramdas et al., 2023) gives exact coverage for normal observations with known variances, but the worked example in Figure 1 uses an inverse-variance fixed-effect estimate whose weights are computed from SE(logHR) values estimated from each trial, and the same aggregate-summary setting is used throughout Sections 2.1 and 2.2. Replacing known variances by estimates breaks the martingale property on which the exact guarantee rests, so the coverage guarantee is at best approximate or asymptotic in the actual workflow. The authors should state the known-variance condition explicitly and add a caution about small or sparse meta-analyses where estimation error in the weights is non-negligible.
- [Sections 2.1 and 2.2] Section 2.1 claims that anytime-valid intervals 'are guaranteed to cover the true effect under any such rule,' referring to result-dependent decisions to start new trials, and Section 2.2 says ALL-IN 'absorbs all of these scenarios' for trial stopping. The anytime-validity literature directly guarantees validity under arbitrary stopping times, but the transfer to a meta-analytic setting with estimated trial variances and potential heterogeneity or selection mechanisms that change the distribution of the included effect estimates requires modeling assumptions that the paper does not state. The authors should either specify the exact model under which the claim is a theorem (for example, fixed-effect normal summary statistics with known variances and decisions that depend only on past summaries) or rephrase the claim as a heuristic extension of anytime-valid validity.
minor comments (5)
- [Abstract and Section 2] The words 'meta-analyist' and 'analysist' in the Abstract and the Introduction should be corrected to 'meta-analyst' and 'analyst'.
- [Figure 1 caption] The country abbreviations in the Figure 1 caption use inconsistent spelling (for example, 'South-Africa' versus 'South Africa' elsewhere) and the list is difficult to parse; consider adding semicolons or making the structure more uniform.
- [References] The reference 'Moldvay, J. er Schure, J.A., (2026)' appears to have a typo; this should be 'ter Schure'.
- [Section 2.2] The phrase 'the disappointing effect of the BCG vaccine' in Section 2.2 is informal for a methods commentary; consider replacing it with 'the observed null result' or similar.
- [Section 3] The limitation section acknowledges that the wider intervals are an avoidable cost when there is nothing to update, but it does not mention the known-variance approximation issue raised above; adding a sentence there would help readers who focus on the limitations section.
Circularity Check
No significant circularity: the anytime-valid coverage guarantee is imported from published external theory, not derived from the paper's own conclusion.
full rationale
The paper's central claim—type-I error control and confidence-interval coverage under unlimited updating—is not defined into existence, fitted from data, or renamed from its own inputs. It is explicitly imported from externally published anytime-valid inference theory: Ramdas et al. (2023) and ter Schure & Grünwald (2025). Those works establish the e-value/confidence-sequence theorems on which the ALL-IN widening factor rests. The present commentary applies that theory to aggregate meta-analysis in Figure 1, where log hazard ratios are treated as normal with fixed standard errors; this is an application of an external result, and the possible gap between the exact known-variance theory and estimated-SE practice is a correctness/approximation concern, not a circular reduction. No parameter is fitted to a subset of data and then 'predicted'; the 3.037 half-width factor is an anytime-valid critical value from the cited theory, not an empirical fit. The self-citations are numerous, but they point to published methods, proofs, and replication packages, and the load-bearing mathematical content is the independently citable anytime-validity literature. The examples using BCG data are retrospective illustrations, not predictions derived from the method's assumptions. No specific equation or definition in the paper equates the conclusion with its input by construction, so no circular step can be exhibited under the review criteria.
Assumptions & free parameters
free parameters (1)
- smallest effect size of interest
assumptions (4)
- standard math Anytime-valid confidence sequences have coverage under arbitrary stopping times (Ville's inequality / game-theoretic statistics).
- domain assumption Log hazard ratio estimates are approximately normal with known standard error.
- domain assumption A fixed-effects model is the correct combination model for the trials.
- ad hoc to paper A smallest effect size of interest and an alpha level are fixed in advance.
Cite this review
Pith. "Pith review of ALL-IN meta-analysis for flexibility and validity in prospective and retrospective evidence synthesis." pith.science (2026). https://pith.science/paper/HUHSCVJO
@misc{pith2026260802105,
author = {Pith},
title = {Pith review of: ALL-IN meta-analysis for flexibility and validity in prospective and retrospective evidence synthesis},
year = {2026},
howpublished = {\url{https://pith.science/paper/HUHSCVJO}},
note = {Machine review of arXiv:2608.02105}
}
read the original abstract
ALL-IN meta-analysis was developed and first applied during the COVID-19 pandemic. While this setting inspired its name, ALL-IN can also benefit non-pandemic circumstances. Conventional meta-analysis loses its coverage when updated repeatedly over time and when the decisions to initiate new trials and synthesize them depend on the results within the meta-analysis (accumulation bias). ALL-IN meta-analysis is anytime-valid. In its simplest form, ALL-IN meta-analysis stays familiar to run and read based on forest plots with confidence intervals that are wider than standard ones. Within collaborative prospective meta-analysis, the payoff is flexibility and speed, with fast sharing of individual participant data or harmonized aggregate data. Outside of this setting, the payoff is validity, when the decisions that shape the evidence (when to stop trials and whether new ones start) are typically outside the control of the meta-analyst. On the one hand, ALL-IN meta-analysis becomes inefficient when there is a maximum sample size or a stopping rule that the meta-analyist can control. On the other hand, it enables adaptations for any evidence synthesis to halfway become living, prospective or even real-time on interim trial results, without complicating the statistics.
Figures
Reference graph
Works this paper leans on
-
[1]
Bauer, P., Koenig, F., Brannath, W., & Posch, M. (2010). Selection and bias—two hostile brothers. Statistics in Medicine, 29(1), 1-13. https://doi.org/10.1002/sim.3716 Carrero Longlax, S., Koster, K. J., Kamat, A. M., Lozano, M., Lerner, S. P., Hannigan, R., ... & DiNardo, A. R. (2025). BCG-induced DNA methylation changes improve coronavirus disease 2019 ...
-
[48]
https://doi.org/10.1016/j.eclinm.2022.101414 Van den Hoogen, G., Upton, C.M., ter Schure, J.A. (2026), Reproducibility Data for SA trial in ALL-IN-META-BCG-CORONA, DANS Data Station Life Sciences, V1. https://doi.org/10.17026/LS/WKPOX8 van Werkhoven, C.H., Bonten, M.J.M., ter Schure, J.A.. (2025) Reproducibility Data for NL trial in ALL-IN-META-BCG-CORONA...
-
[355]
https://doi.org/10.1136/bmj.i5440 Madsen, A. M. R., Schaltz-Buchholzer, F., Benfield, T., Bjerregaard-Andersen, M., Dalgaard, L. S., Dam, C., ... & Benn, C. S. (2020). Using BCG vaccine to enhance non-specific protection of health care workers during the COVID-19 pandemic: A structured summary of a study protocol for a randomised controlled trial in Denma...
-
[481]
https://doi.org/10.1186/s13063-020-04389-w Ten Doesschate, T., van der Vaart, T. W., Debisarun, P. A., Taks, E., Moorlag, S. J., Paternotte, N., ... & van Werkhoven, C. H. (2022). Bacillus Calmette-Guérin vaccine to reduce healthcare worker absenteeism in COVID-19 pandemic, a randomized controlled trial. Clinical Microbiology and Infection, 28(9), 1278-12...
-
[799]
https://doi.org/10.1186/s13063-020-04714-3 Madsen, A. M. R., Schaltz-Buchholzer, F., Nielsen, S., Benfield, T., Bjerregaard-Andersen, M., Dalgaard, L. S., ... & Benn, C. S. (2024). Using BCG vaccine to enhance nonspecific protection of health care workers during the COVID-19 pandemic: a randomized controlled trial. The Journal of infectious diseases, 229(...
-
[881]
https://doi.org/10.1186/s13063-020-04822-0 Lund, H., Brunnhuber, K., Juhl, C., Robinson, K., Leenaars, M., Dorch, B. F., ... & Chalmers, I. (2016). Towards evidence based research. Bmj,
-
[2002]
WHO. (2022) CORE PROTOCOL - An international adaptive multi-country randomized,placebo-controlled, double-blinded trial of the safety and efficacy of treatments for patients with monkeypox virus disease. Available from: https://www.who.int/publications/m/item/core-protocol---an-international-adaptive-multi- country-randomized-placebo-controlled--double-bl...
work page 2022
-
[2023]
BCG Vaccination of Health Care Workers Does Not Reduce SARS-CoV-2 Infections nor Infection Severity or Duration: a Randomized Placebo-Controlled Trial. mBio 14:e00356-23. https://doi.org/10.1128/mbio.00356-23 Dos Anjos, L. R. B., da Costa, A. C., Cardoso, A. D. R. O., Guimarães, R. A., Rodrigues, R. L., Ribeiro, K. M., ... & Junqueira-Kipnis, A. P. (2022)...
arXiv 2022
Show all 9 references
-
[2024]
Junqueira-Kipnis, A
Available from www.cochrane.org/handbook. Junqueira-Kipnis, A. P., Dos Anjos, L. R. B., Barbosa, L. C. D. S., da Costa, A. C., Borges, K. C. M., Cardoso, A. D. R. O., ... & Kipnis, A. (2020). BCG revaccination of health workers in Brazil to improve innate immune responses agai...
2020
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.