REVIEW 5 minor 26 references
Anytime-Valid Evidence for Prespecified Predictive Corrections
T0 review · 0 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper proves that a prespecified predictive correction can be confirmed sequentially by multiplying corrected-to-source likelihood ratios, with a martingale boundary crossing that gives anytime-valid relative confirmation under…
desk verdict A sound, honest packaging of likelihood-ratio monitoring for prespecified predictive corrections; the anytime-valid guarantee is real but strictly conditional on the source predictive being exactly right, and the paper says so itself. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the normalized tilt $e(x,y)=h(x,y)/Z_h(x,D_{\mathrm{tr}})$, where $Z_h(x,D_{\mathrm{tr}})=\int h(x,y)p_0(y\mid x,D_{\mathrm{tr}})\,dy$ is the normalizer of the tilt under the source predictive. Under the source predictive null, $e_i$ is a conditional e-value, and the product $M_t=\prod_{i=1}^t e_i$ is a nonnegative martingale, so Ville's inequality supplies the anytime-valid crossing bound. The same ratio equals the corrected-to-source predictive likelihood ratio, making $\log M_t$ exactly the cumulative predictive log-score advantage of the corrected predictive over the source predictive.
What would settle it
Generate outcomes from a Gaussian with variance $c_{\mathrm{mis}}\sigma^2$ while monitoring the variance-correction e-process computed with model variance $\sigma^2$ and correction factor $c=1.8$; under the paper's Table 4, the false-confirmation rate jumps from $0.027$ at $c_{\mathrm{mis}}=1$ to $0.892$ at $c_{\mathrm{mis}}=1.5$, so a reader can directly check whether the bound holds only when the source predictive null is true. Alternatively, simulate a target $q$ with $E_q[h(x,Y)]>Z_h(x,D_{\mathrm{tr}})$ and verify whether the crossing probability exceeds $\alpha$ when the moment condition of Proposition 3 fails.
Extended reading notes
Core claim
The central discovery is that any fixed nonnegative tilt $h$ with finite positive normalizer turns the source predictive into a corrected predictive, and the corrected-to-source likelihood ratio $e_i=h(X_i,Y_i)/Z_h(X_i,D_{\mathrm{tr}})$ is a conditional e-value under the source predictive null. Consequently, $M_t=\prod_{i=1}^t e_i$ is a nonnegative martingale, and Ville's inequality gives $\sup_{P\in P_0^{\mathrm{pred}}} P(\sup_{t\ge 0}M_t>1/\alpha)\le \alpha$. This makes the stopping time $\tau^*=\inf\{t:M_t>1/\alpha\}$ an anytime-valid test for relative confirmation of the correction, valid under continuous monitoring, optional stopping, and arbitrary or adaptively selected inputs. The paper further shows that the expected log-growth under a target predictive $q$ is $\Gamma_h(x)=D_{\mathrm{KL}}(q(\cdot\mid x)\|p_0(\cdot\mid x,D_{\mathrm{tr}}))-D_{\mathrm{KL}}(q(\cdot\mid x)\|p_h(\cdot\mid x,D_{\mathrm{tr}}))$, so positive drift means the corrected predictive is closer to the target in conditional KL divergence. It also identifies a correction-dependent half-space of misspecified targets where the same $\alpha$ bound persists, and derives a reciprocal boundary for refutation plus an overshoot identity explaining why the realized null crossing probability is often below $\alpha$.
Load-bearing premise
The result hinges on the null that each outcome is generated exactly from the fixed source predictive $p_0(y\mid x,D_{\mathrm{tr}})$; if the source predictive is miscalibrated, the same false-confirmation bound can fail, as the paper's own sweep illustrates with confirmation rates rising from $0.027$ to $0.892$ when the true variance is $1.5$ times the modeled variance.
Editorial extensions
If this is right
- A boundary crossing means the corrected predictive has accumulated more than $\log(1/\alpha)$ nats of observed log-score advantage, so the procedure justifies relative confirmation of the prespecified correction without estimating the full target distribution.
- Validity holds for arbitrary and adaptively selected input sequences, so an experimenter can actively choose informative inputs to accelerate evidence accumulation without changing the source-null error guarantee.
- Under i.i.d. or stationary-ergodic sampling, $\frac1t\log M_t\to \Gamma(q;h)$ almost surely, so the correction is eventually confirmed whenever the corrected predictive has strictly smaller expected log loss, and is eventually refuted when the source predictive does.
- Finite-horizon crossing bounds in Corollary 2 control the probability of delayed confirmation, translating the asymptotic growth rate into explicit lower bounds on confirmation by a given time.
- For a correctly specified correction, the asymptotic growth rate is the expected KL divergence between the corrected and source predictives, and eventual confirmation is almost sure whenever that divergence has positive expectation over the input distribution.
Reading between the lines
- The conditional nature of the e-process makes it insensitive to pure covariate shift, so a practitioner who wants a full distribution-shift alarm should pair this conditional-outcome monitor with a separate input-distribution monitor; the paper itself notes this separation.
- The overshoot identity suggests a practical calibration check: report the conditional mean overshoot $E[M_{\tau^*}\mid \tau^*<\infty]$ alongside the crossing, since the realized null crossing probability equals the corrected-predictive crossing probability divided by that mean overshoot.
- The protected half-space criterion gives a pre-deployment robustness test: before monitoring, one can check whether plausible misspecified target distributions keep $E_q[h(x,Y)]\le Z_h(x,D_{\mathrm{tr}})$ at the inputs likely to be seen; if not, the anytime-valid bound is not guaranteed.
- The mixture construction provides a family-level evidence claim, not a license to select the best component; a data-driven choice among corrections would require a separately prespecified error allocation or a different rule.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops an anytime-valid framework for evaluating a prespecified predictive correction against a fixed source predictive distribution. A nonnegative tilt h transforms the source predictive p0 into a corrected predictive ph, and the normalized likelihood ratio h/Z_h is shown to be a conditional e-value whose running product is a nonnegative martingale under the source predictive null. This yields a Ville-type bound on false confirmation under continuous monitoring, optional stopping, and arbitrary or adaptively selected input sequences. The paper also derives a conditional drift decomposition for evidence growth under an arbitrary target, identifies a correction-dependent half-space in which the false-confirmation bound persists under misspecification, gives reciprocal refutation boundaries and an overshoot identity, and covers label-shift, Gaussian mean and variance corrections, exponential-family tilts, prespecified mixtures, predictable tilts, and beyond-tolerance comparisons. Synthetic experiments verify the analytic drift predictions and quantify failure modes under source miscalibration and cross-family targets.
Significance. If the results are taken as stated, the paper provides a clean and useful extension of e-process methodology from simple likelihood-ratio monitoring to prespecified predictive corrections, with explicit treatment of adaptive inputs, drift rates, and a well-characterized robustness region. The proofs are standard martingale and Ville arguments, and the experiments reproduce the analytic drift rates to Monte Carlo accuracy. The paper is unusually transparent about its central limitation: the anytime-valid guarantee is conditional on the fitted source predictive being the true conditional law, and Section 6 as well as the miscalibration experiments state this clearly. The contribution is incremental rather than revolutionary, but it is coherent, reproducible, and likely to be of practical interest for sequential model-monitoring and calibration-transfer problems.
minor comments (5)
- [Abstract and Section 1] The abstract and introduction state the result as 'anytime-valid evidence' without immediately qualifying that the guarantee is conditional on the fitted source predictive p0 being the exact conditional law of Y_i given X_i and Dtr; Section 6 does acknowledge this, but the abstract should carry the same qualification to prevent overstatement in the paper's main public-facing claim.
- [Section 3.6, Algorithm 1] The pseudocode line 'else if exact two-boundary mode and S_i < log α' would be clearer if it explicitly noted that the lower refutation boundary is available only in exact-normalization mode and when h is strictly positive p0-almost surely; the surrounding text states this, but a comment in the algorithm would help avoid misuse.
- [Section 5.1] The Monte Carlo standard error for the confirmation rates is correctly stated as at most 0.0071, but the paper does not report standard errors for the mean final log wealth values; adding them would make it easier to judge the agreement with the analytic drift predictions.
- [Section 5.10] The statement that the cumulative confirmation rate is 'identical at t = 1000 and t = 5000' is an empirical observation; the text could state the first time at which the rate stabilizes or note explicitly that all crossings occurred early, which the null drift implies.
- [References] The four self-citations to Choi (2026a-d) are preprints or workshop papers; the published version should verify availability and provide arXiv identifiers or DOIs where applicable.
Circularity Check
No significant circularity: the e-process, drift decomposition, and protected half-space follow directly from the stated definitions and standard martingale arguments; self-citations are contextual only.
full rationale
Theorem 1 constructs M_t as a product of h/Z_h with Z_h defined as the expectation of h under p0; the conditional e-value property E[e_i|G_i]=1 is exactly the normalization, and Ville's inequality then yields the bound. This is a direct mathematical derivation from stated definitions, not a fit disguised as prediction. Proposition 2's drift decomposition is a calculation of E_q[log e_i|G_i] and the KL identity; no parameter is fitted. Proposition 3's protected half-space is explicitly the condition E_q[h]<=Z_h, which is equivalent to the one-step conditional mean bound; the paper states it is 'exactly the condition' rather than claiming an external robustness result. The Section 6 limitation that guarantees are conditional on the fitted source predictive is disclosed, and the experiments use an oracle Gaussian p0 to isolate the e-process from estimation error. Self-citations (Choi 2026a-d) appear only as context for label-shift, covariate-balance, and conformal Bayes special cases; none is load-bearing for Theorem 1 or its corollaries. The paper even avoids a potential circularity in Proposition 4 by using a change-of-measure argument instead of uniform integrability, explicitly noting that the latter route would be circular (Section A.13). Accordingly, no step reduces by construction to its own inputs.
Assumptions & free parameters
assumptions (7)
- domain assumption Source predictive null is exactly specified: Y_i | G_i ~ p0(cdot | X_i, Dtr) for every i (Section 3.1, Eq. (5)).
- domain assumption Normalizer finiteness and positivity: 0 < Z_h(x,Dtr) = integral h(x,y) p0(y|x,Dtr) dy < infinity for every relevant x (Eq. (7)).
- standard math Measurability: h is jointly measurable and x maps to Z_h(x) is measurable (Section 3.1).
- domain assumption For Proposition 2: the expected absolute log e-value is finite at each step and a conditional second-moment bound holds (Section 3.3).
- domain assumption For Corollary 1: the pair process is i.i.d. or stationary-ergodic (Section 3.3).
- domain assumption For Proposition 7: the exponential-tilt family has a common interval I on which every log-normalizer psi_x is finite (Section 4.5).
- domain assumption For reciprocal refutation: h(x,y) > 0 for p0-almost every y (Section 3.5.1).
Cite this review
Pith. "Pith review of Anytime-Valid Evidence for Prespecified Predictive Corrections." pith.science (2026). https://pith.science/paper/TUGPEYSR
@misc{pith2026260808174,
author = {Pith},
title = {Pith review of: Anytime-Valid Evidence for Prespecified Predictive Corrections},
year = {2026},
howpublished = {\url{https://pith.science/paper/TUGPEYSR}},
note = {Machine review of arXiv:2608.08174}
}
read the original abstract
A predictive correction is a prespecified modification of an existing predictive distribution intended to reflect an anticipated change in future outcomes given their inputs, motivated, for example, by instrument recalibration, assay drift, or a known intervention. We study how to accumulate anytime-valid evidence that such a correction predicts incoming target outcomes better than the uncorrected source predictive distribution. A fixed nonnegative tilt transforms the source predictive into a corrected predictive, and the corrected-to-source predictive likelihood ratio is a conditional e-value whose running product forms an e-process. This process remains valid under optional stopping and arbitrary input sequences, including adaptively selected ones, while its logarithm equals the cumulative predictive log-score advantage of the correction. A conditional drift decomposition characterizes evidence growth under an arbitrary target predictive distribution, and a correction-dependent half-space identifies misspecified target distributions for which the same false-confirmation bound continues to hold. When the predictive likelihood ratio is strictly positive, its reciprocal yields an anytime-valid refutation boundary, while an overshoot identity explains why the realized null crossing probability may fall below the nominal level. Label-shift, conditional mean and variance, subgroup-specific, and exponential-family corrections arise as special cases. Prespecified mixtures accommodate uncertainty over corrections, predictable tilts permit adaptive betting, and beyond-tolerance comparisons target changes large enough to justify action. Cross-family calculations and synthetic experiments show that a boundary crossing supports the proposed correction relative to its reference but does not uniquely identify the mechanism responsible for the shift.
Figures
Reference graph
Works this paper leans on
-
[1]
M., Kundaje, A., and Shrikumar, A
Alexandari, A. M., Kundaje, A., and Shrikumar, A. (2020). Maximum likelihood with bias-corrected calibration is hard-to-beat at label shift adaptation. InProceedings of the International Conference on Machine Learning (ICML)
work page 2020
-
[2]
Angelopoulos, A. N. and Bates, S. (2023). Conformal prediction: A gentle introduction.Foundations and Trends® in Machine Learning, 16(4):494–591
work page 2023
-
[3]
Choi, S. (2026a). Anytime-valid confirmation of covariate balance for prespecified corrections.Preprint arXiv:2607.23157
work page Pith review arXiv 2026
-
[4]
Choi, S. (2026b). Anytime-valid confirmation of label-shift corrections. InICML 2026 Workshop on Hypothesis Testing
work page 2026
-
[5]
Choi, S. (2026c). Conformal Bayes for two-sided censored Gaussian regression under label shift.Preprint arXiv:2607.02173
work page Pith review arXiv 2026
-
[6]
Dawid, A. P. (1984). Present position and potential developments: Some personal views: Statistical theory: The prequential approach.Journal of the Royal Statistical Society Series A, 147(2):278–292. 40 Anytime-Valid Evidence for Prespecified Predictive Corrections
work page 1984
-
[7]
Fong, E. and Holmes, C. (2021). Conformal Bayesian computation. InAdvances in Neural Information Processing Systems (NeurIPS)
work page 2021
-
[8]
Garg, S., Wu, Y., Balakrishnan, S., and Lipton, Z. C. (2020). A unified view of label shift estimation. In Advances in Neural Information Processing Systems (NeurIPS)
work page 2020
Show all 26 references
-
[9]
E., Li, C., and Rabinovic, A
Johnson, W. E., Li, C., and Rabinovic, A. (2007). Adjusting batch effects in microarray expression data using empirical Bayes methods.Biostatistics, 8(1):118–127
2007
-
[10]
J., Karthikesalingam, A., Suleyman, M., Corrado, G., and King, D
Kelly, C. J., Karthikesalingam, A., Suleyman, M., Corrado, G., and King, D. (2019). Key challenges for delivering clinical impact with artificial intelligence.BMC Medicine, 17(1)
2019
-
[11]
Kennedy, M. C. and O’Hagan, A. (2001). Bayesian calibration of computer model.Journal of the Royal Statistical Society Series B, 63(3):425–464
2001
-
[12]
Baggerly, K., and Irizarry, R. A. (2010). Tackling the widespread and critical impact of batch effects in high-throughput data.Nature Review Genetics, 11:733–739
2010
-
[13]
C., Wang, Y.-X., and Smola, A
Lipton, Z. C., Wang, Y.-X., and Smola, A. J. (2018). Detecting and correcting for label shift with black box predictors. InProceedings of the International Conference on Machine Learning (ICML)
2018
-
[14]
and Ramdas, A
Podkopaev, A. and Ramdas, A. (2021). Distribution-free uncertainty quantification for classification under label shift. InProceedings of the Annual Conference on Uncertainty in Artificial Intelligence (UAI)
2021
-
[15]
and Ramdas, A
Podkopaev, A. and Ramdas, A. (2022). Tracking the risk of a deployed model and detecting harmful distribution shifts. InProceedings of the International Conference on Learning Representations (ICLR)
2022
-
[16]
Qin, S. J. (2012). Survey on data-driven industrial process monitoring and diagnosis.Annual Reviews in Control, 36(2):220–234. Qui˜ nonero-Candela, J., Sugiyama, M., Schwaighofer, A., and Lawrence, N. D., editors (2009).Dataset Shift in Machine Learning. MIT Press
2012
-
[17]
Ramdas, A., Gr¨ unwald, P., Vovk, V., and Shafer, G. (2023). Game-theoretic statistics and safe anytime- valid inference.Statistical Science, 38(4):576–601
2023
-
[18]
Shafer, G. (2021). Testing by betting: A strategy for statistical and scientific communication.Journal of the Royal Statistical Society Series A, 184(2):407–431
2021
-
[19]
and Vovk, V
Shafer, G. and Vovk, V. (2019).Game-Theoretic Foundations for Probability and Finance. Wiley
2019
-
[20]
and Saria, S
Subbaswamy, A. and Saria, S. (2020). From development to deployment: dataset shift, causality, and shift-stable models in health AI.Biostatistics, 21(2):345–352
2020
-
[21]
and Kawanabe, M
Sugiyama, M. and Kawanabe, M. (2012).Machine Learning in Non-Stationary Environments: Introduction to Covariate Shift Adaptation. MIT Press
2012
-
[22]
J., Barber, R
Tibshirani, R. J., Barber, R. F., Cand` es, E. J., and Ramdas, A. (2019). Conformal prediction under covariate shift. InAdvances in Neural Information Processing Systems (NeurIPS)
2019
-
[23]
Ville, J. (1939). ´Etude Critique de la Notion de Collectif. PhD thesis, Universit´ e de Paris
1939
-
[24]
(2005).Algorithmic Learning in a Random World
Vovk, V., Gammerman, A., and Shafer, G. (2005).Algorithmic Learning in a Random World. Springer
2005
-
[25]
and Wang, R
Vovk, V. and Wang, R. (2021). E-values: Calibration, combination and applications.The Annals of Statistics, 49(3):1736–1754
2021
-
[26]
Wald, A. (1945). Sequential tests of statistical hypotheses.The Annals of Mathematical Statistics, 16(2):117–186. Workman Jr., J. J. (2018). A review of calibration transfer practices and instrument differences in spectroscopy.Applied Spectroscopy, 72(3):340–365. 41 Anytime-Va...
1945
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.