REVIEW 2 major objections 4 minor 29 references
A prespecified covariate correction can be certified ε-balanced by sequential monitoring, with false-confirmation probability at most a prespecified δ.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 03:25 UTC pith:ZXTBHMIO
load-bearing objection Sound anytime-valid confirmation result, but the finite-source δ+η guarantee rests on a source-interval construction that is assumed and never supplied — worth refereeing, likely major revision. the 2 major comments →
Anytime-Valid Confirmation of Covariate Balance for Prespecified Corrections
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that confirmation—not estimation—is the right inferential target after a correction is fixed. For a finite class of balancing functions f1,...,fm and tolerances ε_j, it constructs simultaneous time-uniform confidence sequences Cn,j for the target means μ_{t,j}, and declares ε-balance confirmed at the first time every Cn,j is contained in [μ_{w,j}−ε_j, μ_{w,j}+ε_j]. Theorem 2 states that if |μ_{t,j}−μ_{w,j}| > ε_j for some j, then the probability that this stopping time is ever finite is at most δ, because false confirmation can only occur when the simultaneous coverage event fails. With finite source data, the band is replaced by the intersection over all plausib
What carries the argument
The load-bearing object is a simultaneous time-uniform confidence sequence: intervals Cn,j such that P(μ_{t,j} ∈ Cn,j for all n≥1 and all j≤m) ≥ 1−δ. The stopping rule is containment—stop at the first n when every Cn,j sits inside the tolerance band centered at the weighted source moment μ_{w,j}. This converts a running empirical diagnostic into a population-level certificate because on the simultaneous coverage event, an out-of-tolerance mean can never be contained. Finite source uncertainty enters through the contracted confirmation band B^conf_j = [u_{w,j}−ε_j, ℓ_{w,j}+ε_j], the intersection of all tolerance bands around plausible source moments; the expanded union band supports only comp
Load-bearing premise
The finite-source guarantee (δ+η) presupposes simultaneous source confidence intervals for self-normalized weighted source moments with coverage at least 1−η, and the paper does not supply a validated finite-sample construction for them; Section 7.5 explicitly admits that the normal-approximation intervals used in experiments are not validated, and reusing the same source data both to build w and to form the intervals can break the guarantee.
What would settle it
Run the finite-source experiment of Section 7.5 exactly (ns=200, true first-coordinate discrepancy 0.35, ε=0.25, nominal η=0.10) and record how often the normal-approximation intervals [ℓ_{w,j}, u_{w,j}] fail to contain the true self-normalized weighted source moment simultaneously for all five coordinates; if the failure rate exceeds η=0.10, the δ+η bound of Corollary 2 is not achieved for that construction, which the paper already concedes is unvalidated.
If this is right
- A practitioner can monitor a target stream at arbitrary stopping times and, on stopping, hold a formal ε-balance certificate for the prespecified functions and tolerances: the probability of ever falsely confirming an out-of-tolerance correction is at most δ.
- The certificate is local by design: it says nothing about balancing functions outside the prespecified class, so a rich enough class must be chosen in advance; a too-narrow class can pass even while the global monitor warns of harm.
- With finite source data, only contracted confirmation bands should be used for formal confirmation; the expanded compatibility band must be reported only as compatibility evidence, not as confirmation, because it can falsely confirm at a rate far above δ.
- In weighted conformal prediction, the residual score-CDF mismatch ε enters directly as a coverage loss of 1−α−γ−ε, so confirmed balance on score-relevant balancing functions is the right gate for downstream deployment.
- Confirmation is not a power guarantee: a within-tolerance correction may take arbitrarily long to confirm when tolerances are tight, variance is large, or the monitoring stream is short.
Where Pith is reading between the lines
- A natural extension the author leaves open is adaptive enrichment of the balancing-function class during monitoring; preserving anytime validity would require a predictable construction that adds functions without peeking at current target data, possibly via sample splitting or a mixture over an expanding family.
- The same simultaneous-confidence-sequence containment logic applies beyond covariate shift: any fixed correction or recalibration—label-shift weights, predictive tilts, drift corrections—can be monitored for equivalence with a target stream under the same δ false-confirmation guarantee.
- An empty confirmation band in finite-source settings is itself a useful, explicitly non-failure signal: it quantifies how much more source information, wider tolerance, or lower weight variability would be needed before formal confirmation becomes possible.
- One could test the practical value of the relative-evidence monitor as a screening tool: run it alongside a deliberately narrow balancing-function class and check whether sustained negative log-growth flags harmful correction directions that the local certificate misses; the paper's own experiments already point to such overlapping blind spots.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops an anytime-valid procedure to confirm that a fixed covariate-shift correction w makes the corrected source input distribution ε-close to the target distribution on a prespecified finite class of balancing functions, using time-uniform confidence sequences for target means (Theorem 2). It also proposes a source-calibrated likelihood-ratio e-process with a KL-drift identity (Theorem 1, Proposition 1), an exponential-tilt test for an acceptable correction region (Section 6.1), finite-source adjustments via contracted 'confirmation bands' (Corollary 2), and an application to weighted split conformal prediction (Proposition 3). Under the stated simultaneous-coverage assumption, Theorem 2 is correct: an out-of-tolerance target mean cannot be contained in a tolerance band by any interval that contains it. The paper's main weakness is that the finite-source guarantee is conditional on a source-interval construction that is assumed but never supplied or validated.
Significance. If the finite-source component is completed, this is a useful contribution to the anytime-valid inference and covariate-shift literature. The central observation—that containment of simultaneous confidence sequences in tolerance bands yields a level-δ false-confirmation guarantee without requiring the balance statistic itself to be an e-process—is clean and correct. The explicit Hoeffding and sub-Gaussian confidence-sequence radii in Section 5.3 are reproducible, and the KL-drift identity and acceptable-region e-process are correct. The experiments illustrate false-confirmation control and the locality of the certificate. However, the abstract's finite-source promise is not yet supported by an actual construction, and the experiments do not validate the source-side normal approximation; this is the main obstacle to accepting the paper in its current form.
major comments (2)
- [§5.4.2, Corollary 2, Algorithm 2] The finite-source guarantee is load-bearing for the abstract and for Algorithm 2's else-branch, but the required source-side interval construction is assumed rather than supplied. The quantity μw,j = Es[\tilde w f_j]/Es[\tilde w] is a self-normalized ratio, so standard Hoeffding or sub-Gaussian bounds for a single expectation do not directly yield simultaneous finite-sample intervals for this ratio. The text only says 'Suppose a source split independent of the data used to construct the correction, or another conditionally valid construction' and gives no theorem or algorithm producing intervals with coverage ≥1−η. Section 7.5 explicitly states that the normal-approximation intervals used in the experiments are not validated, and the Table 9 caption says the experiment 'does not itself validate the normal approximation.' Consequently, the δ+η false-confirmation bound is not demonstrated
- [§3 and §5.4.2 (data allocation)] Corollary 2 requires the source intervals to be independent of the data used to construct w, but the paper does not specify how a single source dataset Dsrc is split among correction construction, normalization, and source-moment estimation. If the same source data are reused to form w and to build the intervals, the coverage event Es in the proof of Corollary 2 need not hold. This is not a technicality; the main use case described in Section 3 is a correction obtained from source data. The procedure should state explicitly how the split is made and what coverage level is guaranteed after that allocation.
minor comments (4)
- [Algorithm 2] Lines 15–16 update the joint confidence sequence inside the loop over j. Since C_{n,1},...,C_{n,m} are updated jointly, the pseudo-code should move the update outside the j-loop or otherwise clarify the intended ordering.
- [References] Wellek (2010) title contains a typo: 'Noninferiorit' should be 'Noninferiority'.
- [§4] Duplicate word in the sentence 'We do not turn decay into a formal formal refutation procedure here'; 'formal' appears twice.
- [§7.5 / Table 9] The caption is appropriately cautious, but the main text should more prominently warn that Table 9 does not provide empirical support for the δ+η guarantee. In its current placement, readers may take the normal-approximation source intervals as validated.
Circularity Check
No significant circularity: the confirmation guarantee is a direct corollary of an externally supplied confidence-sequence coverage property; the only self-citations are non-load-bearing, and the finite-source gap is an acknowledged assumption rather than a circular reduction.
full rationale
The paper's central claim (Theorem 2) is not circular: it assumes the simultaneous target confidence-sequence event (12), P{∀n≥1, ∀j≤m: μ_t,j ∈ C_{n,j}} ≥ 1−δ, and then proves that if |μ_t,j−μ_w,j|>ε_j for some j, P{τ_bal<∞}≤δ. The proof is a direct logical consequence: on the coverage event, any interval containing the true mean cannot be contained in the tolerance band, so stopping cannot occur. This is a standard confidence-sequence containment rule, and it does not define balance in terms of the intervals or fit any parameter to the quantity being predicted. Similarly, the global e-process rests on the definition E_s[w(X)]=1 and Ville's inequality; Proposition 1's KL drift identity is an exact Radon–Nikodym chain-rule identity, so no circularity arises there. The finite-source extension (Corollary 2) is conditional on an explicitly stated assumption: 'a source split independent of the data used to construct the correction, or another conditionally valid construction' delivering simultaneous source intervals with coverage ≥1−η. The theorem's δ+η bound follows algebraically from that assumption and the target coverage event; it is not obtained by renaming a fitted quantity. The manuscript itself flags the limitation that the normal-approximation source intervals used in experiments are not validated (§7.5) and that the experiment 'illustrates the band geometry and does not itself validate the normal approximation' (Table 9 caption). That is an acknowledged unverified assumption/correctness risk, not a circular reduction. The only self-citation, Choi (2026), appears as a conceptual predecessor for label-shift confirmation and as a contrast for pure covariate shift; none of the covariate-balance proofs rely on it. Hence the derivation is self-contained against an external benchmark, and there is no fitted input called prediction or self-citation chain that forces the result.
Axiom & Free-Parameter Ledger
free parameters (3)
- Tolerance vector ε_j =
0.20-0.30 in experiments; user-specified otherwise
- Test direction λ and acceptable radius κ =
λ=±0.4, κ=0.6 in §7.3
- Confidence levels δ, η, α =
δ=0.05, η=0.10, α=0.05
axioms (7)
- domain assumption Pure covariate shift: P_t(Y|X)=P_s(Y|X)
- domain assumption w is fixed, G0-measurable, and normalized with Es[w(X)]=1
- domain assumption Target inputs are iid from P_t^X, or a valid time-uniform confidence sequence exists
- domain assumption A source split independent of the data used to construct w yields simultaneous source intervals with coverage ≥1−η
- domain assumption Balancing functions are bounded a.s., or sub-Gaussian with known variance proxy for the explicit CS constructions
- domain assumption Acceptable-region test: Θ0 and λ are prespecified before monitoring and ψ_Θ0(λ)<∞
- domain assumption Score-relevant functions f_q are contained in or approximated by F for downstream conformal coverage
read the original abstract
Many covariate-shift adaptation methods construct a correction $w(x)$, but users must still determine whether the corrected distribution is sufficiently balanced for the target stream. We study anytime-valid confirmation of prespecified corrections from sequential unlabeled target inputs. Our primary contribution is a procedure for confirming covariate balance. For a prespecified class of balancing functions and tolerances, time-uniform confidence sequences permit continuous monitoring and data-dependent stopping once all plausible target moments lie within their tolerance bands. If the correction is out of tolerance for at least one function, the probability of ever incorrectly confirming balance is at most the prescribed level. Upon stopping, the procedure yields a certificate local to the chosen functions and tolerances, yet providing an absolute downstream-adequacy statement that ordinary shift diagnostics generally do not. With finite source data, contracted bands preserve this guarantee while accounting for uncertainty in weighted source moments, whereas expanded bands support only compatibility diagnostics. As complementary information, we study a source-calibrated likelihood-ratio e-process whose KL-drift identity characterizes correction directions relative to the source. Under the source-reference distribution, the probability of ever crossing its evidence threshold is controlled, but crossing does not confirm balance. We also give an exponential-tilt test for departures beyond an acceptable correction region and deploy balance-confirmed corrections in weighted conformal prediction. Experiments illustrate false-confirmation control, locality to the balancing-function class, KL-drift diagnostics, acceptable-region monitoring, finite-source effects, and downstream conformal coverage under covariate shift.
Figures
Reference graph
Works this paper leans on
-
[1]
Angelopoulos, A. N. and Bates, S. (2023). Conformal prediction: A gentle introduction.Foundations and Trends® in Machine Learning, 16(4):494–591
2023
-
[2]
W., and Wager, S
Athey, S., Imbens, G. W., and Wager, S. (2018). Approximate residual balancing: Debiased inference of average treatment effects in high dimensions.Journal of the Royal Statistical Society Series B, 80(4):597–623
2018
-
[3]
F., Cand` es, E
Barber, R. F., Cand` es, E. J., Ramdas, A., and Tibshirani, R. J. (2023). Conformal prediction beyond exchangeability.The Annals of Statistics, 51(2):816–845
2023
-
[4]
Ben-Michael, E., Feller, A., Hirshberg, D. A., and Zubizarreta, J. R. (2021). The balancing act in causal inference. arXiv:2110.14831
Pith/arXiv arXiv 2021
-
[5]
Chan, K. C. G., Yam, S. C. P., and Zhang, Z. (2016). Globally efficient non-parametric inference of average treatment effects by empirical balancing calibration weighting.Journal of the Royal Statistical Society Series B, 78(3):673–700
2016
-
[6]
Choi, S. (2026). Anytime-valid confirmation of label-shift corrections. InICML 2026 Workshop on Hypothesis Testing
2026
-
[7]
Cortes, C., Mansour, Y., and Mohri, M. (2010). Learning bounds for importance weighting. InAdvances in Neural Information Processing Systems (NeurIPS)
2010
-
[8]
M., Rasch, M
Gretton, A., Borgwardt, K. M., Rasch, M. J., and Sch¨ olkopf, B. (2012). A kernel two-sample test.Journal of Machine Learning Research, 13:723–773. Gr¨ unwald, P., de Heide, R., and Koolen, W. (2024). Safe testing.Journal of the Royal Statistical Society Series B, 86(5):1091–1128
2012
-
[9]
Hainmueller, J. (2012). Entropy balancing for causal effects: A multivariate reweighting method to produce balanced samples in observational studies.Political Analysis, 20(1):25–46
2012
-
[10]
R., Ramdas, A., McAuliffe, J., and Sekhon, J
Howard, S. R., Ramdas, A., McAuliffe, J., and Sekhon, J. (2021). Time-uniform, nonparametric, nonasymptotic confidence sequences.The Annals of Statistics, 49(2):1055–1080
2021
-
[11]
and Ratkovic, M
Imai, K. and Ratkovic, M. (2014). Covariate balancing propensity score.Journal of the Royal Statistical Society Series B, 76(1):243–263
2014
-
[12]
Kanamori, T., Hido, S., and Sugiyama, M. (2009). A least-squares approach to direct importance estimation.Journal of Machine Learning Research, 10:1391–1445
2009
-
[13]
and Ramdas, A
Manole, T. and Ramdas, A. (2023). Martingale methods for sequential estimation of convex functionals and divergences.IEEE Transactions on Information Theory, 69(7):4641–4658. M¨ uller, A. (1997). Integral probability metrics and their generating classes of functions.Advances in Applied Probability, 29(2):429–443
2023
-
[14]
Papadopoulos, H., Proedrou, K., Vovk, V., and Gammerman, A. (2002). Inductive confidence machines for regression. InProceedings of the European Conference on Machine Learning (ECML)
2002
-
[15]
Ramdas, A., Gr¨ unwald, P., Vovk, V., and Shafer, G. (2023). Game-theoretic statistics and safe anytime- valid inference.Statistical Science, 38(4):576–601. 28 Anytime-Valid Confirmation of Covariate Balance for Prespecified Corrections
2023
-
[16]
Schuirmann, D. J. (1987). A comparison of the two one-sided tests procedure and the power approach for assessing the equivalence of average bioavailability.Journal of Pharmacokinetics and Biopharmaceutics, 15:657–680
1987
-
[17]
and Vovk, V
Shafer, G. and Vovk, V. (2008). A tutorial on conformal prediction.Journal of Machine Learning Research, 9:371–421
2008
-
[18]
and Ramdas, A
Shekhar, S. and Ramdas, A. (2024). Nonparametric two-sample testing by betting.IEEE Transactions on Information Theory, 70(2):1178–1203
2024
-
[19]
Shimodaira, H. (2000). Improving predictive inference under covariate shift by weighting the log-likelihood function.Journal of Statistical Planning and Inference, 90:227–244
2000
-
[20]
K., Fukumizu, K., Gretton, A., Sch¨ olkopf, B., and Lanckriet, G
Sriperumbudur, B. K., Fukumizu, K., Gretton, A., Sch¨ olkopf, B., and Lanckriet, G. R. G. (2012). On the empirical estimation of integral probability metrics.Electronic Journal of Statistics, 6:1550–1599
2012
-
[21]
Sugiyama, M., Nakajima, S., Kashima, H., von B¨ unau, P., and Kawanabe, M. (2007). Direct importance estimation with model selection and its application to covariate shift adaptation. InAdvances in Neural Information Processing Systems (NeurIPS)
2007
-
[22]
(2012).Density Ratio Estimation in Machine Learning
Sugiyama, M., Suzuki, T., and Kanamori, T. (2012).Density Ratio Estimation in Machine Learning. Cambridge University Press
2012
-
[23]
J., Barber, R
Tibshirani, R. J., Barber, R. F., Cand` es, E. J., and Ramdas, A. (2019). Conformal prediction under covariate shift. InAdvances in Neural Information Processing Systems (NeurIPS)
2019
-
[24]
Ville, J. (1939). ´Etude Critique de la Notion de Collectif. PhD thesis, Universit´ e de Paris
1939
-
[25]
(2005).Algorithmic Learning in a Random World
Vovk, V., Gammerman, A., and Shafer, G. (2005).Algorithmic Learning in a Random World. Springer
2005
-
[26]
Wald, A. (1945). Sequential tests of statistical hypotheses.The Annals of Mathematical Statistics, 16(2):117–186
1945
-
[27]
and Ramdas, A
Waudby-Smith, I. and Ramdas, A. (2024). Estimating means of bounded random variables by betting. Journal of the Royal Statistical Society Series B, 86:1–27
2024
-
[28]
(2010).Testing Statistical Hypotheses of Equivalence and Noninferiorit
Wellek, S. (2010).Testing Statistical Hypotheses of Equivalence and Noninferiorit. Chapman & Hall/CRC, 2nd edition
2010
-
[29]
Zubizarreta, J. R. (2015). Stable weights that balance covariates for estimation with incomplete outcome data.Journal of the American Statistical Association, 110(511):910–922. 29 Anytime-Valid Confirmation of Covariate Balance for Prespecified Corrections A Proofs A.1 Proof of Theorem 1 Proof. Fix a realization of G0. By Assumption 1, w is then fixed and...
2015
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.