REVIEW 3 major objections 5 minor 1 cited by
Selective Inference for Time-Varying Effect Moderation
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read After using a randomized lasso to choose moderators, conditioning on a slice of the selection event yields uniformly asymptotically valid confidence intervals for time-varying causal effect moderation, with bounded length in settings…
desk verdict A serious, technically strong extension of randomized selective inference to time-varying causal moderation, but the main coverage theorem is stated for population H and K while the implemented procedure must estimate them—that gap needs fixing before the claimed guarantee covers the algorithm. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the pivot $$$P^{{E\cdot j}}$(\hat $b^{{E\cdot j}}$_n,\hat $g^{{E\cdot j}}$_n;\$beta^{{E\cdot j}}$_n)= \frac{\int_{-\infty}^{\sqrt{n}\hat $b^{{E\cdot j}}$_n}\$\varphi$(x;\sqrt{n}\$beta^{{E\cdot j}}$_n,(\$sigma^{{E\cdot j}}$)^2)F(x,\sqrt{n}\hat $g^{{E\cdot j}}$_n)\,dx}{\int_{-\infty}^{\infty}\$\varphi$(x;\sqrt{n}\$beta^{{E\cdot j}}$_n,(\$sigma^{{E\cdot j}}$)^2)F(x,\sqrt{n}\hat $g^{{E\cdot j}}$_n)\,dx},$$ where $F$ integrates the Gaussian randomization density over the truncation interval $[I^{E\cdot j}_-,I^{E\cdot j}_+]$ determined by the conditioning event. The randomized lasso's K.K.T. stationarity conditions are what connect selection to the master statistics; the independent Gaussian noise makes the change of variables from the randomization to the selection variables exact in the Gaussian case, and Assumption 3 extends that Gaussian weight to general distributions. Inverting the pivot gives the selective confidence intervals.
What would settle it
Take the low-signal, Laplace-error simulation setting of Section 6 ($n=120$, $T=30$, $p=50$) and repeat the 500-replicate coverage study, but draw the randomization $\sqrt{n}\omega_n$ from a heavy-tailed $t$ distribution with the same covariance instead of a Gaussian. If the empirical conditional coverage of the 90% intervals falls clearly below 0.9, then Assumption 3's uniform closeness of the randomization density fails in exactly the regime the simulations showcase, and Theorem 4.2's guarantee is not what produces the reported coverage.
Extended reading notes
Core claim
The central discovery is that the randomized-lasso selection event can be sliced into a one-dimensional truncation region plus a lower-dimensional conditioning statistic, and that slicing is enough for valid inference. Using the K.K.T. conditions of the randomized lasso, the paper rewrites selection in terms of master statistics; conditioning on $\hat S_n=S$ and $\hat V^{E\cdot j}_n=V^{E\cdot j}$ truncates the statistic $\hat U^{E\cdot j}_n$ to an interval $[I^{E\cdot j}_-,I^{E\cdot j}_+]$. After further conditioning on the nuisance statistic $\hat g^{E\cdot j}_n$, marginalizing the truncated Gaussian randomization density, and applying a probability integral transform, the pivot in (12) is exactly uniform in the Gaussian fixed-design case and asymptotically uniform in the semi-parametric case. Theorem 4.2 states that, under Assumptions 2, 3, and 4, the selective interval $(L^{E\cdot j}_{n,\alpha}, U^{E\cdot j}_{n,\alpha})$ satisfies $$\lim_{n\to\infty}\sup_{F_n\in\mathcal{F}_n} \left|P\left(\$beta^{{E\cdot j}}$_n\in ($L^{{E\cdot j}}$_{n,\$\alpha$}, $U^{{E\cdot j}}$_{n,\$\alpha$})\mid \hat S_n=S,\hat $V^{{E\cdot j}}$_n=$V^{{E\cdot j}}$\right)-(1-\$\alpha$)\right|=0.$$ Because the conditioning event is a strict subset of the full selection event $\{\hat E=E\}$, the tower property of expectation transfers the conditional guarantee to the unconditional coverage and to false-coverage-rate control.
Load-bearing premise
The coverage guarantee rests on Assumption 3, that the density of the perturbed randomization variable $\sqrt{n}\tilde{\omega}_n$ is uniformly close to the Gaussian density inside the pivot's selection-probability integral; if that approximation fails, the pivot is no longer uniform and the claimed $(1-\alpha)$ conditional coverage can break down, and the theorem is also stated for fixed $p$ with growing $n$, so the high-dimensional settings mentioned in the abstract are not covered by the proof.
Editorial extensions
If this is right
- Practitioners can use the full dataset for both moderator selection and inference, so intervals reflect the strength of the observed signal instead of inheriting the fixed width of a split sample.
- Coverage remains at the nominal level under non-Gaussian, autocorrelated errors and low signal strengths, the regime where the polyhedral method undercovers and commonly produces infinitely long intervals.
- The uniform guarantee holds for parameter sequences growing as $r_n=o(n^{1/6})$, so the conditional coverage statement covers more than local alternatives fixed as $n$ grows.
- The same pivot construction extends to other causal contrasts with a linear loss formulation, including relative-risk excursion effects for binary outcomes, via the loss-framework reformulation in the appendix.
Reading between the lines
- A natural next test, not performed in the paper, is to push the method into $p>n$ regimes: the theorem is proved for fixed $p$ with growing $n$, while the motivating language describes high-dimensional moderation, so coverage in that regime is an open question that simulations with $n=120$, $p=50$ do not settle.
- The same conditioning construction suggests a tuning principle: the analyst can choose how much randomization variance $\tau^2$ to inject, trading selection accuracy against interval length; the paper's $\Omega=\tau^2 I_p$ choice is one point on that frontier, and data-dependent choices could shorten intervals further.
- If the uniform-closeness assumption on the perturbation density can be replaced by a coupling or Edgeworth-type correction, the pivot approach could extend to discrete or non-Gaussian randomization schemes without losing the bounded-interval property.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a two-step selective inference procedure for time-varying causal effect moderation in mobile-health settings. In the first step, a Gaussian-randomized lasso is applied to a weighted centered least-squares (WCLS) criterion to select a sparse set of effect moderators. In the second step, the authors construct a pivot for each selected coefficient by conditioning on a subset of the selection event, marginalizing over the randomization, and applying a probability integral transform. The main theoretical result, Theorem 4.2, claims that the resulting selective confidence intervals achieve uniform asymptotic conditional coverage over a class of data-generating distributions satisfying Assumptions 2–4. The paper also reports simulations under Gaussian, Laplace, and exponential errors showing that the method attains nominal false coverage rates with shorter bounded intervals than the polyhedral approach of Zhao et al. (2021), and it applies the method to data from the VALENTINE mHealth study.
Significance. If the main theorem is correct, the paper offers a genuinely useful extension of Gaussian-randomization selective inference from fixed-X Gaussian regression to semiparametric WCLS/EMEE-type estimators for causal excursion effects. The pivot construction is carefully motivated and the proofs, especially the Stein-based uniform asymptotic arguments in the appendix, are nontrivial. The empirical comparison is informative and suggests the method can improve on both data splitting and polyhedral selective inference in low-signal, heavy-tailed settings. The claim of bounded intervals is a practical advantage. However, the significance is tempered by a gap between the theoretical object for which coverage is proved (with population H and K) and the procedure as implemented (which must estimate these matrices), and by the fixed-p nature of the asymptotic guarantee relative to the high-dimensional motivation.
major comments (3)
- [§4.2, §4.4, §5.2] The main theorem does not cover the procedure as implemented because H and K are treated as known population matrices. Section 4.2 defines H and K as population expectations, and the matrices P_1, P_2, Λ, Q, η, the truncation interval in Proposition 4.4, and the pivot F in Eq. (11) all depend on H and K. Theorem 4.2 and the proofs in Theorem 11.2 and Lemma 11.7 treat these matrices as fixed. In the WCLS and EMEE settings of Section 5.1, H and K contain population expectations over the unknown data-generating distribution and must be estimated from data. Section 5.2 discusses nuisance function estimation only, not estimation of H and K. If plug-in estimates are used, the joint distribution of the master statistics and the change-of-variables density analyzed in Theorem 4.1 are no longer the ones analyzed, and no uniform perturbation bound is supplied. The coverage guarantee therefore applies to an oracle version of the method, not to the algorithm used in the simulations and data analysis.
- [§4.2, §6, abstract] The asymptotic theory is stated in the fixed-p, growing-n regime, but the paper frames the problem as high-dimensional. Section 4.2 explicitly says the master statistics have an asymptotic normal distribution 'in the fixed p and growing n regime,' and Theorem 4.2's uniformity is over distributions F_n with p fixed. The abstract and Section 2.3 motivate the method by high-dimensional moderation analysis, and the simulations use n=120, p=50, which is not a regime covered by the theorem. Unless the authors either restrict their claims to fixed p or extend the theory to p=p_n growing with n under explicit conditions, the central 'high-dimensional' claim is not supported by the stated uniform asymptotic guarantee.
- [§4.5, Assumption 3] Assumption 3 is load-bearing but its verification is asserted rather than proved. The text states that since √n(\tilde ω_n − ω_n)=o_p(1), the condition is 'automatically satisfied' when r_n does not grow with n. This does not follow: o_p(1) does not imply the existence of a Lebesgue density q_n for \tilde ω_n, nor the uniform sup-norm closeness of the ratio G_n/F that Assumption 3 requires. The proof of Lemma 11.6 uses Assumption 3 to replace G_n by F, so if this assumption fails, the pivot's uniformity and hence the coverage guarantee break down. The authors should either prove that their leading examples satisfy Assumption 3 under explicit conditions or state it as a high-level primitive and verify it in the WCLS and EMEE settings.
minor comments (5)
- [§6.1] The text says the intervals aim to achieve an FCR of '0.1%' but the nominal level is 0.10 (10%); please correct the typo.
- [§4.4] The definition of T^{E·j} contains a notational typo: 'H,EV^{E·j}' should presumably be 'H_E V^{E·j}'. Please clarify.
- [§4.5, Assumption 3] In Assumption 3, q_n is written as q_n(·; 0_p, Ω), but q_n is a general Lebesgue density and is not necessarily Gaussian; the second argument is unexplained and should be removed or redefined.
- [§7.3] There is an incomplete sentence: 'This estimator is then used as a plug-in value in the penalized estimation to identify a more focused subset of important moderators. to the remaining 70% of the data, where we' — the second half appears to be a fragment and should be completed.
- [Theorem 4.2] The notation F_n is used both for a single data-generating distribution and for the collection of distributions, which is confusing; please use a different symbol for the collection, e.g., \mathcal{F}_n.
Circularity Check
No material circularity: the pivot is a probability integral transform of a derived conditional density, and the asymptotic coverage theorem is proven from stated distributional assumptions rather than fitted to the target intervals.
full rationale
The central claim is that the pivot in (12) is asymptotically uniform conditional on the selection event, yielding the selective confidence intervals in Theorem 4.2. The pivot is constructed by a probability integral transform of a conditional density obtained from the joint distribution of the master statistics and the randomized lasso variables; no free parameter is fitted to force coverage. The proof proceeds through relative-difference bounds, a leave-one-out Stein bound (Lemma 11.7, cited from Panigrahi 2023), and explicit regularity conditions (Assumptions 2–4). The Stein-bound citation is a general technical lemma from prior work by an overlapping author, but it does not assert the target coverage result and is used as a supporting estimate; it is not a definitional reduction. The exact-Gaussian pivot is attributed to Panigrahi et al. (2023a), but the paper independently derives the conditional density and pivot before extending asymptotically. The main theorem is stated for population H and K while the implementation estimates them; this is a gap between the theorem's hypotheses and the algorithm, not a circular derivation. No equation is equivalent to its inputs by construction, and no fitted parameter is renamed as a prediction.
Assumptions & free parameters
free parameters (3)
- lambda (lasso penalty) =
not specified
- tau^2 (randomization variance) =
not specified
- nuisance estimation split ratio =
30% / 70% in data application
assumptions (6)
- domain assumption Consistency, positivity, sequential ignorability (Assumption 1)
- domain assumption WCLS provides a consistent estimator and Neyman orthogonality holds (Shi and Dempsey, 2023)
- ad hoc to paper Assumption 2: sub-Gaussian score variables
- ad hoc to paper Assumption 3: density q_n of the perturbed randomization is uniformly close to F
- ad hoc to paper Assumption 4: tail bound on R_{n,1}
- domain assumption Fixed p, n growing regime
Cite this review
Pith. "Pith review of Selective Inference for Time-Varying Effect Moderation." pith.science (2026). https://pith.science/paper/2JNGOBLO
@misc{pith2026241115908,
author = {Pith},
title = {Pith review of: Selective Inference for Time-Varying Effect Moderation},
year = {2026},
howpublished = {\url{https://pith.science/paper/2JNGOBLO}},
note = {Machine review of arXiv:2411.15908}
}
read the original abstract
Causal effect moderation investigates how the effect of interventions (or treatments) on outcome variables changes based on observed characteristics of individuals, known as potential effect moderators. With advances in data collection, datasets containing many observed features as potential moderators have become increasingly common. High-dimensional analyses often lack interpretability, with important moderators masked by noise, while low-dimensional, marginal analyses yield many false positives due to strong correlations with true moderators. In this paper, we propose a two-step method for selective inference on time-varying causal effect moderation that addresses the limitations of both high-dimensional and marginal analyses. Our method first selects a relatively smaller, more interpretable model to estimate a linear causal effect moderation using a Gaussian randomization approach. We then condition on the selection event to construct a pivot, enabling uniformly asymptotic semi-parametric inference in the selected model. Through simulations and real data analyses, we show that our method consistently achieves valid coverage rates, even when existing conditional methods and common sample splitting techniques fail. Moreover, our method yields shorter, bounded intervals, unlike existing methods that may produce infinitely long intervals.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 1 Pith paper
-
Flexible Inference for Winners with Conditional Validity
A data-adaptive exponential randomization scheme yields conditionally valid confidence intervals for top-k winners, with selection quality close to standard top-k and shorter intervals than polyhedral methods.
Reference graph
Works this paper leans on
-
[1]
Benjamini, Y. and Yekutieli, D. (2005). False discovery rate--adjusted multiple confidence intervals for selected parameters. Journal of the American Statistical Association , 100(469):71--81
work page 2005
-
[2]
Boruvka, A., Almirall, D., Witkiewitz, K., and Murphy, S. (2018). Assessing time-varying causal effect moderation in mobile health. J Am Stat Assoc. , 113(523):1112-1121
work page 2018
-
[3]
Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., and Robins, J. (2018). Double/debiased machine learning for treatment and structural parameters . The Econometrics Journal , 21(1):C1--C68
2018
-
[4]
Dempsey, W., Liao, P., Kumar, S., and Murphy, S. A. (2020). The stratified micro-randomized trial design: sample size considerations for testing nested causal effects of time-varying treatments. The annals of applied statistics , 14(2):661
work page 2020
-
[5]
R., Gupta, K., Stevens, R., Jeganathan, V
Golbus, J. R., Gupta, K., Stevens, R., Jeganathan, V. S. E., Luff, E., Shi, J., Dempsey, W., Boyden, T., Mukherjee, B., Kohnstamm, S., Taralunga, V., Kheterpal, V., Murphy, S., Klasnja, P., Kheterpal, S., and Nallamothu, B. K. (2023). A randomized trial of a mobile health intervention to augment cardiac rehabilitation. npj Digital Medicine , 6(1):173
work page 2023
-
[6]
Huang, Y., Pirenne, S., Panigrahi, S., and Claeskens, G. (2023). Selective inference using randomized group lasso estimators for general models. arXiv preprint arXiv:2306.13829
arXiv 2023
-
[7]
D., Sun, D
Lee, J. D., Sun, D. L., Sun, Y., and Taylor, J. E. (2016). Exact post-selection inference, with application to the lasso. The Annals of Statistics , 44(3)
2016
-
[8]
Newey, W. K. and Robins, J. R. (2018). Cross-fitting and fast remainder rates for semiparametric estimation
work page 2018
Show all 22 references
-
[9]
Panigrahi, S. (2023). Carving model-free inference . The Annals of Statistics , 51(6):2318 -- 2341
2023
-
[10]
Panigrahi, S., Fry, K., and Taylor, J. (2023a). Exact selective inference with randomization
2023
-
[11]
W., and Kessler, D
Panigrahi, S., MacDonald, P. W., and Kessler, D. (2023b). Approximate post-selective inference for regression with the group lasso. Journal of machine learning research , 24(79):1--49
2023
-
[12]
Panigrahi, S., Mohammed, S., Rao, A., and Baladandayuthapani, V. (2023c). Integrative bayesian models using post-selective inference: A case study in radiogenomics. Biometrics , 79(3):1801--1813
2023
-
[13]
Qian, T., Yoo, H., Klasnja, P., Almirall, D., and Murphy, S. A. (2020). Estimating time-varying causal excursion effects in mobile health with binary outcomes . Biometrika , 108(3):507--527
2020
-
[14]
Robins, J. M. (1997). Causal inference from complex longitudinal data. In Berkane, M., editor, Latent Variable Modeling and Applications to Causality , pages 69--117, New York, NY. Springer New York
1997
-
[15]
Robinson, P. M. (1988). Root-n-consistent semiparametric regression. Econometrica , 56(4):931--954
1988
-
[16]
Rubin, D. B. (2005). Causal inference using potential outcomes: Design, modeling, decisions. Journal of the American Statistical Association , 100(469):322--331
2005
-
[17]
Schick, A. (1986). On asymptotically efficient estimation in semiparametric models. The Annals of Statistics , 14(3):1139--1151
1986
-
[18]
and Dempsey, W
Shi, J. and Dempsey, W. (2023). A meta-learning method for estimation of causal excursion effects to assess time-varying moderation
2023
-
[19]
Shi, J., Wu, Z., and Dempsey, W. (2022). Assessing time-varying causal effect moderation in the presence of cluster-level treatment effect heterogeneity and interference . Biometrika , 110(3):645--662
2022
-
[20]
J., Rinaldo, A., Tibshirani, R., and Wasserman, L
Tibshirani, R. J., Rinaldo, A., Tibshirani, R., and Wasserman, L. (2018). Uniform asymptotic inference and the bootstrap after model selection . The Annals of Statistics , 46(3):1255 -- 1287
2018
-
[21]
van der Laan, L., Carone, M., and Luedtke, A. (2024). Combining t-learning and dr-learning: a framework for oracle-efficient estimation of causal contrasts
2024
-
[22]
S., and Ertefaie, A
Zhao, Q., Small, D. S., and Ertefaie, A. (2021). Selective inference for effect modification via the lasso
2021
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.