Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Selective Inference for Time-Varying Effect Moderation

T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read After using a randomized lasso to choose moderators, conditioning on a slice of the selection event yields uniformly asymptotically valid confidence intervals for time-varying causal effect moderation, with bounded length in settings…

desk verdict A serious, technically strong extension of randomized selective inference to time-varying causal moderation, but the main coverage theorem is stated for population H and K while the implemented procedure must estimate them—that gap needs fixing before the claimed guarantee covers the algorithm. read the letter →

arxiv 2411.15908 v1 pith:2JNGOBLO submitted 2024-11-24 stat.ME math.STstat.MLstat.TH

classification stat.MEmath.STstat.MLstat.TH MSC 62F2562G2062J07
keywords effectmoderationselectiveinferencerandomizationsemiparametrictime-varyingcausaleffectsrandomizedlassoconfidenceintervalsmicro-randomizedtrials
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Choosing which features moderate a time-varying treatment effect creates a statistical bind: a high-dimensional analysis of all candidate moderators is hard to interpret and masks the true moderators, while separate marginal analyses produce many false positives from correlated features. This paper proposes a two-step solution for causal effect moderation. A lasso with added Gaussian noise selects a smaller working model, and inference then conditions on a carefully chosen part of the selection event through a truncated-normal pivot, using all of the data rather than discarding a split-half. The paper's main theorem gives uniformly asymptotic validity of the resulting confidence intervals across a broad class of non-Gaussian distributions, and simulations show nominal coverage with bounded, signal-adaptive intervals where polyhedral selective inference undercovers or returns infinitely long intervals and data splitting undercovers at low signal.

What carries the argument

The load-bearing object is the pivot $$$P^{{E\cdot j}}$(\hat $b^{{E\cdot j}}$_n,\hat $g^{{E\cdot j}}$_n;\$beta^{{E\cdot j}}$_n)= \frac{\int_{-\infty}^{\sqrt{n}\hat $b^{{E\cdot j}}$_n}\$\varphi$(x;\sqrt{n}\$beta^{{E\cdot j}}$_n,(\$sigma^{{E\cdot j}}$)^2)F(x,\sqrt{n}\hat $g^{{E\cdot j}}$_n)\,dx}{\int_{-\infty}^{\infty}\$\varphi$(x;\sqrt{n}\$beta^{{E\cdot j}}$_n,(\$sigma^{{E\cdot j}}$)^2)F(x,\sqrt{n}\hat $g^{{E\cdot j}}$_n)\,dx},$$ where $F$ integrates the Gaussian randomization density over the truncation interval $[I^{E\cdot j}_-,I^{E\cdot j}_+]$ determined by the conditioning event. The randomized lasso's K.K.T. stationarity conditions are what connect selection to the master statistics; the independent Gaussian noise makes the change of variables from the randomization to the selection variables exact in the Gaussian case, and Assumption 3 extends that Gaussian weight to general distributions. Inverting the pivot gives the selective confidence intervals.

What would settle it

Take the low-signal, Laplace-error simulation setting of Section 6 ($n=120$, $T=30$, $p=50$) and repeat the 500-replicate coverage study, but draw the randomization $\sqrt{n}\omega_n$ from a heavy-tailed $t$ distribution with the same covariance instead of a Gaussian. If the empirical conditional coverage of the 90% intervals falls clearly below 0.9, then Assumption 3's uniform closeness of the randomization density fails in exactly the regime the simulations showcase, and Theorem 4.2's guarantee is not what produces the reported coverage.

Watch

Extended reading notes

Core claim

The central discovery is that the randomized-lasso selection event can be sliced into a one-dimensional truncation region plus a lower-dimensional conditioning statistic, and that slicing is enough for valid inference. Using the K.K.T. conditions of the randomized lasso, the paper rewrites selection in terms of master statistics; conditioning on $\hat S_n=S$ and $\hat V^{E\cdot j}_n=V^{E\cdot j}$ truncates the statistic $\hat U^{E\cdot j}_n$ to an interval $[I^{E\cdot j}_-,I^{E\cdot j}_+]$. After further conditioning on the nuisance statistic $\hat g^{E\cdot j}_n$, marginalizing the truncated Gaussian randomization density, and applying a probability integral transform, the pivot in (12) is exactly uniform in the Gaussian fixed-design case and asymptotically uniform in the semi-parametric case. Theorem 4.2 states that, under Assumptions 2, 3, and 4, the selective interval $(L^{E\cdot j}_{n,\alpha}, U^{E\cdot j}_{n,\alpha})$ satisfies $$\lim_{n\to\infty}\sup_{F_n\in\mathcal{F}_n} \left|P\left(\$beta^{{E\cdot j}}$_n\in ($L^{{E\cdot j}}$_{n,\$\alpha$}, $U^{{E\cdot j}}$_{n,\$\alpha$})\mid \hat S_n=S,\hat $V^{{E\cdot j}}$_n=$V^{{E\cdot j}}$\right)-(1-\$\alpha$)\right|=0.$$ Because the conditioning event is a strict subset of the full selection event $\{\hat E=E\}$, the tower property of expectation transfers the conditional guarantee to the unconditional coverage and to false-coverage-rate control.

Load-bearing premise

The coverage guarantee rests on Assumption 3, that the density of the perturbed randomization variable $\sqrt{n}\tilde{\omega}_n$ is uniformly close to the Gaussian density inside the pivot's selection-probability integral; if that approximation fails, the pivot is no longer uniform and the claimed $(1-\alpha)$ conditional coverage can break down, and the theorem is also stated for fixed $p$ with growing $n$, so the high-dimensional settings mentioned in the abstract are not covered by the proof.

Editorial extensions

If this is right

  • Practitioners can use the full dataset for both moderator selection and inference, so intervals reflect the strength of the observed signal instead of inheriting the fixed width of a split sample.
  • Coverage remains at the nominal level under non-Gaussian, autocorrelated errors and low signal strengths, the regime where the polyhedral method undercovers and commonly produces infinitely long intervals.
  • The uniform guarantee holds for parameter sequences growing as $r_n=o(n^{1/6})$, so the conditional coverage statement covers more than local alternatives fixed as $n$ grows.
  • The same pivot construction extends to other causal contrasts with a linear loss formulation, including relative-risk excursion effects for binary outcomes, via the loss-framework reformulation in the appendix.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural next test, not performed in the paper, is to push the method into $p>n$ regimes: the theorem is proved for fixed $p$ with growing $n$, while the motivating language describes high-dimensional moderation, so coverage in that regime is an open question that simulations with $n=120$, $p=50$ do not settle.
  • The same conditioning construction suggests a tuning principle: the analyst can choose how much randomization variance $\tau^2$ to inject, trading selection accuracy against interval length; the paper's $\Omega=\tau^2 I_p$ choice is one point on that frontier, and data-dependent choices could shorten intervals further.
  • If the uniform-closeness assumption on the perturbation density can be replaced by a coupling or Edgeworth-type correction, the pivot approach could extend to discrete or non-Gaussian randomization schemes without losing the bounded-interval property.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper develops a two-step selective inference procedure for time-varying causal effect moderation in mobile-health settings. In the first step, a Gaussian-randomized lasso is applied to a weighted centered least-squares (WCLS) criterion to select a sparse set of effect moderators. In the second step, the authors construct a pivot for each selected coefficient by conditioning on a subset of the selection event, marginalizing over the randomization, and applying a probability integral transform. The main theoretical result, Theorem 4.2, claims that the resulting selective confidence intervals achieve uniform asymptotic conditional coverage over a class of data-generating distributions satisfying Assumptions 2–4. The paper also reports simulations under Gaussian, Laplace, and exponential errors showing that the method attains nominal false coverage rates with shorter bounded intervals than the polyhedral approach of Zhao et al. (2021), and it applies the method to data from the VALENTINE mHealth study.

Significance. If the main theorem is correct, the paper offers a genuinely useful extension of Gaussian-randomization selective inference from fixed-X Gaussian regression to semiparametric WCLS/EMEE-type estimators for causal excursion effects. The pivot construction is carefully motivated and the proofs, especially the Stein-based uniform asymptotic arguments in the appendix, are nontrivial. The empirical comparison is informative and suggests the method can improve on both data splitting and polyhedral selective inference in low-signal, heavy-tailed settings. The claim of bounded intervals is a practical advantage. However, the significance is tempered by a gap between the theoretical object for which coverage is proved (with population H and K) and the procedure as implemented (which must estimate these matrices), and by the fixed-p nature of the asymptotic guarantee relative to the high-dimensional motivation.

major comments (3)
  1. [§4.2, §4.4, §5.2] The main theorem does not cover the procedure as implemented because H and K are treated as known population matrices. Section 4.2 defines H and K as population expectations, and the matrices P_1, P_2, Λ, Q, η, the truncation interval in Proposition 4.4, and the pivot F in Eq. (11) all depend on H and K. Theorem 4.2 and the proofs in Theorem 11.2 and Lemma 11.7 treat these matrices as fixed. In the WCLS and EMEE settings of Section 5.1, H and K contain population expectations over the unknown data-generating distribution and must be estimated from data. Section 5.2 discusses nuisance function estimation only, not estimation of H and K. If plug-in estimates are used, the joint distribution of the master statistics and the change-of-variables density analyzed in Theorem 4.1 are no longer the ones analyzed, and no uniform perturbation bound is supplied. The coverage guarantee therefore applies to an oracle version of the method, not to the algorithm used in the simulations and data analysis.
  2. [§4.2, §6, abstract] The asymptotic theory is stated in the fixed-p, growing-n regime, but the paper frames the problem as high-dimensional. Section 4.2 explicitly says the master statistics have an asymptotic normal distribution 'in the fixed p and growing n regime,' and Theorem 4.2's uniformity is over distributions F_n with p fixed. The abstract and Section 2.3 motivate the method by high-dimensional moderation analysis, and the simulations use n=120, p=50, which is not a regime covered by the theorem. Unless the authors either restrict their claims to fixed p or extend the theory to p=p_n growing with n under explicit conditions, the central 'high-dimensional' claim is not supported by the stated uniform asymptotic guarantee.
  3. [§4.5, Assumption 3] Assumption 3 is load-bearing but its verification is asserted rather than proved. The text states that since √n(\tilde ω_n − ω_n)=o_p(1), the condition is 'automatically satisfied' when r_n does not grow with n. This does not follow: o_p(1) does not imply the existence of a Lebesgue density q_n for \tilde ω_n, nor the uniform sup-norm closeness of the ratio G_n/F that Assumption 3 requires. The proof of Lemma 11.6 uses Assumption 3 to replace G_n by F, so if this assumption fails, the pivot's uniformity and hence the coverage guarantee break down. The authors should either prove that their leading examples satisfy Assumption 3 under explicit conditions or state it as a high-level primitive and verify it in the WCLS and EMEE settings.
minor comments (5)
  1. [§6.1] The text says the intervals aim to achieve an FCR of '0.1%' but the nominal level is 0.10 (10%); please correct the typo.
  2. [§4.4] The definition of T^{E·j} contains a notational typo: 'H,EV^{E·j}' should presumably be 'H_E V^{E·j}'. Please clarify.
  3. [§4.5, Assumption 3] In Assumption 3, q_n is written as q_n(·; 0_p, Ω), but q_n is a general Lebesgue density and is not necessarily Gaussian; the second argument is unexplained and should be removed or redefined.
  4. [§7.3] There is an incomplete sentence: 'This estimator is then used as a plug-in value in the penalized estimation to identify a more focused subset of important moderators. to the remaining 70% of the data, where we' — the second half appears to be a fragment and should be completed.
  5. [Theorem 4.2] The notation F_n is used both for a single data-generating distribution and for the collection of distributions, which is confusing; please use a different symbol for the collection, e.g., \mathcal{F}_n.

Circularity Check

0 steps flagged · score 0.0 of 10

No material circularity: the pivot is a probability integral transform of a derived conditional density, and the asymptotic coverage theorem is proven from stated distributional assumptions rather than fitted to the target intervals.

full rationale

The central claim is that the pivot in (12) is asymptotically uniform conditional on the selection event, yielding the selective confidence intervals in Theorem 4.2. The pivot is constructed by a probability integral transform of a conditional density obtained from the joint distribution of the master statistics and the randomized lasso variables; no free parameter is fitted to force coverage. The proof proceeds through relative-difference bounds, a leave-one-out Stein bound (Lemma 11.7, cited from Panigrahi 2023), and explicit regularity conditions (Assumptions 2–4). The Stein-bound citation is a general technical lemma from prior work by an overlapping author, but it does not assert the target coverage result and is used as a supporting estimate; it is not a definitional reduction. The exact-Gaussian pivot is attributed to Panigrahi et al. (2023a), but the paper independently derives the conditional density and pivot before extending asymptotically. The main theorem is stated for population H and K while the implementation estimates them; this is a gap between the theorem's hypotheses and the algorithm, not a circular derivation. No equation is equivalent to its inputs by construction, and no fitted parameter is renamed as a prediction.

Assumptions & free parameters 3 free parameters · 6 assumptions · 0 invented entities

The central claim rests on standard causal assumptions, the WCLS consistency theory, and three ad hoc uniformity assumptions (2, 3, 4) that are not verified in the simulations or data application. The method also depends on analyst choices (lambda, tau^2, split ratio) that are not fully specified. There are no invented physical or conceptual entities beyond the statistical construction.

free parameters (3)
  • lambda (lasso penalty) = not specified
    The penalization parameter in the randomized lasso is chosen by data/tuning in the simulations and data application, but the paper does not specify how lambda is selected or its values.
  • tau^2 (randomization variance) = not specified
    The Gaussian randomization covariance Omega = tau^2 * I_p is a user-chosen lever affecting the selection versus inference trade-off, but the simulations do not state the value of tau^2.
  • nuisance estimation split ratio = 30% / 70% in data application
    The paper uses a random one-third sample to estimate nuisance parameters and the rest for selection/inference; this split ratio is an analyst choice that affects the results.
assumptions (6)
  • domain assumption Consistency, positivity, sequential ignorability (Assumption 1)
    Standard causal inference assumptions needed to express the causal excursion effect in terms of observed data. Cited in Section 2.1.
  • domain assumption WCLS provides a consistent estimator and Neyman orthogonality holds (Shi and Dempsey, 2023)
    The WCLS estimator is used as the post-selection estimator, and its asymptotic normality is a key input to Proposition 4.2.
  • ad hoc to paper Assumption 2: sub-Gaussian score variables
    The score X_i^T grad psi(X_i,E beta_E^n; Y_i) must be sub-Gaussian uniformly over the class F_n. This is stronger than standard CLT conditions and is not verified in the simulations.
  • ad hoc to paper Assumption 3: density q_n of the perturbed randomization is uniformly close to F
    This ensures the change-of-variables density can be replaced by the Gaussian form in the pivot. It is a high-level condition not directly checkable from the data.
  • ad hoc to paper Assumption 4: tail bound on R_{n,1}
    Need for the case r_n to infinity. Controls the remainder in the master statistics representation.
  • domain assumption Fixed p, n growing regime
    All asymptotic statements are for fixed p as n -> infinity, but the motivating examples are high-dimensional with p comparable to n. This regime mismatch is unstated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Selective Inference for Time-Varying Effect Moderation." pith.science (2026). https://pith.science/paper/2JNGOBLO

@misc{pith2026241115908,
  author       = {Pith},
  title        = {Pith review of: Selective Inference for Time-Varying Effect Moderation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2JNGOBLO}},
  note         = {Machine review of arXiv:2411.15908}
}
read the original abstract

Causal effect moderation investigates how the effect of interventions (or treatments) on outcome variables changes based on observed characteristics of individuals, known as potential effect moderators. With advances in data collection, datasets containing many observed features as potential moderators have become increasingly common. High-dimensional analyses often lack interpretability, with important moderators masked by noise, while low-dimensional, marginal analyses yield many false positives due to strong correlations with true moderators. In this paper, we propose a two-step method for selective inference on time-varying causal effect moderation that addresses the limitations of both high-dimensional and marginal analyses. Our method first selects a relatively smaller, more interpretable model to estimate a linear causal effect moderation using a Gaussian randomization approach. We then condition on the selection event to construct a pivot, enabling uniformly asymptotic semi-parametric inference in the selected model. Through simulations and real data analyses, we show that our method consistently achieves valid coverage rates, even when existing conditional methods and common sample splitting techniques fail. Moreover, our method yields shorter, bounded intervals, unlike existing methods that may produce infinitely long intervals.

Figures

Figures reproduced from arXiv: 2411.15908 by the authors.

Figure 1
Figure 1. Average Coverage (left), Average Length of Finite CIs (mid), % infinitely long [PITH_FULL_IMAGE:figures/full_fig_p029_1.png] view at source ↗
Figure 2
Figure 2. Average Coverage (left), Average Length of Finite CIs (mid), % infinitely long [PITH_FULL_IMAGE:figures/full_fig_p030_2.png] view at source ↗
Figure 3
Figure 3. Average Coverage (left), Average Length of Finite CIs (mid), % infinitely long [PITH_FULL_IMAGE:figures/full_fig_p031_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: All the three method selected the same 4 variables. The 90% confidence intervals [PITH_FULL_IMAGE:figures/full_fig_p034_4.png]
Figure 5
Figure 5. Figure 5: Post-selective inference of an MRT in cardiac rehab population reveals several [PITH_FULL_IMAGE:figures/full_fig_p074_5.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Flexible Inference for Winners with Conditional Validity

    stat.ME 2026-07 conditional novelty 6.0 of 10

    A data-adaptive exponential randomization scheme yields conditionally valid confidence intervals for top-k winners, with selection quality close to standard top-k and shorter intervals than polyhedral methods.

Reference graph

Works this paper leans on

22 extracted references · 17 canonical work pages · cited by 1 Pith paper

  1. [1]

    and Yekutieli, D

    Benjamini, Y. and Yekutieli, D. (2005). False discovery rate--adjusted multiple confidence intervals for selected parameters. Journal of the American Statistical Association , 100(469):71--81

  2. [2]

    Boruvka, A., Almirall, D., Witkiewitz, K., and Murphy, S. (2018). Assessing time-varying causal effect moderation in mobile health. J Am Stat Assoc. , 113(523):1112-1121

  3. [3]

    Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., and Robins, J. (2018). Double/debiased machine learning for treatment and structural parameters . The Econometrics Journal , 21(1):C1--C68

  4. [4]

    Dempsey, W., Liao, P., Kumar, S., and Murphy, S. A. (2020). The stratified micro-randomized trial design: sample size considerations for testing nested causal effects of time-varying treatments. The annals of applied statistics , 14(2):661

  5. [5]

    R., Gupta, K., Stevens, R., Jeganathan, V

    Golbus, J. R., Gupta, K., Stevens, R., Jeganathan, V. S. E., Luff, E., Shi, J., Dempsey, W., Boyden, T., Mukherjee, B., Kohnstamm, S., Taralunga, V., Kheterpal, V., Murphy, S., Klasnja, P., Kheterpal, S., and Nallamothu, B. K. (2023). A randomized trial of a mobile health intervention to augment cardiac rehabilitation. npj Digital Medicine , 6(1):173

  6. [6]

    Huang, Y., Pirenne, S., Panigrahi, S., and Claeskens, G. (2023). Selective inference using randomized group lasso estimators for general models. arXiv preprint arXiv:2306.13829

  7. [7]

    D., Sun, D

    Lee, J. D., Sun, D. L., Sun, Y., and Taylor, J. E. (2016). Exact post-selection inference, with application to the lasso. The Annals of Statistics , 44(3)

  8. [8]

    Newey, W. K. and Robins, J. R. (2018). Cross-fitting and fast remainder rates for semiparametric estimation

Show all 22 references
  1. [9]

    Panigrahi, S. (2023). Carving model-free inference . The Annals of Statistics , 51(6):2318 -- 2341

  2. [10]

    Panigrahi, S., Fry, K., and Taylor, J. (2023a). Exact selective inference with randomization

  3. [11]

    W., and Kessler, D

    Panigrahi, S., MacDonald, P. W., and Kessler, D. (2023b). Approximate post-selective inference for regression with the group lasso. Journal of machine learning research , 24(79):1--49

  4. [12]

    Panigrahi, S., Mohammed, S., Rao, A., and Baladandayuthapani, V. (2023c). Integrative bayesian models using post-selective inference: A case study in radiogenomics. Biometrics , 79(3):1801--1813

  5. [13]

    Qian, T., Yoo, H., Klasnja, P., Almirall, D., and Murphy, S. A. (2020). Estimating time-varying causal excursion effects in mobile health with binary outcomes . Biometrika , 108(3):507--527

  6. [14]

    Robins, J. M. (1997). Causal inference from complex longitudinal data. In Berkane, M., editor, Latent Variable Modeling and Applications to Causality , pages 69--117, New York, NY. Springer New York

  7. [15]

    Robinson, P. M. (1988). Root-n-consistent semiparametric regression. Econometrica , 56(4):931--954

  8. [16]

    Rubin, D. B. (2005). Causal inference using potential outcomes: Design, modeling, decisions. Journal of the American Statistical Association , 100(469):322--331

  9. [17]

    Schick, A. (1986). On asymptotically efficient estimation in semiparametric models. The Annals of Statistics , 14(3):1139--1151

  10. [18]

    and Dempsey, W

    Shi, J. and Dempsey, W. (2023). A meta-learning method for estimation of causal excursion effects to assess time-varying moderation

  11. [19]

    Shi, J., Wu, Z., and Dempsey, W. (2022). Assessing time-varying causal effect moderation in the presence of cluster-level treatment effect heterogeneity and interference . Biometrika , 110(3):645--662

  12. [20]

    J., Rinaldo, A., Tibshirani, R., and Wasserman, L

    Tibshirani, R. J., Rinaldo, A., Tibshirani, R., and Wasserman, L. (2018). Uniform asymptotic inference and the bootstrap after model selection . The Annals of Statistics , 46(3):1255 -- 1287

  13. [21]

    van der Laan, L., Carone, M., and Luedtke, A. (2024). Combining t-learning and dr-learning: a framework for oracle-efficient estimation of causal contrasts

  14. [22]

    S., and Ertefaie, A

    Zhao, Q., Small, D. S., and Ertefaie, A. (2021). Selective inference for effect modification via the lasso

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.