REVIEW 3 major objections 5 minor 38 references
Reluctant Interaction Inference after Additive Modeling
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper constructs valid p-values for data-adaptive hypotheses about linear interaction effects by conditioning on a randomized sparse additive model fit.
desk verdict A useful extension of randomized group-lasso selective inference to interaction testing, with a solid empirical case but an under-scrutinized Laplace approximation near small selected group norms. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine is external Gaussian randomization added to the group-lasso objective in (1). Because the randomization vector $\omega$ is independent of $y$, a change of variables from $\omega$ to the group-lasso estimators $(\hat\gamma, \hat U, \hat Z)$ turns the opaque event $\{\hat M = M\}$ into simple sign constraints $\hat\gamma \succ 0$ together with fixed unit-norm and subgradient constraints (Lemma 1). The key statistics $\hat S^{M}_{jk} = (\hat\theta^{M}_{jk}, (\hat\beta^{M}_{jk})^\top, (\hat A^{M}_{jk})^\top)^\top$ have a Gaussian law for fixed $M$, and Proposition 1 gives their density conditional on the selection event as a product of that Gaussian, the randomization density evaluated through a mapping $\Pi_S(\gamma,U,Z)$, and a Jacobian determinant. Theorem 1 marginalizes $\gamma$ out, yielding a selective log-likelihood whose normalizing constant $c(\theta^{M}_{jk}, \beta^{M}_{jk})$ is approximated by Laplace's method with a barrier function; Theorem 2 then solves a low-dimensional convex problem for the approximate MLE and observed Fisher information, from which Wald pivots $\Phi(z_{jk})$ are formed.
What would settle it
Generate many null datasets with $\theta^{M}_{jk}=0$ and a large selected main-effect set, compute the proposed Wald p-values, and compare the empirical CDF of the pivots with $\mathrm{Uniform}(0,1)$; in a small-$|M|$ design, also compute the normalizing constant exactly by numerical integration and substitute it for the Laplace approximation. If the approximate-pivot ECDF deviates beyond Monte Carlo tolerance while the exact version stays uniform, the Laplace step—not the conditioning argument—is what fails.
Extended reading notes
Core claim
On the paper's own terms, the central claim is that after solving the randomized group-lasso problem (1) to fit a SPAM and observing the selected main-effect set $M$, valid inference for the linear interaction coefficient $\theta^{M}_{jk}$ in the model $y \sim N(\Psi_M \beta^{M}_{jk} + I_{jk}\,\theta^{M}_{jk}, \sigma^2 I_n)$ can be obtained from the conditional distribution of the OLS statistics given the event that the SPAM selected $M$. Theorem 1 gives the resulting selective log-likelihood, and Theorem 2 gives the MLE and observed Fisher information that turn it into approximate Gaussian pivots and Wald confidence intervals. The paper claims these p-values properly account for the data-adaptive selection of $M$, are valid where the naive z-test is not, and are more powerful than data splitting because inference uses the full dataset rather than a holdout. With small randomization variance (chosen as $r=0.9$), the randomized SPAM fit closely matches the nonrandomized fit, so the validity guarantee costs almost nothing in main-effect recovery.
Load-bearing premise
Valid p-values rest on the assumption that the Laplace approximation to the normalizing constant $c(\theta^{M}_{jk}, \beta^{M}_{jk})$ in Section 3.4 is accurate for the data at hand; the paper gives no finite-sample error bound for this approximation, only a reference to prior work and simulation evidence.
Editorial extensions
If this is right
- Analysts can screen interactions under a weak-hierarchy rule after a SPAM fit and trust the reported p-values, because the selection of main effects is explicitly conditioned on.
- Using the full data for inference yields shorter confidence intervals and better F1 recovery of true interactions than data splitting, whose holdout can become rank-deficient.
- A small randomization variance keeps the randomized SPAM almost identical to the plain SPAM, so the validity guarantee does not require a visibly different first-stage model.
- The same selective likelihood also produces valid p-values and confidence intervals for main-effect coefficients in the fitted SPAM, a byproduct noted in the paper.
Reading between the lines
- The paper fixes $r=0.9$ throughout; a natural extension is to tune $r$ as a deliberate trade-off between main-effect selection fidelity and inferential power, or to choose it data-adaptively.
- The selective likelihood should extend to joint tests of several interactions at once, which the paper names as future work; if the Laplace approximation holds there, an analyst could test an entire interaction set rather than one pair at a time.
- Because the conditioning argument relies only on the Gaussian law of $y$ and $\omega$, the same derivation could be carried over to non-Gaussian responses through an asymptotic selective likelihood, an extension the authors flag.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a two-step procedure for testing whether linear interaction terms should be added to a sparse additive model (SPAM). In Step 1, a SPAM is fitted via a randomized group lasso with Gaussian external randomization; in Step 2, for each interaction in a data-adaptive set T^M, the paper constructs p-values for H_0: θ^M_jk = 0 under the model y ~ N(Ψ_M β + I_jk θ, σ² I). The p-values are derived from the conditional distribution of the key statistics given the observed SPAM selection, using a change of variables from the randomization (Sections 3.1–3.3). Theorem 1 gives an exact selective log-likelihood; Section 3.4 replaces its normalizing constant by a Laplace approximation with a barrier; Section 3.5 obtains an MLE and observed Fisher information and forms Wald p-values. Simulations and a JFK flight delay analysis compare the method with naive z-tests and data splitting.
Significance. If the approximation is reliable, the paper fills a gap by providing selective inference for interactions after additive-model selection, with the advantage of using the full data rather than a holdout. The exact conditional likelihood derivation in Theorem 1 is coherent and builds sensibly on the randomized group-lasso machinery of Panigrahi et al. (2023) and Huang et al. (2023). The empirical results show nearly uniform pivots in many settings and improved F1 over data splitting. However, the central validity claim is approximate, and the Laplace approximation is not quantified; this is the main obstacle to accepting the paper in its current form.
major comments (3)
- [Section 3.4, Eq. (12)] The validity of the proposed p-values rests on the Laplace approximation of the normalizing constant c(θ^M_jk, β^M_jk). The manuscript provides no finite-sample error bound for replacing the integral by its supremum plus a barrier, and the reference to Huang et al. (2023) is not directly transferable because the integrand here includes log(1 + 1/g_k) and -log det D_ΠS, which are strongly non-quadratic when any selected group has small norm γ_j. In that regime the Laplace mode need not represent the integral, and the Wald pivot from Section 3.5 need not be Uniform(0,1) under H0. Since the simulations in Section 4 do not include such weak-boundary selected groups, the central claim of valid p-values is not established for this regime. Please add a quantitative approximation bound under interpretable conditions, or at minimum a simulation study that deliberately generates selected groups with small γ_j and reports the empirical distribution of the pivots.
- [Section 3.5, Theorem 2] Theorem 2 is the computational core of the proposed inference, but its proof is omitted with the statement that it is similar to Theorem 4.1 in Huang et al. (2023). Because the expressions for the MLE and observed Fisher information in the approximate selective likelihood are not identical to that theorem, and because the Wald p-values are built on them, a proof or a detailed derivation should be included in the paper or appendix.
- [Section 4.2 vs. Section 3] The theory treats σ as known: Lemma 2 defines Σ^M_jk in terms of σ², and the selective likelihood in Theorem 1 inherits this. In the numerical experiments and JFK application, σ is replaced by a plug-in estimator bσ, but no adjustment or theory is provided for the effect of this estimation on the selective likelihood or the Wald pivot. This is a second source of approximation error beyond Laplace that is not reflected in the stated validity claims.
minor comments (5)
- [Section 2, Algorithm 1] The text in Section 2 and Section 3 refers to 'Algorithm 2' and 'Step 2.1 of Algorithm 2', but the displayed algorithm is labeled Algorithm 1; the numbering should be made consistent.
- [Section 3.5 vs. Section 4.2] In Section 3.5 the pivot is written as z_jk = (bθ^M_jk − θ^M_jk)/sqrt(I^{-1}), while Section 4.2 defines it via bθ^M_mle; the notation should be aligned so that the reader can see that the OLS estimator is replaced by the selective MLE.
- [Table 2 caption] The caption of Table 2 says 'by both naive inference and data splitting', but the columns are headed 'Naive' and 'Proposed'; the caption should be corrected to refer to the proposed method.
- [Table 1 and Figure 7] The heading 'Precison' appears in Table 1 and Figure 7; it should read 'Precision'.
- [Section 4.2] The notation σ_jk is used for both the naive pivot and the selective pivot; different symbols would avoid confusion.
Circularity Check
No construction-level circularity: the selective likelihood is derived from the joint distribution of data and external randomization; the approximate p-values are not fitted to force uniformity. Main caveat is a load-bearing self-citation for the Laplace approximation's formal justification, which is a correctness risk rather than a circular reduction.
full rationale
The paper's central derivation is not circular. Step 1 defines a randomized group lasso SPAM fit (Eq. 1). Lemma 1 characterizes the selection event via KKT conditions; Lemma 2 gives the marginal distribution of the key statistics; Lemma 3 and Proposition 1 build the joint conditional density; Theorem 1 then gives an exact selective log-likelihood for the interaction coefficient. The p-value in Section 3.5 is obtained by plugging an MLE and observed Fisher information from an approximate version of this likelihood into a Wald pivot. At no point is a parameter fitted to force the resulting pivot to be uniform, and the likelihood is not defined in terms of the target null hypothesis. The paper is also transparent that the pivot is 'approximate Uniform(0,1)', not exactly uniform. The one legitimate concern is that the formal justification for the Laplace approximation replacing the normalizing constant is outsourced to Huang et al. (2023), a previous paper with overlapping authorship, and no finite-sample error bound is supplied. This is a load-bearing reliance on self-citation for the approximation's accuracy, and the weak-boundary regime of small selected group norms is not covered by the simulations. However, this is an unsupported approximation and a correctness risk, not a circular derivation: the exact selective likelihood is proven in the appendix, and the approximate pivot is not constructed by design to equal its own input. The paper is therefore essentially self-contained against external benchmarks, with only minor self-citation in the approximation justification, supporting a low circularity score of 2.
Assumptions & free parameters
free parameters (5)
- r (randomization proportion) =
0.9
- lambda_j (group lasso tuning) =
0.5 sigma sqrt(n) sqrt(B_j) sqrt(2 log q)
- epsilon (ridge penalty in group lasso) =
not specified
- B-spline basis degree and knots =
degree 2, 6 knots
- t0 (true interaction threshold in F1 score) =
0.1
assumptions (5)
- standard math The selection event {Mhat = M} has the equivalent characterization {bgamma > 0, bU = U, bZ = Z} from Lemma 1, borrowed from Panigrahi et al. (2023) and Huang et al. (2023).
- standard math The change-of-variables mapping Pi from the randomization omega to (bgamma, bU, bZ) is one-to-one with the Jacobian determinant given in Proposition 1.
- domain assumption The Laplace approximation to the normalizing constant c(theta, beta) in Section 3.4 is accurate enough for the resulting p-values to be approximately valid.
- domain assumption The response y follows the normal linear model y ~ N(Psi_M beta + I_jk theta, sigma^2 I) when testing the interaction I_jk.
- domain assumption The external randomization omega is independent of y and has a known Gaussian distribution N(0, Omega).
Cite this review
Pith. "Pith review of Reluctant Interaction Inference after Additive Modeling." pith.science (2026). https://pith.science/paper/IMOWCGHZ
@misc{pith2026250601219,
author = {Pith},
title = {Pith review of: Reluctant Interaction Inference after Additive Modeling},
year = {2026},
howpublished = {\url{https://pith.science/paper/IMOWCGHZ}},
note = {Machine review of arXiv:2506.01219}
}
read the original abstract
Additive models enjoy the flexibility of nonlinear models while still being readily understandable to humans. By contrast, other nonlinear models, which involve interactions between features, are not only harder to fit but also substantially more complicated to explain. Guided by the principle of parsimony, a data analyst therefore may naturally be reluctant to move beyond an additive model unless it is truly warranted. To put this principle of interaction reluctance into practice, we formulate the problem as a hypothesis test with a fitted sparse additive model (SPAM) serving as the null. Because our hypotheses on interaction effects are formed after fitting a SPAM to the data, we adopt a selective inference approach to construct p-values that properly account for this data adaptivity. Our approach makes use of external randomization to obtain the distribution of test statistics conditional on the SPAM fit, allowing us to derive valid p-values, corrected for the over-optimism introduced by the data-adaptive process prior to the test. Through experiments on simulated and real data, we illustrate that--even with small amounts of external randomization--this rigorous modeling approach enjoys considerable advantages over naive methods and data splitting.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
(2024), Inference with Randomized Regression Trees, arXiv preprint arXiv:2412.20535\/
Bakshi, S., Huang, Y., Panigrahi, S., and Dempsey, W. (2024), Inference with Randomized Regression Trees, arXiv preprint arXiv:2412.20535\/
-
[2]
(2013), Valid post-selection inference, The Annals of Statistics\/ , 802--837
Berk, R., Brown, L., Buja, A., Zhang, K., and Zhao, L. (2013), Valid post-selection inference, The Annals of Statistics\/ , 802--837
work page 2013
-
[3]
(2013), A lasso for hierarchical interactions, Annals of statistics\/ , 41, 1111
Bien, J., Taylor, J., and Tibshirani, R. (2013), A lasso for hierarchical interactions, Annals of statistics\/ , 41, 1111
work page 2013
-
[4]
Breiman, L. (2001), Statistical modeling: The two cultures (with comments and a rejoinder by the author), Statistical science\/ , 16, 199--231
work page 2001
-
[5]
Caruana, R., Lou, Y., Gehrke, J., Koch, P., Sturm, M., and Elhadad, N. (2015), Intelligible models for healthcare: Predicting pneumonia risk and hospital 30-day readmission, in Proceedings of the 21th ACM SIGKDD international conference on knowledge discovery and data mining\/
work page 2015
-
[6]
Chouldechova, A. and Hastie, T. (2015), Generalized additive model selection, arXiv preprint arXiv:1506.03850\/
arXiv 2015
-
[7]
Couch, S. P. (2025), anyflights: Query 'nycflights13'-Like Air Travel Data for Given Years and Airports\/ , ://CRAN.R-project.org/package=anyflights. R package version 0.3.5
work page 2025
-
[8]
D'Amour, A. et al. (2020), Underspecification Presents Challenges for Credibility in Modern Machine Learning, J. Mach. Learn. Res.\/ , 23, 226:1--226:61
work page 2020
Show all 38 references
-
[9]
L., Witten, D., and Bien, J
Dharamshi, A., Neufeld, A., Motwani, K., Gao, L. L., Witten, D., and Bien, J. (2025), Generalized data thinning using sufficient statistics, Journal of the American Statistical Association\/ , 120, 511--523
2025
-
[10]
Fisher, A., Rudin, C., and Dominici, F. (2019), All models are wrong, but many are useful: Learning a variable's importance by studying an entire class of prediction models simultaneously, Journal of Machine Learning Research\/ , 20, 1--81
2019
-
[11]
(2025), Selective Inference in Graphical Models via Maximum Likelihood, arXiv preprint arXiv:2503.24311\/
Guglielmini, S., Claeskens, G., and Panigrahi, S. (2025), Selective Inference in Graphical Models via Maximum Likelihood, arXiv preprint arXiv:2503.24311\/
2025 arXiv
-
[12]
(2022), Generalized sparse additive models, Journal of machine learning research\/ , 23, 1--56
Haris, A., Simon, N., and Shojaie, A. (2022), Generalized sparse additive models, Journal of machine learning research\/ , 23, 1--56
2022
-
[13]
(2016), Convex modeling of interactions with strong heredity, Journal of Computational and Graphical Statistics\/ , 25, 981--1004
Haris, A., Witten, D., and Simon, N. (2016), Convex modeling of interactions with strong heredity, Journal of Computational and Graphical Statistics\/ , 25, 981--1004
2016
-
[14]
and Mazumder, R
Hazimeh, H. and Mazumder, R. (2020), Learning hierarchical interactions at scale: A convex optimization approach, in International Conference on Artificial Intelligence and Statistics\/ , PMLR
2020
-
[15]
(2023), Selective inference using randomized group lasso estimators for general models, arXiv:2306.13829\/
Huang, Y., Pirenne, S., Panigrahi, S., and Claeskens, G. (2023), Selective inference using randomized group lasso estimators for general models, arXiv:2306.13829\/
2023 arXiv
-
[16]
and Leeb, H
Kivaranovic, D. and Leeb, H. (2024), A (tight) upper bound for the length of confidence intervals with conditional coverage, Electronic Journal of Statistics\/ , 18, 1677--1701
2024
-
[17]
D., Sun, D
Lee, J. D., Sun, D. L., Sun, Y., and Taylor, J. E. (2016), Exact post-selection inference, with application to the lasso, The Annals of Statistics\/ , 44, 907--927
2016
-
[18]
(2023), Data fission: splitting a single data point, Journal of the American Statistical Association\/ , 1--12
Leiner, J., Duan, B., Wasserman, L., and Ramdas, A. (2023), Data fission: splitting a single data point, Journal of the American Statistical Association\/ , 1--12
2023
-
[19]
and Zhang, H
Lin, Y. and Zhang, H. H. (2006), Component selection and smoothing in smoothing spline analysis of variance models, Annals of Statistics\/ , 34, 2272--2297
2006
-
[20]
(2016), Sparse partially linear additive models, Journal of Computational and Graphical Statistics\/ , 25, 1126--1140
Lou, Y., Bien, J., Caruana, R., and Gehrke, J. (2016), Sparse partially linear additive models, Journal of Computational and Graphical Statistics\/ , 25, 1126--1140
2016
-
[21]
(2024), Hybrid confidence intervals for informative uniform asymptotic inference after model selection, Biometrika\/ , 111, 109--127
McCloskey, A. (2024), Hybrid confidence intervals for informative uniform asymptotic inference after model selection, Biometrika\/ , 111, 109--127
2024
-
[22]
(1977), A reformulation of linear models, Journal of the Royal Statistical Society Series A: Statistics in Society\/ , 140, 48--63
Nelder, J. (1977), A reformulation of linear models, Journal of the Royal Statistical Society Series A: Statistics in Society\/ , 140, 48--63
1977
-
[23]
L., and Witten, D
Neufeld, A., Dharamshi, A., Gao, L. L., and Witten, D. (2024), Data thinning for convolution-closed distributions, Journal of Machine Learning Research\/ , 25, 1--35
2024
-
[24]
C., Gao, L
Neufeld, A. C., Gao, L. L., and Witten, D. M. (2022), Tree-values: selective inference for regression trees, Journal of Machine Learning Research\/ , 23, 1--43
2022
-
[25]
(2024), Exact selective inference with randomization, Biometrika\/ , 111, 1109--1127
Panigrahi, S., Fry, K., and Taylor, J. (2024), Exact selective inference with randomization, Biometrika\/ , 111, 1109--1127
2024
-
[26]
W., and Kessler, D
Panigrahi, S., MacDonald, P. W., and Kessler, D. (2023), Approximate post-selective inference for regression with the group lasso, Journal of Machine Learning Research\/ , 24, 1--49
2023
-
[27]
(2017), An MCMC-free approach to post-selective inference, arXiv preprint arXiv:1703.06154\/
Panigrahi, S., Markovic, J., and Taylor, J. (2017), An MCMC-free approach to post-selective inference, arXiv preprint arXiv:1703.06154\/
2017 arXiv
-
[28]
and Taylor, J
Panigrahi, S. and Taylor, J. (2023), Approximate selective inference via maximum likelihood, Journal of the American Statistical Association\/ , 118, 2810--2820
2023
-
[29]
(2021), Integrative methods for post-selection inference under convex constraints, The Annals of Statistics\/ , 49, 2803--2824
Panigrahi, S., Taylor, J., and Weinstein, A. (2021), Integrative methods for post-selection inference under convex constraints, The Annals of Statistics\/ , 49, 2803--2824
2021
-
[30]
Peixoto, J. L. (1987), Hierarchical variable selection in polynomial regression models, The American Statistician\/ , 41, 311--313
1987
-
[31]
Rasines, D. G. and Young, G. A. (2023), Splitting strategies for post-selection inference, Biometrika\/ , 110, 597--614
2023
-
[32]
(2009), Sparse additive models, Journal of the Royal Statistical Society Series B: Statistical Methodology\/ , 71, 1009--1030
Ravikumar, P., Lafferty, J., Liu, H., and Wasserman, L. (2009), Sparse additive models, Journal of the Royal Statistical Society Series B: Statistical Methodology\/ , 71, 1009--1030
2009
-
[33]
(2022), Interpretable machine learning: Fundamental principles and 10 grand challenges , Statistics Surveys\/ , 16, 1 -- 85, ://doi.org/10.1214/21-SS133
Rudin, C., Chen, C., Chen, Z., Huang, H., Semenova, L., and Zhong, C. (2022), Interpretable machine learning: Fundamental principles and 10 grand challenges , Statistics Surveys\/ , 16, 1 -- 85, ://doi.org/10.1214/21-SS133
2022 doi
-
[34]
(2017), Selective inference for sparse high-order interaction models, in International Conference on Machine Learning\/ , PMLR
Suzumura, S., Nakagawa, K., Umezu, Y., Tsuda, K., and Takeuchi, I. (2017), Selective inference for sparse high-order interaction models, in International Conference on Machine Learning\/ , PMLR
2017
-
[35]
Tay, J. K. and Tibshirani, R. (2020), Reluctant generalised additive modelling, International Statistical Review\/ , 88, S205--S224
2020
-
[36]
(2019), Reluctant interaction modeling, arXiv:1907.08414\/
Yu, G., Bien, J., and Tibshirani, R. (2019), Reluctant interaction modeling, arXiv:1907.08414\/
2019 arXiv
-
[37]
R., and Zou, H
Yuan, M., Joseph, V. R., and Zou, H. (2009), Structured variable selection and estimation, The Annals of Applied Statistics\/ , 1738--1757
2009
-
[38]
and Fithian, W
Zrnic, T. and Fithian, W. (2024), Locally simultaneous inference, The Annals of Statistics\/ , 52, 1227--1253
2024
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.