REVIEW 2 major objections 5 minor 3 references
Performance of variable and function selection methods for estimating the non-linear health effects of correlated chemical mixtures: a simulation study
T0 review · 2 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read The paper claims that flexible Bayesian methods BKMR and BSTARSS identify harmful chemicals in correlated mixtures as well as lasso when effects are linear and better when effects are non-linear.
desk verdict A careful, reproducible simulation study that fills a real gap, but the headline conclusion overstates the case for BKMR/BSTARSS over lasso beyond the symmetric inverse-U shape. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine of the study is a simulation design in which outcomes are generated from an additive main-effects model $Y_i = \sum_{j=1}^{J^*} f_j(x_{ij}) + \varepsilon_i$, with no interactions and only $J^*=4$ of the $J=6$ or $J=12$ exposures affecting the response. Each associated exposure's function $f_j$ is one of four shapes—linear, S-shaped (log-logistic CDF), symmetric inverse-U (quadratic), or asymmetric inverse-U (Dawson function)—scaled to equal area under the curve. Exposure vectors are drawn from a multivariate t copula with truncated kernel-smoothed empirical margins and the observed Spearman correlation of the biomonitoring sample, giving both an observed-correlation and a half-correlation version. The comparison is carried by a battery of metrics: sensitivity, specificity, precision, negative predictive value, the F1-statistic, the proportion of replications with perfect exposure ranking, mean-squared error relative to an oracle GAM, and 90% credible-interval coverage, across 32 scenarios formed by two model sizes, two signal-to-noise ratios, two correlation structures, and four curve shapes.
What would settle it
Repeat the same 32-scenario simulation with an interaction term between two outcome-associated exposures, or with curves scaled to equal peak height instead of equal area, and check whether BSTARSS still ranks all true exposures above all null exposures as often as reported.
Extended reading notes
Core claim
The paper's central discovery is comparative: across 32 simulation scenarios built from copula-simulated phthalate and phenol exposures, BKMR and BSTARSS consistently balance true positives and false positives, while lasso fails specifically when the exposure-response curve is a symmetric inverted U and BART selects too few exposures. BSTARSS had the best F1-statistic in 25 of 32 scenarios, and it or BKMR ranked the truly associated exposures above the null exposures most reliably in the majority of scenarios. Lasso was highly sensitive for linear, S-shaped, and asymmetric inverse-U relationships but had sensitivity at or below 0.20 for symmetric inverse-U relationships. In estimation, BKMR and BSTARSS matched the mean-squared error of an oracle generalized additive model fitted to the true model, whereas BART was less accurate; BSTARSS and BART produced excessively wide credible intervals, while BKMR's coverage stayed close to 90% for most shapes. The authors conclude that there is little cost to using BKMR or BSTARSS instead of lasso when relationships are linear, and a distinct advantage when they are non-linear.
Load-bearing premise
The entire ranking rests on the simulation's data-generating process being additive with no interactions, exactly four exposures driving the outcome, and every curve scaled to equal total area; if real chemical mixtures act through interactions or differently shaped curves, the reported rankings may change.
Editorial extensions
If this is right
- Environmental-health studies that suspect non-monotonic dose-response relationships should not rely on linear penalised regression alone; a symmetric U-shaped effect can be invisible to lasso.
- BKMR and BSTARSS give researchers two things at once: a selected set of chemicals and an estimated curve shape, with accuracy close to a correctly specified GAM.
- BART can serve as a conservative screen: nearly every exposure it selects is likely a true positive, but it will miss many real associations, especially in low-signal or low-sparsity settings.
- Signal-to-noise ratio matters more than exposure correlation for ranking chemicals, so studies with weak effects need larger samples or stronger priors, not just better variable-selection software.
- Choosing BKMR or BSTARSS over lasso costs little even when the true relationships are linear, so flexibility is nearly free in the settings tested.
Reading between the lines
- Not tested here: a two-stage strategy that screens with lasso and then re-fits flexibly would still miss a U-shaped component at the screening stage, so the flexible method would need to be the primary analysis.
- I would infer that adding pairwise interactions to the simulation would improve BKMR and BART relative to BSTARSS, because those methods model multivariate functions by default while the paper's additive design is acknowledged to favour BSTARSS.
- The equal-area curve scaling may make non-monotonic effects easier to detect than in real data, where low-dose effects can concentrate in a narrow exposure window; simulating equal-peak or equal-slope curves would stress-test the conclusion.
- The fixed 0.5 inclusion-probability threshold is a free choice; an adaptive threshold or a threshold tuned to posterior predictive performance could raise BART's sensitivity without giving up its near-perfect specificity.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a simulation study comparing four methods (BKMR, BART, BSTARSS, and lasso) for variable and function selection in correlated chemical mixture analyses. Exposures are simulated from a t copula fitted to NHANES data, with 12 phthalates and phenols; outcomes are generated from additive main-effects models with four associated exposures, using linear, S-shaped, symmetric inverse-U (quadratic), and asymmetric inverse-U shapes, two signal-to-noise ratios, two correlation structures, and two model sizes (J=6 and J=12). Performance is measured by sensitivity, specificity, precision, F1, exposure ranking, mean-squared error relative to an oracle GAM, and credible interval coverage. The main finding is that BKMR and BSTARSS perform well across non-linear scenarios, while lasso performs well for linear and approximately linear relationships but fails for symmetric inverse-U shapes; BART is highly specific but less sensitive. The authors conclude by recommending BKMR and BSTARSS for studies of non-monotonic relationships.
Significance. The study addresses a practical gap in the environmental mixtures literature: how non-monotonicity affects exposure selection. Its strengths include a realistic simulation design grounded in the observed NHANES correlation structure, clearly described data-generating processes, the use of an oracle method for estimation accuracy, and publicly available code. The authors also honestly acknowledge the additive main-effects-only DGP that favors BSTARSS, the limited number of replications, and the manual tuning of Bayesian hyperparameters. If the conclusions are appropriately qualified, the paper provides useful guidance for environmental epidemiologists choosing among statistical methods. However, the 'distinct advantage' claim in the conclusion is currently overstated relative to the paper's own results, which is a load-bearing issue for the headline recommendation.
major comments (2)
- [Section 4.2 (Conclusions) and Sections 3.1.1, 3.1.3] The conclusion states that BKMR and BSTARSS had a 'distinct advantage' over lasso when exposure-response relationships are non-linear. This is not supported by the reported results for all non-linear shapes. Section 3.1.3 says that for monotonic (S-shaped) and asymmetric inverse-U relationships, lasso 'tended to perform comparably and sometimes marginally better than BSTARSS and BKMR' in terms of F1, and Section 3.1.1 reports lasso sensitivity of 0.75-0.99 for those shapes. The only shape where lasso clearly failed is the symmetric inverse-U (quadratic) relationship, with sensitivity 0.13-0.20. The abstract includes the qualifier 'except for symmetric inverse-U-shaped relationships,' but the conclusion drops it. This overgeneralization directly inflates the headline recommendation and should be fixed by revising the conclusion to specify that the advantage is specific to symmetric inverse-U shapes or by adding the same qualifier used in the abstract.
- [Section 4.1 (Limitations) and Discussion, first paragraph] The paper's comparative claims about non-linear performance are derived from a data-generating process with additive main effects only and no interactions. The authors acknowledge this in the Discussion and in Section 4.1, noting that this favors BSTARSS. However, because BART and BKMR are often advocated for their ability to detect interactions, the conclusions should explicitly state that the reported rankings are conditional on the absence of interactions; otherwise readers may over-generalize the 'distinct advantage' recommendation to settings with interaction effects. This scoping is important for the central message and should be stated in the conclusion, not only in the limitations.
minor comments (5)
- [Section 2.3] The equation for the data-generating process contains garbled symbols in the manuscript text; please ensure it renders as y_i = sum_j f_j(x_ij) + epsilon_i.
- [Sections 2.4.1-2.4.3 and Discussion] The prior sensitivity analyses are referenced as 'Section 3 of the Supplementary Material' in the Methods sections and as 'Section 4' in the Discussion; the numbering should be harmonized.
- [Reproducibility statement] The R code is given as 'https://github.com/n-lazarevic/' which is a user page rather than a direct repository link; please provide the full URL to the specific repository.
- [Section 3.1.1] Given that only 100 replications are used, reporting Monte Carlo standard errors or 95% confidence intervals for the mean sensitivity and specificity values would help readers assess the strength of the observed differences between methods.
- [Figure 5] The boxplot legend appears to use the same symbol for means and outliers; please clarify the symbology in the caption.
Circularity Check
No significant circularity: the simulation targets are generated from known exposure–response functions and the methods are evaluated against those external targets.
full rationale
This paper is a simulation benchmark. The outcome data in each replication are generated from explicitly specified functions (linear, log-logistic S-shaped, quadratic inverse-U, and Dawson asymmetric inverse-U) with known association strengths, and the target quantities (sensitivity, specificity, precision, F1, ranking, MSE, credible interval coverage) are defined relative to those known simulated targets. No method's fitted parameter is used to define the truth, and no prediction is equivalent by construction to an input. The only self-reference is the introductory sentence 'Building on our previous work,4 we focused on exposure to EDC-mixtures during pregnancy', which is a study-design choice and not load-bearing for any reported performance comparison. The cited prior review is not used to justify the simulation outcomes or to forbid alternative interpretations. The authors' own discussion acknowledges limitations (additive data-generating process, no interactions, fixed curve shapes), but those are scope limitations, not circularity. The central comparison is self-contained against independent simulated data and an oracle GAM benchmark.
Assumptions & free parameters
free parameters (5)
- t copula parameters (correlation matrix and degrees of freedom) =
Not reported; fitted to NHANES data by maximum likelihood
- BKMR slab prior Gamma(1,4) =
mean 0.25, sd 0.25
- BKMR MCMC proposal standard deviations =
0.5/1 for rho, 0.1/0.2 for lambda depending on SNR
- BART hyperparameters (number of trees, alpha, beta, nu, q) =
m=50, alpha=0.95, beta=2, nu=3, q=0.9
- BSTARSS spike-slab and spline parameters =
c0=0.025, a0=5, b0=40, a1=1, b1=1, 20 B-spline basis functions
assumptions (4)
- domain assumption The t copula adequately represents the dependence structure of the NHANES exposure data.
- domain assumption The four exposure-response functions represent plausible non-monotonic EDC dose-response shapes.
- domain assumption The data-generating process is additive with no interactions.
- standard math MCMC samples for BKMR, BART, and BSTARSS have converged to the posterior.
Cite this review
Pith. "Pith review of Performance of variable and function selection methods for estimating the non-linear health effects of correlated chemical mixtures: a simulation study." pith.science (2026). https://pith.science/paper/Y4MROU5V
@misc{pith2026190801583,
author = {Pith},
title = {Pith review of: Performance of variable and function selection methods for estimating the non-linear health effects of correlated chemical mixtures: a simulation study},
year = {2026},
howpublished = {\url{https://pith.science/paper/Y4MROU5V}},
note = {Machine review of arXiv:1908.01583}
}
abstract
Statistical methods for identifying harmful chemicals in a correlated mixture often assume linearity in exposure-response relationships. Non-monotonic relationships are increasingly recognised (e.g., for endocrine-disrupting chemicals); however, the impact of non-monotonicity on exposure selection has not been evaluated. In a simulation study, we assessed the performance of Bayesian kernel machine regression (BKMR), Bayesian additive regression trees (BART), Bayesian structured additive regression with spike-slab priors (BSTARSS), and lasso penalised regression. We used data on exposure to 12 phthalates and phenols in pregnant women from the U.S. National Health and Nutrition Examination Survey to simulate realistic exposure data using a multivariate copula. We simulated datasets of size N = 250 and compared methods across 32 scenarios, varying by model size and sparsity, signal-to-noise ratio, correlation structure, and exposure-response relationship shapes. We compared methods in terms of their sensitivity, specificity, and estimation accuracy. In most scenarios, BKMR and BSTARSS achieved moderate to high specificity (0.56--0.91 and 0.57--0.96, respectively) and sensitivity (0.49--0.98 and 0.25--0.97, respectively). BART achieved high specificity ($\geq$ 0.96), but low to moderate sensitivity (0.13--0.66). Lasso was highly sensitive (0.75--0.99), except for symmetric inverse-U-shaped relationships ($\leq$ 0.2). Performance was affected by the signal-to-noise ratio, but not substantially by the correlation structure. Penalised regression methods that assume linearity, such as lasso, may not be suitable for studies of environmental chemicals hypothesised to have non-monotonic relationships with outcomes. Instead, BKMR and BSTARSS are attractive methods for flexibly estimating the shapes of exposure-response relationships and selecting among correlated exposures.
Reference graph
Works this paper leans on
-
[5]
Stafoggia M, Breitner S, Hampel R, Basagaña X. Statistical approaches to address multi-pollutant mixtures and multiple exposures: the state of the science. Curr Environ Health Rep. 2017;4(4):481–90. 6. Sun Z, Tao Y, Li S, Ferguson KK, Meeker JD, Park SK, et al. Statistical strategies for constructing health risk models with multiple pollutants and their i...
work page 2017
-
[23]
Special functions in R: introducing the gsl package
Hankin RKS. Special functions in R: introducing the gsl package. R News. 2006;6(4):24–6. 24. Bobb JF, Valeri L, Claus Henn B, Christiani DC, Wright RO, Mazumdar M, et al. Bayesian kernel machine regression for estimating the health effects of multi-pollutant mixtures. Biostatistics. 2015;16(3):493–508. 25. Liu D, Lin X, Ghosh D. Semiparametric Regression ...
work page 2006
-
[44]
Fahrmeir L, Kneib T, Konrath S. Bayesian regularisation in structured additive regression: a unifying perspective on shrinkage, smoothing and predictor selection. Stat Comput. 2009;20(2):203–19. 45. Richardson S. Discussion on the paper "Stability selection" by Meinshausen N, Bühlmann P. J R Stat Soc: Series B Stat Methodol. 2010;72(4):417–73. 46. Kirk P,...
work page 2009
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.