Pith. sign in

REVIEW 2 major objections 5 minor 3 references

Performance of variable and function selection methods for estimating the non-linear health effects of correlated chemical mixtures: a simulation study

T0 review · 2 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read The paper claims that flexible Bayesian methods BKMR and BSTARSS identify harmful chemicals in correlated mixtures as well as lasso when effects are linear and better when effects are non-linear.

desk verdict A careful, reproducible simulation study that fills a real gap, but the headline conclusion overstates the case for BKMR/BSTARSS over lasso beyond the symmetric inverse-U shape. read the letter →

arxiv 1908.01583 v1 pith:Y4MROU5V submitted 2019-08-05 stat.AP

classification stat.AP MSC 62F1562J0762P10
keywords chemicalmixturesvariableselectionnon-linearexposure-responseBKMRBARTBSTARSSlassopenalisedregressionendocrine-disruptingchemicals
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks a practical question: when a health outcome is influenced by several correlated chemicals, and some of the dose-response curves bend, which statistical method should an environmental epidemiologist trust to pick out the harmful chemicals? The authors simulate realistic exposure data for twelve phthalates and phenols, generate outcomes with four different exposure-response shapes (linear, S-shaped, symmetric inverse U, asymmetric inverse U), and compare three flexible Bayesian methods (BKMR, BART, BSTARSS) with lasso penalised regression across 32 scenarios. Their central claim is that the flexible Bayesian methods, especially BKMR and BSTARSS, identify the truly outcome-associated exposures with little loss relative to lasso when the truth is linear and with a clear advantage when the truth is non-linear. They conclude that methods assuming linearity, such as lasso, are unsuitable when non-monotonic dose-response relationships are plausible, and that BKMR and BSTARSS are attractive because they also recover the shape of the curve.

What carries the argument

The engine of the study is a simulation design in which outcomes are generated from an additive main-effects model $Y_i = \sum_{j=1}^{J^*} f_j(x_{ij}) + \varepsilon_i$, with no interactions and only $J^*=4$ of the $J=6$ or $J=12$ exposures affecting the response. Each associated exposure's function $f_j$ is one of four shapes—linear, S-shaped (log-logistic CDF), symmetric inverse-U (quadratic), or asymmetric inverse-U (Dawson function)—scaled to equal area under the curve. Exposure vectors are drawn from a multivariate t copula with truncated kernel-smoothed empirical margins and the observed Spearman correlation of the biomonitoring sample, giving both an observed-correlation and a half-correlation version. The comparison is carried by a battery of metrics: sensitivity, specificity, precision, negative predictive value, the F1-statistic, the proportion of replications with perfect exposure ranking, mean-squared error relative to an oracle GAM, and 90% credible-interval coverage, across 32 scenarios formed by two model sizes, two signal-to-noise ratios, two correlation structures, and four curve shapes.

What would settle it

Repeat the same 32-scenario simulation with an interaction term between two outcome-associated exposures, or with curves scaled to equal peak height instead of equal area, and check whether BSTARSS still ranks all true exposures above all null exposures as often as reported.

Watch

Extended reading notes

Core claim

The paper's central discovery is comparative: across 32 simulation scenarios built from copula-simulated phthalate and phenol exposures, BKMR and BSTARSS consistently balance true positives and false positives, while lasso fails specifically when the exposure-response curve is a symmetric inverted U and BART selects too few exposures. BSTARSS had the best F1-statistic in 25 of 32 scenarios, and it or BKMR ranked the truly associated exposures above the null exposures most reliably in the majority of scenarios. Lasso was highly sensitive for linear, S-shaped, and asymmetric inverse-U relationships but had sensitivity at or below 0.20 for symmetric inverse-U relationships. In estimation, BKMR and BSTARSS matched the mean-squared error of an oracle generalized additive model fitted to the true model, whereas BART was less accurate; BSTARSS and BART produced excessively wide credible intervals, while BKMR's coverage stayed close to 90% for most shapes. The authors conclude that there is little cost to using BKMR or BSTARSS instead of lasso when relationships are linear, and a distinct advantage when they are non-linear.

Load-bearing premise

The entire ranking rests on the simulation's data-generating process being additive with no interactions, exactly four exposures driving the outcome, and every curve scaled to equal total area; if real chemical mixtures act through interactions or differently shaped curves, the reported rankings may change.

Editorial extensions

If this is right

  • Environmental-health studies that suspect non-monotonic dose-response relationships should not rely on linear penalised regression alone; a symmetric U-shaped effect can be invisible to lasso.
  • BKMR and BSTARSS give researchers two things at once: a selected set of chemicals and an estimated curve shape, with accuracy close to a correctly specified GAM.
  • BART can serve as a conservative screen: nearly every exposure it selects is likely a true positive, but it will miss many real associations, especially in low-signal or low-sparsity settings.
  • Signal-to-noise ratio matters more than exposure correlation for ranking chemicals, so studies with weak effects need larger samples or stronger priors, not just better variable-selection software.
  • Choosing BKMR or BSTARSS over lasso costs little even when the true relationships are linear, so flexibility is nearly free in the settings tested.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Not tested here: a two-stage strategy that screens with lasso and then re-fits flexibly would still miss a U-shaped component at the screening stage, so the flexible method would need to be the primary analysis.
  • I would infer that adding pairwise interactions to the simulation would improve BKMR and BART relative to BSTARSS, because those methods model multivariate functions by default while the paper's additive design is acknowledged to favour BSTARSS.
  • The equal-area curve scaling may make non-monotonic effects easier to detect than in real data, where low-dose effects can concentrate in a narrow exposure window; simulating equal-peak or equal-slope curves would stress-test the conclusion.
  • The fixed 0.5 inclusion-probability threshold is a free choice; an adaptive threshold or a threshold tuned to posterior predictive performance could raise BART's sensitivity without giving up its near-perfect specificity.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper presents a simulation study comparing four methods (BKMR, BART, BSTARSS, and lasso) for variable and function selection in correlated chemical mixture analyses. Exposures are simulated from a t copula fitted to NHANES data, with 12 phthalates and phenols; outcomes are generated from additive main-effects models with four associated exposures, using linear, S-shaped, symmetric inverse-U (quadratic), and asymmetric inverse-U shapes, two signal-to-noise ratios, two correlation structures, and two model sizes (J=6 and J=12). Performance is measured by sensitivity, specificity, precision, F1, exposure ranking, mean-squared error relative to an oracle GAM, and credible interval coverage. The main finding is that BKMR and BSTARSS perform well across non-linear scenarios, while lasso performs well for linear and approximately linear relationships but fails for symmetric inverse-U shapes; BART is highly specific but less sensitive. The authors conclude by recommending BKMR and BSTARSS for studies of non-monotonic relationships.

Significance. The study addresses a practical gap in the environmental mixtures literature: how non-monotonicity affects exposure selection. Its strengths include a realistic simulation design grounded in the observed NHANES correlation structure, clearly described data-generating processes, the use of an oracle method for estimation accuracy, and publicly available code. The authors also honestly acknowledge the additive main-effects-only DGP that favors BSTARSS, the limited number of replications, and the manual tuning of Bayesian hyperparameters. If the conclusions are appropriately qualified, the paper provides useful guidance for environmental epidemiologists choosing among statistical methods. However, the 'distinct advantage' claim in the conclusion is currently overstated relative to the paper's own results, which is a load-bearing issue for the headline recommendation.

major comments (2)
  1. [Section 4.2 (Conclusions) and Sections 3.1.1, 3.1.3] The conclusion states that BKMR and BSTARSS had a 'distinct advantage' over lasso when exposure-response relationships are non-linear. This is not supported by the reported results for all non-linear shapes. Section 3.1.3 says that for monotonic (S-shaped) and asymmetric inverse-U relationships, lasso 'tended to perform comparably and sometimes marginally better than BSTARSS and BKMR' in terms of F1, and Section 3.1.1 reports lasso sensitivity of 0.75-0.99 for those shapes. The only shape where lasso clearly failed is the symmetric inverse-U (quadratic) relationship, with sensitivity 0.13-0.20. The abstract includes the qualifier 'except for symmetric inverse-U-shaped relationships,' but the conclusion drops it. This overgeneralization directly inflates the headline recommendation and should be fixed by revising the conclusion to specify that the advantage is specific to symmetric inverse-U shapes or by adding the same qualifier used in the abstract.
  2. [Section 4.1 (Limitations) and Discussion, first paragraph] The paper's comparative claims about non-linear performance are derived from a data-generating process with additive main effects only and no interactions. The authors acknowledge this in the Discussion and in Section 4.1, noting that this favors BSTARSS. However, because BART and BKMR are often advocated for their ability to detect interactions, the conclusions should explicitly state that the reported rankings are conditional on the absence of interactions; otherwise readers may over-generalize the 'distinct advantage' recommendation to settings with interaction effects. This scoping is important for the central message and should be stated in the conclusion, not only in the limitations.
minor comments (5)
  1. [Section 2.3] The equation for the data-generating process contains garbled symbols in the manuscript text; please ensure it renders as y_i = sum_j f_j(x_ij) + epsilon_i.
  2. [Sections 2.4.1-2.4.3 and Discussion] The prior sensitivity analyses are referenced as 'Section 3 of the Supplementary Material' in the Methods sections and as 'Section 4' in the Discussion; the numbering should be harmonized.
  3. [Reproducibility statement] The R code is given as 'https://github.com/n-lazarevic/' which is a user page rather than a direct repository link; please provide the full URL to the specific repository.
  4. [Section 3.1.1] Given that only 100 replications are used, reporting Monte Carlo standard errors or 95% confidence intervals for the mean sensitivity and specificity values would help readers assess the strength of the observed differences between methods.
  5. [Figure 5] The boxplot legend appears to use the same symbol for means and outliers; please clarify the symbology in the caption.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the simulation targets are generated from known exposure–response functions and the methods are evaluated against those external targets.

full rationale

This paper is a simulation benchmark. The outcome data in each replication are generated from explicitly specified functions (linear, log-logistic S-shaped, quadratic inverse-U, and Dawson asymmetric inverse-U) with known association strengths, and the target quantities (sensitivity, specificity, precision, F1, ranking, MSE, credible interval coverage) are defined relative to those known simulated targets. No method's fitted parameter is used to define the truth, and no prediction is equivalent by construction to an input. The only self-reference is the introductory sentence 'Building on our previous work,4 we focused on exposure to EDC-mixtures during pregnancy', which is a study-design choice and not load-bearing for any reported performance comparison. The cited prior review is not used to justify the simulation outcomes or to forbid alternative interpretations. The authors' own discussion acknowledges limitations (additive data-generating process, no interactions, fixed curve shapes), but those are scope limitations, not circularity. The central comparison is self-contained against independent simulated data and an oracle GAM benchmark.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

The simulation benchmark is self-contained and does not rest on new theoretical entities. The main imported assumptions are the realism of the generated exposure data and the representativeness of the chosen dose-response shapes. Method hyperparameters are numerous and were chosen by the investigators, so they are listed as free parameters because they can affect the relative performance.

free parameters (5)
  • t copula parameters (correlation matrix and degrees of freedom) = Not reported; fitted to NHANES data by maximum likelihood
    Used to simulate realistic correlated exposures in Section 2.2; the dependence structure affects correlation between chemicals and thus method performance.
  • BKMR slab prior Gamma(1,4) = mean 0.25, sd 0.25
    Chosen in Section 2.4.1 by fitting frequentist kernel machine regression and observing which lambda values gave appropriate smoothing; affects BKMR selection and curve estimates.
  • BKMR MCMC proposal standard deviations = 0.5/1 for rho, 0.1/0.2 for lambda depending on SNR
    Tuned to achieve 20-40% acceptance rates; affects posterior mixing and inclusion probabilities.
  • BART hyperparameters (number of trees, alpha, beta, nu, q) = m=50, alpha=0.95, beta=2, nu=3, q=0.9
    Defaults recommended by Chipman et al. (2010) and Kapelner and Bleich (2016); smaller tree count preferred for variable selection (Section 2.4.2).
  • BSTARSS spike-slab and spline parameters = c0=0.025, a0=5, b0=40, a1=1, b1=1, 20 B-spline basis functions
    Set in Section 2.4.3; these priors control selection sparsity and smooth function estimation and may influence BSTARSS performance.
assumptions (4)
  • domain assumption The t copula adequately represents the dependence structure of the NHANES exposure data.
    The simulation realism in Section 2.2 relies on this; the paper selected the t copula by maximum likelihood among several copulas, but this does not guarantee faithful representation of the true data-generating dependence.
  • domain assumption The four exposure-response functions represent plausible non-monotonic EDC dose-response shapes.
    Section 2.3 defines linear, log-logistic S-shaped, quadratic inverse-U, and Dawson asymmetric inverse-U functions. The comparative conclusions about lasso and Bayesian methods depend on these shapes being representative.
  • domain assumption The data-generating process is additive with no interactions.
    Section 2.3 states 'For simplicity, we assumed no confounding by non-exposure variables, no interaction'; Section 4.1 acknowledges this favors BSTARSS. If interactions are present, BKMR and BART's ability to model them might change the ranking.
  • standard math MCMC samples for BKMR, BART, and BSTARSS have converged to the posterior.
    The paper states convergence diagnostics are in Supplementary Material; the reported posterior means and inclusion probabilities assume the chains are representative.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Performance of variable and function selection methods for estimating the non-linear health effects of correlated chemical mixtures: a simulation study." pith.science (2026). https://pith.science/paper/Y4MROU5V

@misc{pith2026190801583,
  author       = {Pith},
  title        = {Pith review of: Performance of variable and function selection methods for estimating the non-linear health effects of correlated chemical mixtures: a simulation study},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Y4MROU5V}},
  note         = {Machine review of arXiv:1908.01583}
}
abstract

Statistical methods for identifying harmful chemicals in a correlated mixture often assume linearity in exposure-response relationships. Non-monotonic relationships are increasingly recognised (e.g., for endocrine-disrupting chemicals); however, the impact of non-monotonicity on exposure selection has not been evaluated. In a simulation study, we assessed the performance of Bayesian kernel machine regression (BKMR), Bayesian additive regression trees (BART), Bayesian structured additive regression with spike-slab priors (BSTARSS), and lasso penalised regression. We used data on exposure to 12 phthalates and phenols in pregnant women from the U.S. National Health and Nutrition Examination Survey to simulate realistic exposure data using a multivariate copula. We simulated datasets of size N = 250 and compared methods across 32 scenarios, varying by model size and sparsity, signal-to-noise ratio, correlation structure, and exposure-response relationship shapes. We compared methods in terms of their sensitivity, specificity, and estimation accuracy. In most scenarios, BKMR and BSTARSS achieved moderate to high specificity (0.56--0.91 and 0.57--0.96, respectively) and sensitivity (0.49--0.98 and 0.25--0.97, respectively). BART achieved high specificity ($\geq$ 0.96), but low to moderate sensitivity (0.13--0.66). Lasso was highly sensitive (0.75--0.99), except for symmetric inverse-U-shaped relationships ($\leq$ 0.2). Performance was affected by the signal-to-noise ratio, but not substantially by the correlation structure. Penalised regression methods that assume linearity, such as lasso, may not be suitable for studies of environmental chemicals hypothesised to have non-monotonic relationships with outcomes. Instead, BKMR and BSTARSS are attractive methods for flexibly estimating the shapes of exposure-response relationships and selecting among correlated exposures.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

3 extracted references · 3 canonical work pages

  1. [5]

    exposome

    Stafoggia M, Breitner S, Hampel R, Basagaña X. Statistical approaches to address multi-pollutant mixtures and multiple exposures: the state of the science. Curr Environ Health Rep. 2017;4(4):481–90. 6. Sun Z, Tao Y, Li S, Ferguson KK, Meeker JD, Park SK, et al. Statistical strategies for constructing health risk models with multiple pollutants and their i...

  2. [23]

    Special functions in R: introducing the gsl package

    Hankin RKS. Special functions in R: introducing the gsl package. R News. 2006;6(4):24–6. 24. Bobb JF, Valeri L, Claus Henn B, Christiani DC, Wright RO, Mazumdar M, et al. Bayesian kernel machine regression for estimating the health effects of multi-pollutant mixtures. Biostatistics. 2015;16(3):493–508. 25. Liu D, Lin X, Ghosh D. Semiparametric Regression ...

  3. [44]

    Stability selection

    Fahrmeir L, Kneib T, Konrath S. Bayesian regularisation in structured additive regression: a unifying perspective on shrinkage, smoothing and predictor selection. Stat Comput. 2009;20(2):203–19. 45. Richardson S. Discussion on the paper "Stability selection" by Meinshausen N, Bühlmann P. J R Stat Soc: Series B Stat Methodol. 2010;72(4):417–73. 46. Kirk P,...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.