Pith. sign in

REVIEW 3 major objections 6 minor 50 references

This paper claims that adding a Markov random field prior that links nutritionally similar food items to a skew-normal censored mixture model raises the detection of true diet–metabolite associations in censored, skewed, high-dimensional me

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-04 23:05 UTC pith:CSCUMM4S

load-bearing objection Solid integration of skew-normal censored mixture with MRF variable selection, but the new eta rule and simulation design need closer scrutiny before the TPR/FDR claim fully lands. the 3 major comments →

arxiv 2509.06779 v1 pith:CSCUMM4S submitted 2025-09-08 stat.ME

A nutritionally informed model for Bayesian variable selection with metabolite response variables

classification stat.ME MSC 62F1562J0762N0162P10
keywords Bayesian variable selectionskew-normal censored mixtureMarkov random field priormetabolomicspoint mass valuesdiet–metabolite associationsleft-censored dataspike-and-slab prior
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to establish that Bayesian variable selection can reliably find which foods are tied to blood metabolites even when many metabolite measurements are missing, left-censored, or highly skewed. Its key idea is to make selection of one food predictor encourage selection of nutritionally similar foods, through a Markov random field prior built on a pre-specified relationship matrix. In simulations with 300 predictors, that prior raised the rate of finding true links, especially for weak signals, while keeping the false discovery rate near its nominal level; a mis-specified relationship matrix did little harm. Applied to two cohorts of healthcare professionals, the model reported 54 food–metabolite associations, including several that the independent-prior version missed.

Core claim

The authors claim that a skew-normal censored mixture model combined with a Markov random field prior over inclusion indicators addresses four data problems at once: point-mass missing values, left-censoring below a detection limit, right skewness, and correlated predictors. The prior writes P(gamma) proportional to exp(omega sum gamma_j + eta gamma^T R gamma), where R encodes nutritional similarity among the 30 foods; eta controls how strongly joint selection is encouraged. The paper proposes a data-independent rule for eta—choose the largest value whose prior 95th percentile model size stays below that of a scaled independent Bernoulli prior—so the prior borrows strength without crossing i

What carries the argument

The central object is the SNCM likelihood paired with a Markov random field (MRF) prior on the selection indicators gamma_1,...,gamma_p. The SNCM likelihood models a metabolite as the product of a presence indicator and a skew-normal latent value, treating values below the detection limit as censored; this internally handles technical and biological missingness without imputation or transformation. The MRF prior P(gamma) ∝ exp(omega sum gamma_j + eta gamma^T R gamma) uses the relationship matrix R to make selection of related predictors mutually reinforcing, and the new rule for eta keeps the prior away from its phase transition while matching the prior model size to an independent-Bernoulli

Load-bearing premise

The load-bearing assumption is that the prior strength is chosen so the Markov random field prior does not suddenly switch from favoring sparse models to selecting nearly all variables; if that switch happens, the reported gains in detection and the extra diet–metabolite associations could come from the prior rather than from the data.

What would settle it

Simulate null datasets from the fitted model with all diet coefficients set to zero, keeping the same relationship matrix and the same rule for choosing the prior strength, and count how often each food is selected. If the Markov random field prior selects variables above the declared Bayesian FDR, the extra associations are artifacts of the prior. A cheaper check in the real data is to move the prior strength slightly above and below the chosen value: if the number of selected associations jumps discontinuously, the prior is operating near its phase transition.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • Analysts can score all 30 foods jointly for each of hundreds of metabolites instead of testing foods one at a time, which is the practical bottleneck the paper removes.
  • Individually weak dietary effects become detectable when they sit in a tightly connected nutritional cluster; the paper reports that one weak-signal link went from being found under 5% of the time with independent priors to nearly 60% of the time with the MRF prior.
  • A mis-specified relationship matrix did not inflate false discoveries in simulation, so crude or imperfect external groupings still appear usable.
  • Skew-normal errors improved model fit more than the MRF prior did, suggesting that modeling skewness directly is the first-order fix for metabolomic responses.
  • The same response model can be applied to other omics or viral-load outcomes that share censoring, missingness, and skewness.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The eta-selection rule is a matching heuristic rather than a theorem: if the chosen eta happens to sit near a phase transition, posterior model size could jump and the MRF-only associations could become fragile. One could test this by re-running the analysis at eta values just above and below the chosen value and checking whether selected-variable counts and posterior inclusion probabilities move
  • Because R is built from nutritional groupings, an association found only with the MRF prior is evidence that the food's signal is consistent with its nutritional group, not independent proof that that specific food drives the metabolite; a follow-up with an alternative grouping matrix or adjustment for total diet pattern would sharpen the causal reading.
  • The paper analyzes each metabolite separately; jointly modeling multiple metabolite responses could exploit correlation among metabolites to reduce multiplicity and further improve selection, a natural extension not attempted here.
  • The nine MRF-only associations in the two cohorts are externally testable claims: a third cohort with the same food frequency data could check whether they replicate before being treated as true signals.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper proposes a Bayesian variable selection framework for metabolite response variables subject to left-censoring (point mass values), skewness, and high-dimensional correlated predictors. The model is a skew-normal censored mixture (SNCM) with spike-and-slab variable selection via Kuo-Mallick priors, and an MRF prior over inclusion indicators that uses a nutritional relationship matrix R among food predictors. The authors introduce a prior-based rule for selecting the MRF interaction parameter eta, compare the method with independent-Bernoulli-prior selection in simulations, and apply it to 244 metabolites in the MBS and MLVS cohorts, reporting additional food-metabolite associations and improved model fit. The central claim is that the MRF prior improves true positive rates without inflating false discovery rates, and that the SNCM model identifies more dietary predictors than alternatives on real data.

Significance. If the central claims hold, the paper contributes a useful and practical framework for metabolomics variable selection: it combines a censored mixture model with skew-normal errors and a structured prior, and it ships an R package (multimetab). The model components are standard and the MCMC scheme is straightforward, which is a strength for reproducibility. However, the claimed advantage of the MRF prior over independent Bernoulli priors rests on an unvalidated hyperparameter-selection heuristic and on simulations that place true signals inside the R-connected structure. The absence of global-null and eta-sensitivity checks leaves open the possibility that the reported TPR/FDR gains and the extra real-data associations are prior artifacts. These are load-bearing gaps for the paper's main message.

major comments (3)
  1. [Section 2.3.1] The new eta-selection rule is not validated against the phase-transition behavior that the paper itself cites. Choosing the largest eta such that the prior 95th percentile model size of MRF(omega0, eta) is below that of an independent Bernoulli prior with inclusion probability 2*omega0 only matches one prior quantile; it does not establish that the MRF prior is in a subcritical, unimodal regime. Near a critical eta, Ising-type MRF priors can become dominated by either very sparse or very dense configurations, and posterior inclusion probabilities can be driven by the prior. The central TPR/FDR claims in Section 3 therefore depend on a hyperparameter position that is asserted, not demonstrated. Please report, for the R used in simulations and in the real data, the full prior model-size distribution and a diagnostic for phase transition (e.g., bimodality of the prior, or free energy / coup
  2. [Section 3.1 / Table 1] The simulation design systematically places all true signals inside the connected blocks of R. In the baseline setting, true predictors are confined to blocks 1-4, each within a block of 20 connected predictors, so the MRF prior is supplied with correct external information by construction. The 'misspecified dependence' scenario partially addresses mis-specified R, but no global-null simulation is reported: no setting has all gamma_j = 0. Without a global-null check, the claim that MRF 'does not inflate FDRs' is incomplete. Please add global-null simulations under both priors, reporting posterior inclusion probabilities, the number of selected variables at the 5% Bayesian FDR threshold, and the empirical FDR. This is necessary to rule out that the MRF prior increases prior mass toward non-null models.
  3. [Section 4.3] The real-data associations found only under the MRF prior (dairy triacylglycerols, carotenoid vegetables, coffee/caffeine metabolites) all lie in connected groups of R. Because eta is chosen by the heuristic in Section 2.3.1 rather than estimated, the extra associations could be sensitive to the exact eta value. No sensitivity analysis is given for the application. Please report whether these MRF-only associations persist over a plausible range of eta (including values below and above the chosen one), and ideally compare against a null or permuted R. Without such a check, the statement that the nutritionally informed model 'identified more dietary predictors' is not yet supported.
minor comments (6)
  1. [Section 1.1 / Figure 1] The text refers to Figure 1B as correlations and Figure 1C as skewness, but the caption lists (B) as skewness and (C) as pairwise correlations. Please fix the mismatch.
  2. [Section 2.1] In the model formulation, U_i is defined as an indicator that the metabolite is present in the sample, with Y_i^* = U_i V_i. However, the text later says 'samples where the metabolite is present (i.e., those with U_i = 0)'. This should be U_i = 1; the current wording is inconsistent with the likelihood.
  3. [Section 5] The text refers to 'the SVSS struggled' and 'SVSS-style sampling schemes'; the acronym should be SSVS (stochastic search variable selection).
  4. [Section 4.2.2] The sentence 'We derived R was derived from this hierarchical structure' contains a grammatical error; should be 'We derived R from this hierarchical structure.'
  5. [Section 4.2.1] The phrase 'we thinned the samples by 95%' is ambiguous; it likely means retaining only 5% of samples. Please state explicitly.
  6. [Section 3.2] The hyperparameter grid for eta is described as '0.01r^{-1}, 0.02r^{-1}, ..., 1r^{-1}', but the application Section 4.2.1 uses a grid 0.01, 0.02, ..., 1.00. Clarify whether the simulation grid is scaled by 1/max(R) while the real-data grid is not, and whether this affects comparability.

Circularity Check

0 steps flagged

No significant circularity: the MRF/SNCM improvements are empirically demonstrated against an independent-Bernoulli baseline, with a misspecified-R control showing the gain is not forced by the prior.

full rationale

The paper's central claims are not equivalent to their inputs by construction. The MRF prior is a standard external construction (Li and Zhang 2010; Zhao et al. 2024), and the key comparison in Section 3 is between the MRF prior and the nested independent-Bernoulli prior (η=0) under identical simulated data; TPR/FDR are estimated from replicated datasets, not defined by the prior. The closest candidate for circularity is the simulation design, where true βs are placed inside R-connected clusters, so the MRF is tested on data matching its assumptions. But this is a favorable-condition experiment rather than a reduction: Section 3.3/Table 1 includes a 'Misspecified R' scenario in which the TPR gain shrinks from 0.75 vs 0.53 to 0.56 vs 0.53, showing the improvement is not guaranteed by construction. The η-selection rule (Section 2.3.1) is prior-based and data-independent, so it does not fit the response and is not a fitted-input-called-prediction; its possible proximity to a phase transition is a calibration-risk concern, not circularity. Real-data findings are compared against an independent-prior baseline from the same paper, and the R matrix is taken from external nutritional classifications (Satija et al. 2016), not from the metabolite outcomes. The paper's self-citations (e.g., Lee et al. 2017, 2020; Li et al. 2022a,b) are descriptive references to prior MCMC/zero-inflated modeling or data availability, and no load-bearing argument is reduced to them. No equality of equations or fitted-parameter-as-prediction step was found, so no circular step is identified.

Axiom & Free-Parameter Ledger

5 free parameters · 3 axioms · 0 invented entities

The central results rest on the mixture decomposition assumptions for PMVs, the MRF prior structure, and hyperparameters that are either hand-set or estimated from the data in weakly informative ways. No invented physical entities are introduced.

free parameters (5)
  • omega (MRF sparsity) = logit(0.05) in application; logit(0.02) in simulations
    Chosen by hand to target prior inclusion rates; affects sparsity of selection.
  • eta (MRF interaction) = searched 0.01 to 1.00, chosen by new prior-model-size rule
    Controls influence of R; selection rule is a heuristic without theoretical guarantee.
  • nu^2 (slab variance) = empirical variance of single-predictor regression coefficients (application); 4 (simulations)
    Set from data or fixed; influences effect size scale under selection.
  • rho0, rho1 = 5*sqrt(W_bar), 5*(1-sqrt(W_bar))
    Weakly informative prior for biological PMV rate is centered on the observed PMV fraction, using the data twice.
  • psi (detection limit) = min observed value
    Estimated from data; the censoring threshold is not known a priori.
axioms (3)
  • domain assumption PMVs decompose into biological absence (U=0) and technical censoring (U=1, V<psi), with U ~ Bernoulli(rho) independent of V and covariates.
    Section 2.1; the paper acknowledges in Section 5 that rho is assumed constant across individuals, which is a simplification.
  • domain assumption Among present samples, metabolite abundance is linear in predictors and confounders with skew-normal errors.
    Section 2.1; this is the core regression model whose validity underlies all coefficient estimates.
  • domain assumption The MRF prior with R derived from nutritional groupings correctly captures dependence among dietary predictors for joint selection.
    Section 2.2 and 4.2.2; if R is uninformative or misleading, the MRF prior can distort selection, though simulation with permuted R suggests limited harm.

pith-pipeline@v1.3.0-alltime-deepseek · 15768 in / 10790 out tokens · 116667 ms · 2026-08-04T23:05:25.701186+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of A nutritionally informed model for Bayesian variable selection with metabolite response variables." pith.science (2026). https://pith.science/paper/CSCUMM4S

@misc{pith2026250906779,
  author       = {Pith},
  title        = {Pith review of: A nutritionally informed model for Bayesian variable selection with metabolite response variables},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CSCUMM4S}},
  note         = {Machine review of arXiv:2509.06779}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Understanding the pathways through which diet affects human metabolism is a central task in nutritional epidemiology. This article proposes novel methodology to identify food items associated with blood metabolites in two cohorts of healthcare professionals. We analyze 30 food intake variables that exhibit relationship structure through their correlations and nutritional attributes. The metabolic responses include 244 compounds measured by mass spectrometry, presenting substantial challenges that include missingness, left-censoring, and skewness. While existing methods can address such factors in low-dimensional settings, they are not designed for high-dimensional regression involving strongly correlated predictors and non-normal outcomes. To address these challenges, we propose a novel Bayesian variable selection framework for metabolite response variables based on a skew-normal censored mixture model. To exploit substantive information on the nutritional similarities among dietary factors, we employ a Markov random field prior that encourages joint selection of related predictors, while introducing a new, efficient strategy for its hyperparameter specification. Applying this methodology to the cohort data identifies multiple metabolite-diet associations that are consistent with previous research as well as several potentially novel associations that were not detected using standard methods. The proposed approach is implemented in the R package multimetab, facilitating its use in high-dimensional metabolomic analyses.

Figures

Figures reproduced from arXiv: 2509.06779 by Brent A Coull, Dylan Clark-Boucher, Fenglei Wang, Harrison T Reeder, Jacqueline R Starr, Kyu Ha Lee, Qi Sun.

Figure 1
Figure 1. Figure 1: MBS and MLVS data challenges. (A) The point mass value (PMV) percentages of 30 metabolites with at least one PMV in the Mind Body Study (MBS) and Men’s Lifestyle Validation Study (MLVS). (B) The skewness index of all 244 metabolites, where more positive values indicate stronger right skewness. (C) Pairwise correlations among the 30 dietary intake predictor variables after combining the two datasets. We use… view at source ↗
Figure 2
Figure 2. Figure 2: Relationships within a block of 20 predictors in the simulation design [PITH_FULL_IMAGE:figures/full_fig_p013_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Variable-specific TPR among clusters of substantively connected [PITH_FULL_IMAGE:figures/full_fig_p016_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Dietary relationship matrix (R). The figure depicts empirical pairwise rela￾tionships among the 30 dietary intake variables considered for variable selection. Variables were grouped hierarchically based on their nutritional properties. R was incorporated into the variable selection scheme via the Markov random field prior. 4.2.3 Model Validation We evaluated the quality of the model fit in two complementar… view at source ↗
Figure 5
Figure 5. Figure 5: Metabolite-dietary associations identified by SNCM model [PITH_FULL_IMAGE:figures/full_fig_p022_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Posterior predictive distributions from the normal and skew-normal [PITH_FULL_IMAGE:figures/full_fig_p024_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

50 extracted references · 49 canonical work pages

  1. [1]

    Abdullah, M. M. H., Cyr, A., Lépine, M.-C., Labonté, M.-E., Couture, P., Jones, P. J. H., and Lamarche, B. (2015). Recommended dairy product intake modulates circulating fatty acid profile in healthy adults: a multi-centre cross-over study. British Journal of Nutrition , 113(3):435–444

  2. [2]

    Albert, J. H. and Chib, S. (1993). Bayesian analysis of binary and polychotomous response data. Journal of the American Statistical Association , 88(422):669–679

  3. [3]

    Baker, S. A. and Rutter, J. (2023). Metabolites as signalling molecules. Nature Reviews Molecular Cell Biology , 24(5):355–374

  4. [4]

    L., Lenart, E

    Bao, Y., Bertoia, M. L., Lenart, E. B., Stampfer, M. J., Willett, W. C., Speizer, F. E., and Chavarro, J. E. (2016). Origin, methods, and evolution of the three nurses’ health studies. American Journal of Public Health , 106(9):1573–1581

  5. [5]

    S., and De Oliveira, E

    Batista-da Silva, B., Limirio, L. S., and De Oliveira, E. P. (2024). Association between caffeine metabolites in urine and muscle strength in young and older adults: A cross-sectional study from nhanes 2011–2012. Clinical Nutrition , 43(6):1584–1592

  6. [6]

    J., Vannucci, M., and Fearn, T

    Brown, P. J., Vannucci, M., and Fearn, T. (1998). Multivariate bayesian variable selection and prediction. Journal of the Royal Statistical Society Series B: Statistical Methodology , 60(3):627–641

  7. [7]

    and Swann, J

    Caspani, G. and Swann, J. (2019). Small talk: microbial metabolites involved in the signaling from microbiota to brain. Current Opinion in Pharmacology , 48:99–106

  8. [8]

    Dagne, G. A. and Huang, Y. (2013). Bayesian semiparametric mixture tobit models with left censoring, skewness, and covariate measurement errors. Statistics in Medicine , 32(22):3881–3898

  9. [9]

    T., Wahl, S., Raffler, J., Molnos, S., Laimighofer, M., Adamski, J., Suhre, K., Strauch, K., Peters, A., Gieger, C., Langenberg, C., Stewart, I

    Do, K. T., Wahl, S., Raffler, J., Molnos, S., Laimighofer, M., Adamski, J., Suhre, K., Strauch, K., Peters, A., Gieger, C., Langenberg, C., Stewart, I. D., Theis, F. J., Grallert, H., Kastenmüller, G., and Krumsiek, J. (2018). Characterization of missing values in untargeted ms-based metabolomics data and evaluation of missing data handling strategies. Me...

  10. [10]

    E., Dey, D

    Gelfand, A. E., Dey, D. K., and Chang, H. (1992). Model Determination using Predictive Distributions with Implementation via Sampling-Based Methods , page 0. Oxford University Press

  11. [11]

    C., Carpenter, B., Yao, Y., Kennedy, L., Gabry, J., Bürkner, P.-C., and Modrák, M

    Gelman, A., Vehtari, A., Simpson, D., Margossian, C. C., Carpenter, B., Yao, Y., Kennedy, L., Gabry, J., Bürkner, P.-C., and Modrák, M. (2020). Bayesian workflow. (arXiv:2011.01808)

  12. [12]

    George, E. I. and McCulloch, R. E. (1997). Approaches for bayesian variable selection. Statistica Sinica , 7(2):339–373

  13. [13]

    Gleiss, A., Dakna, M., Mischak, H., and Heinze, G. (2015). Two-group comparisons of zero-inflated intensity values: the choice of test statistic matters. Bioinformatics , 31(14):2310–2317

  14. [14]

    B., Macklaim, J

    Gloor, G. B., Macklaim, J. M., Pawlowsky-Glahn, V., and Egozcue, J. J. (2017). Microbiome datasets are compositional: And this is not optional. Frontiers in Microbiology , 8:2224

  15. [15]

    D., Sampson, L., Barnett, J

    Gu, X., Wang, D. D., Sampson, L., Barnett, J. B., Rimm, E. B., Stampfer, M. J., Djousse, L., Rosner, B., and Willett, W. C. (2024). Validity and reproducibility of a semiquantitative food frequency questionnaire for measuring intakes of foods and food groups. American Journal of Epidemiology , 193(1):170–179

  16. [16]

    C., Huang, T., Manson, J

    Hamaya, R., Sun, Q., Li, J., Yun, H., Wang, F., Curhan, G. C., Huang, T., Manson, J. E., Willett, W. C., Rimm, E. B., Clish, C., Liang, L., Hu, F. B., and Ma, Y. (2024). 24-h urinary sodium and potassium excretions, plasma metabolomic profiles, and cardiometabolic biomarkers in the united states adults: a cross-sectional study. The American Journal of Cli...

  17. [17]

    and Viant, M

    Hrydziuszko, O. and Viant, M. R. (2012). Missing values in mass spectrometry based metabolomics: an undervalued step in the data processing pipeline. Metabolomics , 8(S1):161–174

  18. [18]

    M., Sawyer, S., Kubzansky, L

    Huang, T., Trudel-Fitzgerald, C., Poole, E. M., Sawyer, S., Kubzansky, L. D., Hankinson, S. E., Okereke, O. I., and Tworoger, S. S. (2019). The mind–body study: study design and reproducibility and interrelationships of psychosocial factors in the nurses’ health study ii. Cancer Causes & Control , 30(7):779–790

  19. [19]

    N., Fan, T

    Huang, Z., Lane, A. N., Fan, T. W.-M., Higashi, R. M., Weiss, H. L., Yin, X., and Wang, C. (2020). Differential abundance analysis with bayes shrinkage estimation of variance (dasev) for zero-inflated proteomic and metabolomic data. Scientific Reports , 10(1):876

  20. [20]

    and Wang, C

    Huang, Z. and Wang, C. (2022). A review on differential abundance analysis methods for mass spectrometry-based metabolomic data. Metabolites , 12(4):305

  21. [21]

    M., Mueller, N

    Jia, Y., Yang, X., Wilson, L. M., Mueller, N. T., Sears, C. L., Treisman, G. J., and Robinson, K. A. (2022). Diet-related and gut-derived metabolites and health outcomes: A scoping review. Metabolites , 12(12):1261

  22. [22]

    S., Huang, T., Chan, A

    Ke, S., Guimond, A.-J., Tworoger, S. S., Huang, T., Chan, A. T., Liu, Y.-Y., and Kubzansky, L. D. (2023). Gut feelings: associations of emotions and emotion regulation with the gut microbiome in women. Psychological Medicine , 53(15):7151–7160

  23. [23]

    and Mallick, B

    Kuo, L. and Mallick, B. (1998). Variable selection for regression models. Sankhyā: The Indian Journal of Statistics, Series B (1960-2002) , 60(1):65–81

  24. [24]

    H., Coull, B

    Lee, K. H., Coull, B. A., Moscicki, A.-B., Paster, B. J., and Starr, J. R. (2020). Bayesian variable selection for multivariate zero-inflated models: Application to microbiome count data. Biostatistics , 21(3):499–517

  25. [25]

    H., Tadesse, M

    Lee, K. H., Tadesse, M. G., Baccarelli, A. A., Schwartz, J., and Coull, B. A. (2017). Multivariate bayesian variable selection exploiting dependence structure among outcomes: Application to air pollution effects on dna methylation: Multivariate bayesian variable selection. Biometrics , 73(1):232–241

  26. [26]

    and Zhang, N

    Li, F. and Zhang, N. R. (2010). Bayesian variable selection in structured high-dimensional covariate spaces with applications in genomics. Journal of the American Statistical Association , 105(491):1202–1214

  27. [27]

    L., Wang, D

    Li, J., Li, Y., Ivey, K. L., Wang, D. D., Wilkinson, J. E., Franke, A., Lee, K. H., Chan, A., Huttenhower, C., Hu, F. B., Rimm, E. B., and Sun, Q. (2022a). Interplay between diet and gut microbiome, and circulating concentrations of trimethylamine n-oxide: findings from a longitudinal cohort of us men. Gut , 71(4):724–733

  28. [28]

    D., Satija, A., Ivey, K

    Li, Y., Wang, D. D., Satija, A., Ivey, K. L., Li, J., Wilkinson, J. E., Li, R., Baden, M., Chan, A. T., Huttenhower, C., Rimm, E. B., Hu, F. B., and Sun, Q. (2021). Plant-based diet index and metabolic risk in men: Exploring the role of the gut microbiome. The Journal of Nutrition , 151(9):2780–2789

  29. [29]

    L., Wilkinson, J

    Li, Y., Wang, F., Li, J., Ivey, K. L., Wilkinson, J. E., Wang, D. D., Li, R., Liu, G., Eliassen, H. A., Chan, A. T., Clish, C. B., Huttenhower, C., Hu, F. B., Sun, Q., and Rimm, E. B. (2022b). Dietary lignans, plasma enterolactone levels, and metabolic risk in men: exploring the role of the gut microbiome. BMC Microbiology , 22(1):82

  30. [30]

    S., Guasch-Ferre, M., Hu, F

    Malik, V. S., Guasch-Ferre, M., Hu, F. B., Townsend, M. K., Zeleznik, O. A., Eliassen, A. H., Tworoger, S. S., Karlson, E. W., Costenbader, K. H., Ascherio, A., Wilson, K. M., Mucci, L. A., Giovannucci, E. L., Fuchs, C. S., and Bao, Y. (2019). Identification of plasma lipid metabolites associated with nut consumption in us men and women. The Journal of Nu...

  31. [31]

    N., Franzosa, E

    Manghi, P., Bhosle, A., Wang, K., Marconi, R., Selma-Royo, M., Ricci, L., Asnicar, F., Golzato, D., Ma, W., Hang, D., Thompson, K. N., Franzosa, E. A., Nabinejad, A., Tamburini, S., Rimm, E. B., Garrett, W. S., Sun, Q., Chan, A. T., Valles-Colomer, M., Arumugam, M., Bermingham, K. M., Giordano, F., Davies, R., Hadjigeorgiou, G., Wolf, J., Strowig, T., Ber...

  32. [32]

    A., Noueiry, A., Deepayan, S., and Paul, A

    Newton, M. A., Noueiry, A., Deepayan, S., and Paul, A. (2004). Detecting differential gene expression with a semiparametric hierarchical mixture method. Biostatistics , 5(2):155–176

  33. [33]

    Rimm, E., Giovannucci, E., Willett, W., Colditz, G., Ascherio, A., Rosner, B., and Stampfer, M. (1991). Prospective study of alcohol consumption and risk of coronary disease in men. The Lancet , 338(8765):464–468

  34. [34]

    K., Dey, D

    Sahu, S. K., Dey, D. K., and Branco, M. D. (2003). A new class of multivariate skew distributions with applications to bayesian regression models. Canadian Journal of Statistics , 31(2):129–150

  35. [35]

    N., Rimm, E

    Satija, A., Bhupathiraju, S. N., Rimm, E. B., Spiegelman, D., Chiuve, S. E., Borgi, L., Willett, W. C., Manson, J. E., Sun, Q., and Hu, F. B. (2016). Plant-based dietary patterns and incidence of type 2 diabetes in us men and women: Results from three prospective cohort studies. PLOS Medicine , 13(6):e1002039

  36. [36]

    E., Rogowska-Wrzesinska, A., and Jensen, O

    Schwämmle, V., Hagensen, C. E., Rogowska-Wrzesinska, A., and Jensen, O. N. (2020). Polystest: Robust statistical testing of proteomics data with missing values improves detection of biologically relevant features. Molecular & Cellular Proteomics , 19(8):1396–1408

  37. [37]

    N., and Gaskins, J

    Shah, J., Brock, G. N., and Gaskins, J. (2019). Bayesmetab: treatment of missing values in metabolomic studies using a bayesian modeling approach. BMC Bioinformatics , 20(S24):673

  38. [38]

    C., Chen, Y

    Stingo, F. C., Chen, Y. A., Tadesse, M. G., and Vannucci, M. (2011). Incorporating biological information into linear models: A bayesian approach to the selection of pathways and genes. The Annals of Applied Statistics , 5(3)

  39. [39]

    L., Leiserowitz, G

    Taylor, S. L., Leiserowitz, G. S., and Kim, K. (2013). Accounting for undetected compounds in statistical analyses of mass spectrometry ‘omic studies. Statistical Applications in Genetics and Molecular Biology , 12(6)

  40. [40]

    Tsikas, D. (2023). Homoarginine in health and disease. Current Opinion in Clinical Nutrition & Metabolic Care , 26(1):42–49

  41. [41]

    J., Cohen, N

    Vauzour, D., Scholey, A., White, D. J., Cohen, N. J., Cassidy, A., Gillings, R., Irvine, M. A., Kay, C. D., Kim, M., King, R., Legido-Quigley, C., Potter, J. F., Schwarb, H., and Minihane, A.-M. (2023). A combined dha-rich fish oil and cocoa flavanols intervention does not improve cognition or brain structure in older adults with memory complaints: result...

  42. [42]

    D., Argiento, R., Guindani, M., Galloway-Pena, J., Shelburne, S

    Wadsworth, W. D., Argiento, R., Guindani, M., Galloway-Pena, J., Shelburne, S. A., and Vannucci, M. (2017). An integrative bayesian dirichlet-multinomial regression model for the analysis of taxonomic abundances in microbiome data. BMC Bioinformatics , 18(1):94

  43. [43]

    D., Nguyen, L

    Wang, D. D., Nguyen, L. H., Li, Y., Yan, Y., Ma, W., Rinott, E., Ivey, K. L., Shai, I., Willett, W. C., Hu, F. B., Rimm, E. B., Stampfer, M. J., Chan, A. T., and Huttenhower, C. (2021). The gut microbiome modulates the protective association between a mediterranean diet and cardiometabolic disease risk. Nature Medicine , 27(2):333–343

  44. [44]

    Watanabe, S. (2010). Asymptotic equivalence of bayes cross validation and widely applicable information criterion in singular learning theory. Journal of Machine Learning Research , 11:3571–3594

  45. [45]

    Wei, R., Wang, J., Su, M., Jia, E., Chen, S., Chen, T., and Ni, Y. (2018). Missing value imputation approach for mass spectrometry-based metabolomics data. Scientific Reports , 8(1):663

  46. [46]

    C., Sampson, L., Stampfer, M

    Willett, W. C., Sampson, L., Stampfer, M. J., Rosner, B., Bain, C., Witschi, J., Hennekens, C. H., and Speizer, F. E. (1985). Reproducibility and reliability of a semiquantitative food frequency questionnaire. American Journal of Epidemiology , 122(1):51–65

  47. [47]

    Wilson, T., Loughran, T., and Brame, R. (2020). Substantial bias in the tobit estimator: Making a case for alternatives. Justice Quarterly , 37(2):231–257

  48. [48]

    B., Rosner, B

    Yuan, C., Spiegelman, D., Rimm, E. B., Rosner, B. A., Stampfer, M. J., Barnett, J. B., Chavarro, J. E., Rood, J. C., Harnack, L. J., Sampson, L. K., and Willett, W. C. (2018). Relative validity of nutrient intakes assessed by questionnaire, 24-hour recalls, and diet records as compared with urinary recovery and plasma concentration biomarkers: Findings fo...

  49. [49]

    R., Do, K., and Peterson, C

    Zhang, L., Shi, Y., Jenq, R. R., Do, K., and Peterson, C. B. (2021). Bayesian compositional regression with structured priors for microbiome feature selection. Biometrics , 77(3):824–838

  50. [50]

    Zhao, Z., Banterle, M., Lewin, A., and Zucknick, M. (2024). Multivariate bayesian structured variable selection for pharmacogenomic studies. Journal of the Royal Statistical Society Series C: Applied Statistics , 73(2):420–443