Pith. sign in

REVIEW 3 major objections 4 minor 38 references

CoLaDAG recovers directed dependence in compositional microbiome counts better than generic graph learners under its aligned simulations, but on real data its output is a reference-dependent ranked hypothesis list, not a causal map.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 00:00 UTC pith:NRF53DVF

load-bearing objection Honest, well-hedged workflow paper: the integration is genuinely useful, the stress tests are unusually candid, but the headline exact-direction metric overstates what is identifiable and the main benchmark is partly stacked. the 3 major comments →

arxiv 2607.23233 v1 pith:NRF53DVF submitted 2026-07-25 stat.AP

CoLaDAG: Compositional Latent Log-ratio DAG Analysis of the Gut Microbiome under Silver Nanoparticle Exposure

classification stat.AP MSC 62H2262P10
keywords compositional data analysisdirected acyclic graphsadditive log-ratio transformationmicrobiome network inferencemultinomial count modelsparse structural equation modelssilver nanoparticle exposurestability selection
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

CoLaDAG aims to turn compositional sequencing counts into sparse directed hypotheses about which gut microbes depend conditionally on which others, without claiming the arrows are identified causal effects. The central claim is that a fixed-reference latent additive log-ratio (ALR) model — log-ratios of each taxon to one chosen reference, wrapped in a multinomial observation layer for read counts and a sparse Gaussian structural equation model with acyclicity constraints, finished by hard thresholding and greedy DAG projection — recovers directed dependence better than the evaluated generic graph learners under simulations generated from that same model class. The paper is equally explicit about the boundaries of its claim: log-ratio transforms densify sparse absolute-scale graphs, orientation is not identifiable from observational Gaussian data, the advantage reverses under continuous-data misspecification, and all methods fall to chance once 10–30% of nonzero counts are dropped. On the 12-mouse silver-nanoparticle data, the defensible output is a ranked follow-up list — 60 of 284 edges survive mouse-block resampling at 0.60 selection frequency and 37 also keep their sign — which matters because it gives toxicology a principled bridge from univariate differential abundance and undirected co-occurrence to directional, testable log-ratio dependence hypotheses.

Core claim

On its own terms, CoLaDAG establishes that directed conditional-dependence can be recovered from compositional counts by jointly modeling multinomial read sampling, latent fixed-reference ALR coordinates, and a sparse acyclic linear Gaussian SEM among them, solved by DC-ADMM on a truncated-L1 acyclicity relaxation and finished by hard thresholding with greedy acyclic projection. In aligned simulations (30 nodes, 500 samples, depth 10,000) it reports the largest mean exact-direction Matthews correlation (0.578) and smallest mean false-discovery rate (0.336) among the six compared implementations. The paper also sets the limits of its object: log-ratio transforms densify sparse absolute-scale

What carries the argument

The load-bearing object is the fixed-reference latent additive log-ratio (ALR) coordinate system, Z_j = log(pi_j/pi_r), carried through a joint objective that couples a multinomial count likelihood for the observed reads with a Gaussian SEM regularizer on the latent coordinates. The mechanism the paper stresses is densification: because every ALR coordinate shares the reference term, sparse dependence on the absolute scale becomes dense in log-ratio space, so sparsity must be re-imposed explicitly — hence the hard threshold followed by greedy acyclic projection that converts the continuous nonconvex solution into a sparse DAG. Acyclicity is enforced by a dual-constraint formulation relaxed t

Load-bearing premise

The load-bearing premise is that the count data actually follow the paper's multinomial-latent model — no dropout or zero-inflation, and repeated samples from the same mouse conditionally independent given latent ALR coordinates — which the paper itself calls 'a substantial limitation' (Section 2, with effective n closer to 12 mice than 48 samples), and which its own supplement shows to be decisive (Table S6: every method at chance after 10–30% of nonzero counts are dropped).

What would settle it

Count the empirical zero rate in the AgNP data (or any target dataset) and compare with the dropout regime in the paper's own Table S6: as nonzero-count dropout rises from 0% to 10%, CoLaDAG's mean exact-direction Matthews correlation falls from 0.477 to −0.008, i.e., chance. A direct experiment: sequence a synthetic microbial community with known composition at realistic depth to induce typical zero patterns, run CoLaDAG, and check whether directed recovery reproduces the aligned-simulation advantage; if the observed dropout exceeds about 10%, the paper's own results predict it will not.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Under its aligned model class — multinomial counts drawn from latent ALR compositions with no dropout — CoLaDAG recovers directed edges with better exact-direction agreement and lower false-discovery rate than the five generic DAG learners evaluated, making it a candidate end-to-end screening pipeline for compositional count data.
  • The operating range is narrow and explicitly mapped: for directly observed continuous SEM data a score-based baseline (HC-Huge) is substantially better, and at 10–30% random dropout of nonzero counts all evaluated methods fall to near-chance recovery — so CoLaDAG should only be deployed where zero-inflation is known to be low.
  • In the 12-mouse silver-nanoparticle study, the defensible output is a ranked hypothesis list: 60 of 284 fitted edges reach 0.60 mouse-block selection frequency, 37 also keep sign agreement, and the leading relations (e.g., Intestinimonas to Lachnospiraceae AC2044 group) are candidates for targeted abundance, metabolite, and perturbation studies — not established interactions.
  • The fitted graph is materially reference-dependent: refitting with different ALR denominators changes the graph from 284 to 684 directed edges, with 92 skeleton edges common to all five references, so skeleton-level overlap across references is a more robust summary than any single directed orientation.
  • The dose-stratified displays are time-adjusted pairwise ALR slopes on a restricted global support — descriptive follow-up candidates, not dose-response or exposure-effect estimates, given three mice per exposure group.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The paper's own dropout stress test implies a deployment rule it leaves implicit: any dataset meant for CoLaDAG should first pass a documented zero-rate audit, because per-taxon zero proportions above roughly 10% put the analysis outside the regime where the method was shown to work.
  • A natural extension the authors do not pursue is a multi-reference consensus estimator: running CoLaDAG under several ALR denominators and keeping only recurring skeleton edges (they find 92 such edges across five references) would produce a conservative core-hypothesis set that sidesteps reference-dependence.
  • Because the effective sample size is 12 mice rather than 48 samples, the minimal upgrade from screening to confirmation is a mouse-level random-effects or block-conditional likelihood; the paper's block resampling is descriptive and does not repair the missing mouse layer.
  • If the 60 stable edges survive targeted quantification in a larger cohort, the method would give environmental toxicology a reusable template for converting 16S count tables into ranked, falsifiable microbial-dependence hypotheses — the paper frames its value as the ranked validation plan, not the graph itself.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper introduces CoLaDAG, a procedure for estimating sparse directed conditional-dependence graphs from compositional count data. It models observed counts as multinomial with latent additive log-ratio (ALR) coordinates and fits a zero-intercept linear Gaussian SEM on these latent coordinates using penalized DC-ADMM optimization with acyclicity constraints, followed by hard thresholding and greedy acyclic projection. Simulation studies compare CoLaDAG with PC, MMHC, HC-Huge, Tabu-Huge, and a VI-NOTEARS-style method under an aligned multinomial-latent generative model and under misspecified continuous-SEM and dropout settings. The paper reports that CoLaDAG achieves the largest exact-direction MCC and lowest FDR in the aligned setting, but not universally. A real-data application to a 12-mouse, four-group silver-nanoparticle experiment yields a 58-node genus-level graph with 284 edges; the authors explicitly treat this graph as an exploratory, coordinate-specific hypothesis list, supported by mouse-block stability, reference sensitivity, and dose-stratified descriptive slopes.

Significance. If taken at face value, the paper provides a well-documented workflow for generating directed log-ratio hypotheses from compositional counts under a specific model class. Its strengths are the multinomial observation layer, explicit thresholding and acyclic projection, extensive robustness/ablation checks, mouse-block resampling, and unusually candid disclosure of limitations (non-identifiability, small effective sample size, reference dependence, nonconvex optimization, reproducibility gaps). The paper does not overclaim causal identification. However, the headline evaluation metrics (exact-direction MCC/FDR) are not aligned with the stated identifiability limitations, and the main simulation benchmark conditions baselines on CoLaDAG's latent recovery, so the practical significance of the claimed advantage is not yet established.

major comments (3)
  1. [Abstract; §3.3; §4; Table 1] The abstract and Table 1 headline exact-direction MCC and FDR, but the paper's own model assumptions rule out orientation identification. §3.3 (Eq. 1) states that in a zero-intercept linear Gaussian SEM, 'true edge orientations cannot be strictly identified without additional assumptions ... which we do not impose,' and §4 states that CPDAG and skeleton results are 'the more defensible structural summaries.' Exact-direction MCC and FDR treat an edge with a reversed orientation as wrong, which is not justified when the true orientation is unidentifiable from the observational likelihood. The large MCC gap (0.578 vs 0.374) may therefore reflect algorithmic orientation bias rather than better recovery. Please either report the primary comparisons at the CPDAG/skeleton level (as the paper itself calls more defensible), or impose and empirically validate an orientation-identifiability assumpt
  2. [§4 (Simulation Study)] The benchmark is not an end-to-end comparison: PC, MMHC, HC-Huge, and Tabu-Huge all receive the same latent ALR matrix produced by CoLaDAG's preliminary recovery wrapper, while only the VI-NOTEARS-style method estimates its own latent representation. The paper acknowledges this, but the central claim 'CoLaDAG obtained the largest mean exact-direction MCC' is therefore conditional on a latent representation estimated by CoLaDAG's own pipeline. It is possible that the wrapper, not the DAG estimator, drives the advantage. Please add an end-to-end version of the comparison in which baseline methods use their own natural compositional inputs (e.g., observed pseudocount ALR, CLR, or a neutral multinomial latent estimator), and report whether the MCC/SHD rankings change.
  3. [§5 (Real Data); Table 1; S3.1/S3.4] The simulation study never covers the regime of the real-data case study. Table 1 uses n=500, d=30; the small-n sensitivity in Table S4 uses d=20 and n=100; the dimension experiment in Table S8 reaches d=50 but still n=500. The real application has 48 samples from 12 mice and d=58 nodes. The paper's own stability diagnostic (Table 6) shows that the 284-edge primary graph is highly sensitive to block resampling, and the text calls the independence treatment 'a substantial limitation.' Given that the proposed method is showcased on exactly this small-n, high-dimensional setting, a simulation at approximately (n=48, d=58) with mouse-block dependence and dropout would be necessary to know what performance to expect in the applied regime. Without it, the real-data graph remains very difficult to interpret beyond an exploratory exercise.
minor comments (4)
  1. [Fig. 3] Figure 3 uses the label 'clrdag' while the text and tables use 'CoLaDAG'; please make the legend consistent.
  2. [§5.2] The term 'conditional sign agreement' is used without a definition in the main text. Please define it explicitly (e.g., the proportion of bootstrap fits in which an edge selected by the bootstrap has the same sign as in the primary fit, conditional on selection).
  3. [§3.3] The zero-intercept, non-centered working SEM is unusual; because ALR coordinates have no natural origin, please state explicitly why centering is omitted and discuss the effect on coefficient interpretation. Currently this appears only as a listed limitation in §6.
  4. [S2.5] The reproducibility section notes that the archive does not currently record the complete contributed-package version set and that the latest driver status 'verified existing' does not claim a clean rerun. Please add a lockfile and rerun logs, or clearly state which portions were actually re-executed.

Circularity Check

0 steps flagged

No significant circularity: the headline simulation result is honestly scoped to the paper's own model class and is paired with misspecified stress tests; no fitted quantity is renamed as a prediction.

full rationale

CoLaDAG's derivation chain is self-contained rather than circular. The estimator is defined by a fixed-reference ALR multinomial observation model, a zero-intercept Gaussian SEM, DC-ADMM with a truncated-L1 relaxed acyclicity constraint, hard thresholding, and greedy acyclic projection (Secs. 3.3–3.5). The headline result is not a scientific prediction from data: it is a finite-sample simulation comparison, and the paper explicitly qualifies it as applying "under simulations aligned with this observation model" and states that "these results describe finite-sample behavior under the stated generating mechanism; they do not establish orientation identifiability or uniform superiority." The aligned simulation is generated from the same latent-multinomial construction CoLaDAG assumes, but this is a standard model-specific operating-characteristic evaluation, not a reduction of the output to the input; moreover, the paper includes unaligned continuous-SEM stress tests where HC-Huge wins and random-dropout tests where all methods fail, providing external checks on the claim. The real-data graph is explicitly labeled exploratory (60 of 284 primary edges at mouse-block frequency 0.60, strong reference sensitivity), and the dose slopes are called a "biological follow-up list rather than as significant exposure-specific effects." The self-citations to Yuan et al. (2019) and Shen et al. (2012) supply standard constrained-likelihood and truncated-L1 machinery; they are not used as an unverified uniqueness theorem and are not load-bearing in a way that makes any result true by construction. The Markov-equivalence caveat about exact-direction MCC is a metric-interpretation and correctness limitation, not circularity: it does not make the measured MCC equal to the fitted threshold or to the simulation inputs. The paper also candidly flags its strongest limitations (conditional-independence working likelihood, effective n closer to 12 mice, nonconvex optimization with no global-optimum claim), and those admissions further reduce any circularity risk because the claims are not over-extended. Overall, no step was found where a "prediction" is by definition the fitted parameter or where the argument closes through a self-citation chain.

Axiom & Free-Parameter Ledger

6 free parameters · 6 axioms · 0 invented entities

The central claim depends on the multinomial-latent-ALR construction and on strong working approximations (independence across repeated samples, no dropout, zero-intercept SEM). All are stated, but none are externally validated on the real data. The method hyperparameters are tuning choices, not quantities fixed by theory.

free parameters (6)
  • truncated-L1 penalty tau = sim: recovery 0.10 / refinement 0.25; real: recovery 0.10 / refinement 0.08
    Controls the surrogate for the edge indicator in the acyclicity constraint; chosen by hand/tuning, not data-derived.
  • sparsity penalty mu = sim: recovery 4 / refinement 1; real: recovery 2 / refinement 1
    Strength of the sparse penalty on U; selected on separate tuning simulations/operational choice.
  • ADMM penalty rho = 1.5
    ADMM step-size hyperparameter for the graph solver; chosen by hand.
  • final edge threshold eta = sim: 0.30; real: 0.08
    Hard threshold after DC-ADMM before greedy DAG projection; presented as an operational screening choice, not error-controlling.
  • pseudocount 0.5 = 0.5
    Used to initialize latent ALR coordinates and to build baseline inputs; because optimization is nonconvex it can influence the fitted solution.
  • DC/ADMM iteration caps = e.g., 3/20 and 10/50 in simulations; 5/20 and 10/50 in primary real-data fit; reduced 3/15 and 6/30 in bootstrap diagnos
    Fixed computational caps; reaching a cap is not a convergence certificate and can affect outputs.
axioms (6)
  • domain assumption Observed counts follow a multinomial distribution conditional on an unobserved composition, X_i | pi_i ~ Multinomial(M_i, pi_i).
    Invoked in §3.1/Equation before (2). If sequencing introduces dropout, zero-inflation, or batch effects, this layer is misspecified.
  • ad hoc to paper The latent ALR coordinates follow a zero-intercept linear Gaussian SEM, Z_j = sum U_kj Z_k + eps_j with Gaussian errors.
    Called a 'working SEM' in §3.3; the zero-intercept and no-centering are specific modeling restrictions that can make coefficients absorb location as well as dependence.
  • domain assumption Samples are conditionally independent given latent ALR representations, ignoring mouse identity and time.
    Stated in §2 as 'a substantial limitation'; effective n is 12 mice, not 48 samples.
  • domain assumption No dropout/zero-inflation component in the observation model.
    Supp. S3.2: all methods near chance after 10-30% dropout; the paper notes the model has no structural-zero or zero-inflation component.
  • standard math The dual-constraint characterization of acyclicity (Eq. 3) is a correct characterization of DAGs.
    Taken from Yuan et al. (2019) and used without proof; with the nonnegativity of Lambda it rules out directed cycles.
  • ad hoc to paper Hard thresholding followed by greedy acyclic projection recovers a sparse, interpretable ALR-coordinate DAG after log-ratio densification.
    §3.5 introduces thresholding to counteract densification; the threshold level is an operational choice and its optimality is not derived.

pith-pipeline@v1.3.0-alltime-deepseek · 22532 in / 16475 out tokens · 138101 ms · 2026-08-01T00:00:58.495106+00:00 · methodology

0 comments
read the original abstract

Directed network analysis of microbiome counts is complicated by compositional sampling, high dimensionality, and limited biological replication. We present CoLaDAG, a fixed-reference latent additive log-ratio (ALR) estimator for generating sparse directed conditional-dependence hypotheses from compositional counts. The method combines a multinomial observation model, a working linear Gaussian structural equation model, nonconvex DC-ADMM optimization, and post-estimation thresholding with greedy acyclic projection. Under simulations aligned with this observation model, CoLaDAG obtained the largest mean exact-direction Matthews correlation and the smallest mean false discovery rate among the evaluated implementations; performance deteriorated under continuous-data and dropout misspecification. In the 12-mouse silver-nanoparticle (AgNP) case study, the 58-node fitted graph was sensitive to block resampling and ALR reference choice: 60 of 284 primary edges attained a mouse-block selection frequency of at least 0.60. The reported orientations and dose-stratified slopes are exploratory, coordinate-specific hypotheses rather than identified causal or exposure effects. The leading stable relations prioritize anaerobic gut taxa for targeted abundance, metabolite, and perturbation studies, but do not establish cross-feeding or toxicological mechanisms.

Figures

Figures reproduced from arXiv: 2607.23233 by Shuyan Chen, Xinlei Wang, Ziliang Shen.

Figure 1
Figure 1. Figure 1: AgNP exposure and gut microbiome disruption. 1.1 Related work High-throughput 16S ribosomal RNA (rRNA) sequencing produces count vectors whose totals are governed by sequencing depth rather than absolute microbial load. This places microbiome ob￾servations in the compositional-data setting introduced by Aitchison (1982); only ratios among components are directly interpretable. The broader compositional-dat… view at source ↗
Figure 2
Figure 2. Figure 2: summarizes the analysis-level structure of CoLaDAG. The key distinction is that sequencing counts are not first converted into an ordinary Euclidean data matrix and then handed to a generic graph learner. Instead, the count layer, latent log-ratio coordinates, sparse acyclic SEM, explicit truncation, and stability summaries are treated as connected parts of one compositional graph-estimation procedure. Seq… view at source ↗
Figure 3
Figure 3. Figure 3: Main simulation performance comparisons. 10 [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Time-adjusted pairwise ALR slopes for the retained dose-overlay relations. Slopes are descriptive and preserve the global orientation only as a labeling convention. 15 [PITH_FULL_IMAGE:figures/full_fig_p015_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Thresholded exposure-group displays of time-adjusted pairwise ALR slopes on restricted global CoLaDAG candidate orientations. 16 [PITH_FULL_IMAGE:figures/full_fig_p016_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

38 extracted references

  1. [1]

    Journal of the American Statistical Association , volume=

    Compositional graphical lasso resolves the impact of parasitic infection on gut microbial interaction networks in a zebrafish model , author=. Journal of the American Statistical Association , volume=. 2023 , publisher=

  2. [2]

    ACS nano , volume=

    Changes in gut microbiota structure: a potential pathway for silver nanoparticles to affect the host metabolism , author=. ACS nano , volume=. 2022 , doi=

  3. [3]

    Environmental Science & Technology , volume=

    Nanoparticle silver released into water from commercially available sock fabrics , author=. Environmental Science & Technology , volume=. 2008 , doi=

  4. [4]

    Environmental Science & Technology , volume=

    Environmental transformations of silver nanoparticles: Impact on stability and toxicity , author=. Environmental Science & Technology , volume=. 2012 , doi=

  5. [5]

    Environmental Science & Technology , volume=

    Behavior of metallic silver nanoparticles in a pilot wastewater treatment plant , author=. Environmental Science & Technology , volume=. 2011 , doi=

  6. [6]

    and Nowack, Bernd , journal=

    Gottschalk, Fadri and Sonderer, Tobias and Scholz, Roland W. and Nowack, Bernd , journal=. Modeled environmental concentrations of engineered nanomaterials (. 2009 , doi=

  7. [7]

    Water Research , volume=

    The inhibitory effects of silver nanoparticles, silver ions, and silver chloride colloids on microbial growth , author=. Water Research , volume=. 2008 , doi=

  8. [8]

    Particle and Fibre Toxicology , volume=

    Dietary silver nanoparticles can disturb the gut microbiota in mice , author=. Particle and Fibre Toxicology , volume=. 2016 , doi=

  9. [9]

    and Gokulan, Kuppan and Cerniglia, Carl E

    Williams, Kimberly and Milner, Justin and Boudreau, Michael D. and Gokulan, Kuppan and Cerniglia, Carl E. and Khare, Salman , journal=. Effects of subchronic exposure of silver nanoparticles on intestinal microbiota and gut-associated immune responses in the ileum of. 2015 , doi=

  10. [10]

    Scientific Reports , volume=

    Gut dysbiosis and neurobehavioral alterations in rats exposed to silver nanoparticles , author=. Scientific Reports , volume=. 2017 , doi=

  11. [11]

    The effects of orally administered

    Chen, Hui and Zhao, Ruixin and Wang, Baoyan and Cai, Caixia and Zheng, Lulu and Wang, Hao and Wang, Mingjun and Ouyang, Hongsheng and Zhou, Xue and Chai, Zhan and others , journal=. The effects of orally administered. 2017 , doi=

  12. [12]

    2012 , publisher=

    Introduction to graphical modelling , author=. 2012 , publisher=

  13. [13]

    Biometrika , volume=

    Constrained likelihood for reconstructing a directed acyclic Gaussian graph , author=. Biometrika , volume=. 2019 , publisher=

  14. [14]

    Journal of the American Statistical Association , volume=

    Likelihood-based selection and sharp parameter estimation , author=. Journal of the American Statistical Association , volume=. 2012 , publisher=

  15. [15]

    Journal of the American Statistical Association , volume=

    Likelihood Ratio Tests for a Large Directed Acyclic Graph , author=. Journal of the American Statistical Association , volume=. 2020 , publisher=

  16. [16]

    Journal of the Royal Statistical Society: Series D (The Statistician) , volume=

    The multinomial-Poisson transformation , author=. Journal of the Royal Statistical Society: Series D (The Statistician) , volume=. 1994 , publisher=

  17. [17]

    Proceedings of the 36th International Conference on Machine Learning (ICML) , pages=

    Variational inference for sparse network reconstruction from count data , author=. Proceedings of the 36th International Conference on Machine Learning (ICML) , pages=. 2019 , organization=

  18. [18]

    Journal of the American Statistical Association , volume=

    Variational inference: A review for statisticians , author=. Journal of the American Statistical Association , volume=. 2017 , publisher=

  19. [19]

    Journal of the Royal Statistical Society: Series B (Methodological) , volume=

    The statistical analysis of compositional data , author=. Journal of the Royal Statistical Society: Series B (Methodological) , volume=. 1982 , publisher=

  20. [20]

    Econometrica: journal of the Econometric Society , pages=

    Consistent estimates based on partially consistent observations , author=. Econometrica: journal of the Econometric Society , pages=. 1948 , publisher=

  21. [21]

    Advances in Neural Information Processing Systems , volume=

    DAGs with NO TEARS: Continuous optimization for structure learning , author=. Advances in Neural Information Processing Systems , volume=

  22. [22]

    Advances in Neural Information Processing Systems , volume=

    Differentiable causal discovery from interventional data , author=. Advances in Neural Information Processing Systems , volume=

  23. [23]

    2000 , publisher=

    Causation, prediction, and search , author=. 2000 , publisher=

  24. [24]

    Machine learning , volume=

    Learning Bayesian networks: The combination of knowledge and statistical data , author=. Machine learning , volume=. 1995 , publisher=

  25. [25]

    Journal of Statistical Software , volume=

    Learning Bayesian networks with the bnlearn R package , author=. Journal of Statistical Software , volume=

  26. [26]

    Machine learning , volume=

    The max-min hill-climbing Bayesian network structure learning algorithm , author=. Machine learning , volume=. 2006 , publisher=

  27. [27]

    Biostatistics , volume=

    Sparse inverse covariance estimation with the graphical lasso , author=. Biostatistics , volume=. 2008 , publisher=

  28. [28]

    Frontiers in microbiology , volume=

    Microbiome datasets are compositional: and this is not optional , author=. Frontiers in microbiology , volume=. 2017 , publisher=

  29. [29]

    PLoS computational biology , volume=

    Inferring correlation networks from genomic survey data , author=. PLoS computational biology , volume=. 2012 , publisher=

  30. [30]

    PLoS computational biology , volume=

    Sparse and compositionally robust inference of microbial ecological networks , author=. PLoS computational biology , volume=. 2015 , publisher=

  31. [31]

    Mathematical Geology , volume=

    Isometric logratio transformations for compositional data analysis , author=. Mathematical Geology , volume=. 2003 , publisher=

  32. [32]

    2015 , publisher=

    Modeling and Analysis of Compositional Data , author=. 2015 , publisher=

  33. [33]

    Proceedings of the Sixth Annual Conference on Uncertainty in Artificial Intelligence , pages=

    Equivalence and synthesis of causal models , author=. Proceedings of the Sixth Annual Conference on Uncertainty in Artificial Intelligence , pages=. 1990 , publisher=

  34. [34]

    The Annals of Statistics , volume=

    Identifiability of Gaussian structural equation models with equal error variances , author=. The Annals of Statistics , volume=. 2014 , publisher=

  35. [35]

    Journal of Machine Learning Research , volume=

    A linear non-Gaussian acyclic model for causal discovery , author=. Journal of Machine Learning Research , volume=

  36. [36]

    Journal of the Royal Statistical Society: Series B (Statistical Methodology) , volume=

    Stability selection , author=. Journal of the Royal Statistical Society: Series B (Statistical Methodology) , volume=. 2010 , publisher=

  37. [37]

    PLoS Computational Biology , volume=

    Waste not, want not: why rarefying microbiome data is inadmissible , author=. PLoS Computational Biology , volume=. 2014 , publisher=

  38. [38]

    Microbiome , volume=

    Simple statistical identification and removal of contaminant sequences in marker-gene and metagenomics data , author=. Microbiome , volume=. 2018 , publisher=