{"id":"dbea1be1-57b2-45f5-ab82-f81cc9cb25fd","arxiv_id":"2411.15400","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Group-wise normalization (G-RLE and FTSS) improves statistical power and false discovery control in differential abundance analysis of microbiome samples, according to extensive simulations.","lead":"This paper proposes two new ways to normalize microbiome count data when comparing two groups of subjects: G-RLE and FTSS. In simulations, the methods find more truly different microbes while keeping false discoveries under control better than existing normalization techniques.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"FTSS performance may hinge on the fixed tuning parameter p* and KDE bandwidth; absent a sensitivity analysis, the central claim that FTSS achieves top power with controlled FDR is not robustly established.","rationale":"We read the paper as making an empirical claim: group-wise normalization, especially FTSS, yields higher power and controlled FDR in realistic simulations. The mathematical derivation in Equation 1 is a motivation, but the actual support is the simulation study. The most vulnerable point in that support is the unexamined tuning parameter p* (and, secondarily, the KDE bandwidth). The reader's weakest_assumption focused on the minority-of-DA-taxa condition; that is an explicit scope condition, and the simulations do test up to 30% DA, a minority. The p* issue is different: it affects whether the reported results are robust even within the tested settings. Without a sensitivity analysis, a skeptical reader cannot tell if FTSS's advantage is a property of the method or an artifact of a particular p*. The paper's own PHACS/DESeq2 result is a warning that method behavior can break down when the reference-set estimation is unstable, and sparsity is exactly where the mode and percentile interval are hardest to pin down. A sweep over p* would settle this cheaply, since the code is available. We therefore keep the verdict at CONDITIONAL, agreeing with the reader that independent validation is needed, but our emphasis is on tuning-parameter robustness rather than the minority assumption alone.","tokens_in":11808,"tokens_out":13314,"duration_ms":123983,"concrete_test":"Run the model-based simulation for the high-variance, 30% signal, n=200 setting using FTSS with MetagenomeSeq (and, secondarily, DESeq2), sweeping p* over {0.1, 0.2, 0.3, 0.4, 0.5, 0.6} and using at least two KDE bandwidth choices (e.g., Silverman's rule and half/double that). Report mean TPR and FDR at nominal 0.05 over 1,000 replicates for each configuration. If the FDR exceeds 0.10 or the TPR ranking relative to GMPR reverses for any p* in the grid, the central claim that FTSS maintains FDR and maximizes power is contingent on an unstated tuning choice.","verdict_should_be":"UNCHANGED","load_bearing_attack":"FTSS selects reference taxa by retaining those whose observed log fold changes lie within a percentile interval of width p* around the kernel-density mode (Section 2.1.4). The value of p* controls a bias-variance trade-off: a small p* gives truncated sums based on few taxa, inflating variance; a large p* admits differentially abundant taxa, reintroducing compositional bias. The paper fixes p* (illustrated as 40% in Figure 1) and never reports a sensitivity analysis. Since the headline result is that FTSS has the highest TPR with FDR at or below 0.05, an unexamined tuning parameter is load-bearing: if the apparent advantage over GMPR depends on p* (or on the KDE bandwidth, which also goes unspecified), the central empirical claim is not robust. The theoretical derivation gives no guidance for choosing p*, and the absence of any robustness check means we cannot rule out that the reported ranking is specific to the chosen value. The PHACS/DESeq2 failure already shows the method's behavior degrades in sparse data, and p* interacts with sparsity: with many zero counts, the mode and the selected reference set become unstable. This concern directly targets the validity of the simulation-based claim, not just its scope.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a group-wise normalization framework for differential abundance analysis of microbiome count data. The authors derive that, under a log-linear model of absolute abundance and a binary covariate, each taxon's observed log fold change converges to the true log fold change plus a common bias term Δ equal to the log ratio of total absolute abundances between groups. They propose two normalization methods: G-RLE, which applies RLE to pooled group-level relative abundances, and FTSS, which estimates Δ as the mode of observed log fold changes and constructs a truncated library size from taxa near that mode. In model-based simulations (18 settings) and synthetic data based on two real microbiome datasets, the methods are compared to TSS, TMM, RLE, GMPR, CSS, and Wrench, paired with edgeR, DESeq2, and metagenomeSeq. The authors report that FTSS and G-RLE achieve higher true positive rates while maintaining FDR near 0.05 in many challenging settings, with FTSS + metagenomeSeq performing best. Code is publicly available.","tokens_in":11987,"tokens_out":7823,"duration_ms":67534,"significance":"If the results are substantiated, the group-wise normalization perspective is a valuable conceptual contribution: it reduces compositional bias to a single group-level parameter and suggests that group-level pooling may be more robust to zero inflation than sample-level normalization. The simulation study is extensive, with 1000 replicates per setting and realistic synthetic data, and the derivation of Equation (1) is clearly presented. The public availability of code is a strength. However, the empirical claims depend on tuning parameters that are not examined, and the framework's scope is explicitly limited to binary covariates.","major_comments":[{"comment":"The FTSS method depends on two tuning parameters: the truncation proportion p* (illustrated with p*=40% in Figure 1) and the bandwidth of the Gaussian kernel density estimator used to estimate the mode of observed log fold changes. No sensitivity analysis is reported for either parameter. Since p* controls the bias-variance trade-off in the reference taxon set and the bandwidth affects the mode estimate, the headline result that FTSS attains the highest true positive rate in every setting (Section 3.1, Figure 3) may be specific to the chosen tuning values. The theoretical derivation gives no guidance for choosing p*, and the absence of a robustness check means the central empirical claim is not fully supported.","section":"Section 2.1.4, Figures 1, 3, 4"},{"comment":"The abstract claims the proposed methods maintain the false discovery rate in challenging scenarios, but in the PHACS synthetic data with DESeq2, no normalization method including G-RLE and FTSS controls the FDR (Section 3.2, Figure 4). The paper acknowledges this in the Discussion but does not qualify the abstract's claim. The conditions under which the proposed methods fail should be stated, or the claim should be restricted to the settings where the methods succeed.","section":"Section 3.2, Discussion"},{"comment":"The framework is formally derived only for a binary covariate and a model where absolute abundances are deterministic within each group up to log-linear terms. The Discussion notes that continuous covariates are outside the scope, but the manuscript does not state the additional assumptions required for Δ to be a single common bias term when subject-level random effects are present (e.g., identical random-effect distributions across groups). The simulations do include random effects, but the derivation is not formally extended; this should be clarified so that readers know the exact conditions under which the proposed methods are guaranteed to remove compositional bias.","section":"Section 2.1.1, Discussion"}],"minor_comments":[{"comment":"In the formula for S_FTSS, the expression uses ρ(α1j) without a hat; it should be ρ(\\hat α1j) to match the definition of ρ as a function of the observed log fold changes.","section":"Section 2.1.4"},{"comment":"The sentence 'the bias term ∆ in would disappear' is missing a reference; it should point to Equation (1).","section":"Section 2.1.4"},{"comment":"The list of varied parameters reads 'β1q, . . . , β1q'; this should be 'β11, . . . , β1q'.","section":"Section 2.2.1"},{"comment":"The phrase 'the the development of methods' contains a duplicated article and should be corrected.","section":"Discussion"},{"comment":"The competing interests statement says 'no competing interest'; it should say 'no competing interests.'","section":"Section 5"},{"comment":"The supplementary materials referenced in Sections 2.1.1 and 2.2 are not included in the arXiv posting; please ensure they are uploaded for review, as the reproducibility of the simulation results depends on them.","section":"Supplementary materials"}],"recommendation":"major_revision","confidential_remarks":"The paper makes a solid conceptual contribution and the simulation study is extensive. The main revision needed is a sensitivity analysis for the FTSS tuning parameters (p* and KDE bandwidth), a qualified abstract that acknowledges the PHACS/DESeq2 failure, and inclusion of the supplementary materials. The binary-covariate limitation is acceptable for a methods paper but should be stated clearly in the abstract or introduction."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid, honest methods paper. The new contribution is real: reframing normalization as a group-level task, with G-RLE (RLE applied to pooled group proportions) and FTSS (truncated sum around the mode of observed log fold changes). The bias decomposition in Equation 1 is standard—LinDA and ANCOM-BC do similar things—but the translation into two concrete normalization factors is new and clean. The simulations are extensive (1,000 replicates, 18 model-based settings plus two synthetic real-data settings), and the code is public. That deserves credit.\n\nWhere I'd push back: the headline claim that FTSS has top power with FDR control depends on the fixed tuning parameter p* and the KDE bandwidth, and there is no sensitivity analysis for either. p* is load-bearing: too small inflates variance, too large reintroduces bias. The illustration uses p*=40%, but we don't know if the ranking survives other choices. That's a real gap, not a nitpick, but it doesn't sink the paper—it means the empirical claim is conditional. Also, the PHACS/DESeq2 setting where no method controls FDR is reported honestly, but it does limit the scope. The model-based simulations are generated from a model very close to the method's assumptions, so they are not independent validation. The supplementary derivation and extra settings couldn't be checked from the preprint alone; that's normal, but worth flagging.\n\nThe derivation itself is sound. The FTSS self-referential step (using observed fold changes to pick references that then re-estimate fold changes) is circular only in the same mild way TMM is; in practice it works because the mode is a robust location estimate under the minority-signal assumption.\n\nWho this is for: statisticians working on compositionality in microbiome DAA and users of edgeR/DESeq2/metagenomeSeq pipelines. It deserves a serious referee. I'd send it out, ask for a sensitivity analysis of p* and bandwidth, a real-data demonstration or at least a clearer statement of what would falsify the method, and a check of the supplementary code.\n\nRecommendation: peer-review accept with major/minor revision depending on the sensitivity analysis.","headline":"A genuinely new group-wise normalization idea with solid simulations, but the FTSS tuning parameter p* needs a sensitivity analysis before the headline claim is robust.","tokens_in":12583,"tokens_out":1391,"would_cite":true,"duration_ms":12820,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Group-wise normalization fixes microbiome abundance bias","keywords":["microbiome","differential abundance analysis","compositional bias","group-wise normalization","false discovery rate","reference taxa","zero-inflation","relative log expression"],"falsifier":"Generate a two-group microbiome dataset under the paper's multinomial model but with 60% of taxa having nonzero true log fold changes, then apply FTSS followed by MetagenomeSeq. If the method works, the estimated $\\hat{\\Delta}$ should still equal the true $\\Delta$ and the false discovery rate should stay near 0.05; a drift in $\\hat{\\Delta}$ and FDR inflation would falsify the mode-based correction.","tokens_in":47,"feed_emoji":"🫀","tokens_out":4218,"duration_ms":89783,"temperature":0.7,"pith_summary":"The paper argues that compositional bias in differential abundance analysis of microbiome data is not a per-sample nuisance but a single group-level quantity: the log ratio of total absolute abundance between the two groups being compared. It derives this bias term from a simple multinomial model and shows that all observed log fold changes are shifted by the same additive constant. On that basis it proposes two normalization methods, G-RLE and FTSS, which estimate this one constant from pooled group-level data instead of estimating separate scaling factors for every sample. In model-based and synthetic-data simulations, both methods achieve higher power for detecting differentially abundant taxa than existing normalizers while keeping false discovery rates near nominal levels in settings where competitors fail.","feed_headline":"Group-wise normalization fixes microbiome abundance bias","feed_subtitle":"New group-level normalizers G-RLE and FTSS turn that bias into a correction, lifting power while holding false discovery rates.","key_machinery":"The load-bearing identity is equation (1): under a log-linear model for absolute abundances followed by multinomial sampling, the pooled observed relative abundance converges to $\\exp(\\beta_{0j}+\\beta_{1j}g)/\\sum_k \\exp(\\beta_{0k}+\\beta_{1k}g)$, so the maximum-likelihood observed log fold change converges to $\\beta_{1j}+\\Delta$. This turns normalization into estimation of the single parameter $\\Delta$. G-RLE estimates it through the median of group-level fold changes, and FTSS estimates it as the mode of the observed log fold-change distribution via Gaussian kernel density, then rescales each sample's library size using only taxa within a percentile window around that mode.","core_discovery":"The central claim is that the observed log fold change $\\hat{\\alpha}_{1j}$ for taxon $j$ converges to the true log fold change $\\beta_{1j}$ plus a taxon-invariant bias $\\Delta = \\log\\left(\\sum_j e^{\\beta_{0j}} / \\sum_j e^{\\beta_{0j}+\\beta_{1j}}\\right)$, which is exactly the log ratio of average total absolute abundance in the two groups. Because $\\Delta$ does not depend on $j$, correcting it is a group-level problem: estimate one number, not $n$ sample fractions. G-RLE does this by applying relative-log-expression normalization to the two pooled group profiles, and FTSS does it by truncating the library size to taxa whose observed log fold changes cluster near the estimated mode, the assumed location of $\\Delta$ when only a minority of taxa are differentially abundant. The paper establishes this derivation and supports the methods with simulations showing improved true positive rate and false discovery rate control, particularly for MetagenomeSeq analysis of zero-inflated, high-variance data.","pith_inferences":["The same decomposition of bias into a taxon-invariant constant likely applies to any compositional count data with a binary covariate, such as RNA-seq or metabolomics, so FTSS could be tested outside the microbiome without modification.","The mode-based estimator for $\\Delta$ could be replaced by a trimmed mean or a robust location estimator to handle cases where the minority-of-signals assumption is only approximately met; the paper does not explore this variant.","For multi-group or continuous covariate designs, the single-constant bias structure breaks down; one could generalize $\\Delta$ to a per-group vector, but this would require a different reference-taxa rule than FTSS.","A natural stress test is to vary the signal proportion continuously from 5% to 50%: the paper tests 10–30%, and the method's performance should degrade smoothly as the mode becomes less identifiable."],"forward_implications":["Using FTSS or G-RLE as a preprocessing step for MetagenomeSeq gives higher true positive rates than TSS, TMM, RLE, GMPR, CSS, and Wrench in the paper's simulations, with FDR held near nominal even at 20–30% differential abundance.","The derivation implies that any DAA method using a library-size offset can incorporate the group-level correction, so edgeR and DESeq2 also gain power when paired with FTSS, except in the PHACS-like sparse-data setting where DESeq2 loses FDR control.","Because the group-level pooled counts are strictly positive, G-RLE and FTSS are expected to be robust to zero-inflation, a common feature of microbiome count matrices.","The bias term $\\Delta$ is a single number, so correcting it is statistically easier than estimating $n$ sample-specific normalization factors, which explains the performance gain in high-variance scenarios."],"supporting_citations":[{"why":"Treats compositional bias as an estimable parameter under a hierarchical multinomial model, the precedent for the group-wise bias-correction view.","marker":"[16]"},{"why":"Defines the relative log expression normalization that G-RLE modifies from sample-level to group-level.","marker":"[12]"},{"why":"Provides GMPR, a zero-inflated robust normalization baseline that the paper compares against and often outperforms.","marker":"[15]"},{"why":"Supplies MetagenomeSeq and cumulative sum scaling, the DAA method with which FTSS achieves its best results.","marker":"[9]"},{"why":"Supplies the zero-inflated, correlated count generation model used in the paper's model-based simulations.","marker":"[20]"},{"why":"Supplies the Poisson-lognormal perspective used to construct correlated positive and zero-valued taxon counts.","marker":"[21]"},{"why":"Documents the reference-taxa strategy and the broader evaluation of DAA methods that motivates the group-wise approach.","marker":"[6]"},{"why":"Represents the compositional data analysis class that corrects bias via parameter estimation, the approach the paper's normalization framework parallels.","marker":"[5]"}],"fun_headline_variants":["Bias in microbiome counts is group-level: new normalizers fix it","Group-wise scaling lifts power in microbiome differential abundance","FTSS and G-RLE: group normalization for unbiased microbiome tests","Microbiome abundance bias shrinks with group-level normalization","One bias, not many: group-wise normalization improves microbiome DA"],"cache_read_input_tokens":14720,"weakest_assumption_plain":"The method depends on the assumption that most taxa are equally abundant between the groups, so that the mode of observed log fold changes sits at the bias term; if a majority of taxa change, the reference set is contaminated and the correction misses.","fun_headline_variants_meta":{"raw":{"variants":["Bias in microbiome counts is group-level: new normalizers fix it","Group-wise scaling lifts power in microbiome differential abundance","FTSS and G-RLE: group normalization for unbiased microbiome tests","Microbiome abundance bias shrinks with group-level normalization","One bias, not many: group-wise normalization improves microbiome DA"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000714,"raw_usage":{"total_tokens":3226,"prompt_tokens":973,"completion_tokens":2253,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":589,"completion_tokens_details":{"reasoning_tokens":2168}},"tokens_in":589,"tokens_out":2253,"duration_ms":12938,"temperature":1.0,"reasoning_tokens":2168,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:20:23.671243+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate a two-group microbiome dataset under the paper's multinomial model but with 60% of taxa having nonzero true log fold changes, then apply FTSS followed by MetagenomeSeq. If the method works, the estimated $\\hat{\\Delta}$ should still equal the true $\\Delta$ and the false discovery rate should stay near 0.05; a drift in $\\hat{\\Delta}$ and FDR inflation would falsify the mode-based correction.","supporting_citations":[{"cited_title":"Analysis and correction of compositional bias in sparse sequencing count data","cited_arxiv_id":null,"evidence_quote":"Treats compositional bias as an estimable parameter under a hierarchical multinomial model, the precedent for the group-wise bias-correction view."},{"cited_title":"Differential expression analysis for sequence count data","cited_arxiv_id":null,"evidence_quote":"Defines the relative log expression normalization that G-RLE modifies from sample-level to group-level."},{"cited_title":"GMPR: A robust nor- malization method for zero-inflated count data with application to microbiome sequencing data","cited_arxiv_id":null,"evidence_quote":"Provides GMPR, a zero-inflated robust normalization baseline that the paper compares against and often outperforms."},{"cited_title":"Differential abundance analysis for microbial marker-gene surveys","cited_arxiv_id":null,"evidence_quote":"Supplies MetagenomeSeq and cumulative sum scaling, the DAA method with which FTSS achieves its best results."},{"cited_title":"Bayesian variable selec- tion for multivariate zero-inflated models: Application to microbiome count data","cited_arxiv_id":null,"evidence_quote":"Supplies the zero-inflated, correlated count generation model used in the paper's model-based simulations."},{"cited_title":"The Poisson-Lognormal Model as a Versatile Framework for the Joint Analysis of Species Abundances","cited_arxiv_id":null,"evidence_quote":"Supplies the Poisson-lognormal perspective used to construct correlated positive and zero-valued taxon counts."},{"cited_title":"A comprehensive evaluation of microbial differential abun- dance analysis methods: current status and potential solutions","cited_arxiv_id":null,"evidence_quote":"Documents the reference-taxa strategy and the broader evaluation of DAA methods that motivates the group-wise approach."},{"cited_title":"Analysis of compositions of microbiomes with bias cor- rection","cited_arxiv_id":null,"evidence_quote":"Represents the compositional data analysis class that corrects bias via parameter estimation, the approach the paper's normalization framework parallels."}],"review_version":1}