{"id":"9ee68123-3c1b-4daf-8d17-c9d3ade732f5","arxiv_id":"2506.15855","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"A Bayesian NMF model with a correlated multivariate normal prior for mutational signatures is proposed, but the claimed accuracy gain is only demonstrated in favorable simulations.","lead":"This paper introduces a Bayesian method for finding cancer mutation signatures that assumes the 96 mutation types are correlated, not independent. The method converges faster in simulations, but the evidence for better accuracy relies on giving the method the correct correlation structure as a prior.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The small-sample accuracy improvement rests on an oracle-like prior: the MVN method is given the generative covariance in §4.3, so the 6/8 vs 4/8 comparison does not support the abstract's claim.","rationale":"I read the paper as proposing two Bayesian NMF extensions and claiming empirical gains in convergence and accuracy. The derivations in §7.3–§7.4 are coherent, the method is a reasonable extension of bayesNMF, and the software contribution is real. The load-bearing weakness is the §4.3 simulation design: the MVN method is effectively handed the true generative covariance, while nonMVN is not. In Bayesian simulation, using the true prior when generating data is standard for checking posterior calibration, but it cannot support the comparative claim made in the abstract. The paper's own hierarchical results on Panc-Endocrine, with near-zero estimated correlations, underscore that the dependence structure may not transfer across cancer types. If the concrete test shows the gain persists under a COSMIC-derived prior, the paper would be strengthened; if not, the advertised advantage is unsupported. I agree with the reader's identification of the weakest assumption, and I would keep the REJECT verdict.","tokens_in":24837,"tokens_out":4247,"duration_ms":45673,"concrete_test":"Re-run the §4.3 simulation with the MVN prior covariance set to the James–Stein shrinkage estimate of the COSMIC correlation matrix described in §3.1.4 (as used in §4.2), while keeping the generative Σ at σ2P=7, ρsame=0.5, ρdiff=-0.1. Repeat across at least 20 simulated datasets and report the fraction of signatures with cosine similarity ≥0.9 for MVN versus nonMVN. If the MVN advantage disappears or diminishes to noise, the abstract's small-sample accuracy claim is an artifact of supplying the generative covariance. A secondary check: run the same comparison with a mis-specified covariance, such as the COSMIC estimate, to test whether the method's benefit survives prior mismatch.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central advertised benefit—'improvements in accuracy, especially on small sample sizes'—is supported almost entirely by §4.3. There, P is generated from the 3-parameter covariance with σ2P=7, ρsame=0.5, ρdiff=-0.1, and the MVN method is, by the text's own narrative, compared after being given this same structure, while nonMVN is forced to assume an identity covariance. The sentence 'Given that this dataset was constructed under the assumption of the 3-parameter hyperprior structure ... Perhaps the Bayesian Hierarchical model will be able to uncover the parameters that generated this data' makes the oracle nature of the comparison apparent. This is an in-sample, correctly-specified-model versus mis-specified-model comparison. It does not test the practical claim that a COSMIC-derived covariance (or any external covariance) improves recovery on real data. The real-data hierarchical fit finds ρsame≈0.05 and ρdiff≈-0.008, suggesting the assumed dependence is weak or absent in at least one cancer type. Additionally, even if the comparison were fair, it rests on only two simulated datasets. The mathematical derivations and software contribution are not in question; the empirical evidence for the headline claim is.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes two Bayesian NMF extensions for mutational signature analysis: (i) a multivariate truncated normal prior on the columns of the signatures matrix, with covariance estimated from COSMIC, and (ii) a hierarchical model in which a three-parameter covariance structure (sigma2_P, rho_same, rho_diff) is learned via Metropolis-Hastings. The authors derive the relevant full conditionals, implement the methods in an open-source R package, and report simulation studies plus an application to Panc-Endocrine PCAWG data. The abstract claims that the MVN model converges in fewer MCMC iterations and improves accuracy, especially on small sample sizes.","tokens_in":25103,"tokens_out":6771,"duration_ms":81041,"significance":"The full conditional derivation in Appendix 7.4 is algebraically sound, and the biological motivation for modeling dependence across mutation types is reasonable. The contribution of open-source code is also useful. However, the main empirical evidence for the headline accuracy claim is an in-sample comparison in which the MVN method is supplied with the generative covariance, and the supporting statistical tests contain methodological flaws. As presented, the paper does not demonstrate that an external or learned covariance improves signature recovery in realistic settings. The significance is therefore currently limited to a proof-of-concept, pending a properly designed evaluation.","major_comments":[{"comment":"The small-sample accuracy comparison is an oracle comparison. The two datasets are generated from the three-parameter covariance with sigma2_P=7, rho_same=0.5, rho_diff=-0.1, and the MVN method is given the same structure as its prior, while the nonMVN method is restricted to an identity covariance. The reported 6/8 versus 4/8 recovered signatures therefore measures the gain from using the true generative covariance against a deliberately misspecified model, not the practical value of a COSMIC-derived or learned covariance. The sentence in Section 4.3, 'Given that this dataset was constructed under the assumption of the 3-parameter hyperprior structure ...', makes this explicit. This experiment cannot support the abstract's claim of 'improvements in accuracy, especially on small sample sizes'. The same issue appears in attenuated form in Section 4.2, where the ground-truth signatures are drawn from COSMIC and the MVN prior covariance is also estimated from COSMIC.","section":"§4.3 (Figure 6)"},{"comment":"All simulation results are based on one dataset in Section 4.2 and two datasets in Section 4.3, with no multiple seeds, no error bars, and no quantification of MCMC variability. The convergence-iteration differences in Figure 4 and the cosine-similarity differences in Figures 5–7 are therefore point estimates with unknown variability. Since the headline efficiency and accuracy claims rest on these figures, the simulations need to be repeated over many simulated datasets and summarized with appropriate measures of spread.","section":"§4.2–§4.4"},{"comment":"The permutation test does not simulate the null hypothesis it claims to test. Under the null, each mutation type is drawn independently from a univariate truncated normal, but the sampled columns are then normalized to sum to one. Normalization induces negative dependence among the 96 entries, so the simulated null correlation distribution is not one of independence. The result that 319/4560 approximately 6.996% of adjusted p-values are significant therefore does not establish that dependence in the COSMIC signatures exceeds chance; it may simply reflect the normalization artifact.","section":"§4.1.2"},{"comment":"The hierarchical model is validated only on data generated from its own three-parameter covariance and on one real dataset. In the simulation, the recovered rho_same=0.21 and rho_diff=0.08 are not close to the generative values 0.5 and -0.1, so the claim that these are 'close to the original parameter values' is inaccurate, even after accounting for the scale non-identifiability of sigma2_P. On Panc-Endocrine the MAP estimates are near zero (rho_same=0.053, rho_diff=-0.008), which undercuts the motivating assumption for that dataset, and the claimed superiority of the hierarchical model is based only on a visual comparison of cosine-similarity heatmaps without a formal test or uncertainty estimate. This does not provide evidence that the hierarchical model learns dependence structure in real applications.","section":"§4.4–§4.5"}],"minor_comments":[{"comment":"The text states N=3 and G=60 for the hierarchical simulation, but the Figure 7 caption reports a sample size of 120, and the following paragraph refers to 'a smaller sample size (of 30 or 60 instead of 120)'. This inconsistency should be resolved.","section":"§4.4 and Figure 7"},{"comment":"There is a duplicated phrase, 'we will still perform the same the same Gibbs updates', which should be corrected.","section":"§3.2.2"},{"comment":"The manuscript repeatedly refers to itself as 'this thesis' and 'this project' (e.g., Sections 2.1 and 5.1). For a journal submission, this framing should be replaced with neutral language.","section":"General"},{"comment":"The implemented sampler uses truncation bounds at mean plus or minus 10 standard deviations rather than the stated [0, infinity) truncation. This changes the effective prior, and the paper should state this approximation in the main text and discuss its potential impact on the posterior.","section":"§7.2"},{"comment":"Some references have incomplete or irregular author formatting, for example 'Wilhelm and G, 2023' and 'V.' in the Favero et al. entry. These should be brought into consistent journal style.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The mathematical derivations are sound and the code contribution is real, but the empirical evaluation does not currently support the headline claims. The oracle-like simulation in Section 4.3 and the invalid permutation test in Section 4.1.2 are the main obstacles. I would be willing to see a revised version with a fair benchmark using external or held-out covariance information, multiple replications, and a corrected null procedure. If the authors cannot provide such evidence, the abstract's accuracy claims should be substantially weakened."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read it. What's actually new is a truncated multivariate normal prior on the signature matrix in Bayesian NMF, with covariance supplied externally (COSMIC shrinkage) or learned via a two-parameter hierarchical model. The full conditional derivations are correct, the code is in an open-source R package, and the authors are honest about per-iteration cost. That is a real, if modest, contribution.\n\nThe soft spots are mostly in the evaluation. The small-sample comparison (8 signatures, 10 samples) rests on two simulated datasets, no replicates and no error bars, and the cosine-similarity \"discovered\" counting is threshold-based. I also think the stress-test note overstates the oracle problem. Section 3.1.4 says the MVN prior used in the simulation studies is the James-Stein shrunk COSMIC correlation matrix, and Section 4.3 does not state that the true generative covariance was supplied to the MVN method. The section is under-specified, though: it never says exactly which covariance matrix each method received. The authors should fix that. If the MVN method was handed the generative covariance, the 6/8 vs 4/8 result is meaningless; if it used COSMIC, the result is still too thin to support the abstract's \"improvements in accuracy\" claim.\n\nThe permutation test in 4.1.2 is genuinely broken: the null matrices are column-normalized after sampling independent mutation types, and normalization induces negative dependence, so the null is no longer independence. The observed excess (about 7% vs 5%) is also small.\n\nThe hierarchical simulation does not recover the generating parameters well: MAP estimates are 0.21 and 0.08 for true values 0.5 and -0.1. Calling that \"close\" is generous, and the Panc-Endocrine fit gives near-zero correlations, which undercuts the biological motivation.\n\nNet: the math and implementation are solid, the empirical case for the headline claim is not. This is a major-revision paper, not a reject-for-incoherence. It deserves a serious referee and a request for more simulations, explicit prior-covariance reporting, and a valid permutation null.","headline":"A mostly sound Bayesian NMF extension whose advertised accuracy gain rests on thin, under-specified simulations.","tokens_in":25580,"tokens_out":5613,"would_cite":true,"duration_ms":65389,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"Covariance-aware NMF recovers more mutational signatures","keywords":["mutational signatures","Bayesian NMF","multivariate truncated normal","covariance structure","Gibbs sampling","hierarchical model","somatic mutations","cancer genomics"],"falsifier":"Generate a test set from independent truncated normals (identity covariance) and run both models: if the covariance-aware model still recovers more signatures than the independent model, its advantage does not come from modelling true dependence; if it recovers fewer, the in-sample advantage of a matched prior is confirmed.","tokens_in":24622,"feed_emoji":"🧬","tokens_out":6481,"duration_ms":68012,"temperature":0.7,"pith_summary":"Mutational signature analysis typically decomposes a tumor's 96-type mutation counts into a product of signatures and exposures, and standard Bayesian treatments assume the 96 mutation probabilities are independent. This paper argues that mutation types are measurably dependent—mutations sharing a center base are positively correlated—and that ignoring this dependence costs accuracy when samples are few. The authors propose two Bayesian NMF extensions: one that places a Multivariate Truncated Normal prior on the signatures matrix using an externally estimated covariance, and a hierarchical version that learns the covariance itself as two correlation parameters. In simulations, the covariance-aware model reaches the same posterior quality in fewer MCMC iterations and, on small samples, recovers six of eight true signatures where the independent model recovers four. If the claims hold, small cancer cohorts could yield more accurate signatures, and the learned covariance could expose biological connections among mutation types.","feed_headline":"Covariance-aware NMF recovers more mutational signatures","feed_subtitle":"Correlated mutation-type priors converge faster and find 6 of 8 true signatures, versus 4 for independent ones.","key_machinery":"The load-bearing object is the 96-by-96 covariance matrix $\\Sigma$ of the signatures matrix $P$. It is inserted as the scale matrix of a Multivariate Truncated Normal prior on each column of $P$, and the Normal likelihood is conjugate enough that the full conditional of each column is again a truncated multivariate normal with precision $\\Sigma^{-1} + (\\sum_g E_{ng}^2)(\\Sigma^M)^{-1}$. In the hierarchical model $\\Sigma$ is parameterized as $\\sigma_P^2$ times a two-correlation exchangeable structure: $\\rho_{\\text{same}}$ for mutation types sharing a center base, $\\rho_{\\text{diff}}$ otherwise, giving a parsimonious, interpretable covariance that can be learned by Metropolis-Hastings within the Gibbs sampler.","core_discovery":"The central claim is that replacing independent truncated-normal priors on each entry of the signatures matrix with a single Multivariate Truncated Normal prior—whose covariance encodes how mutation types co-vary—improves both the speed and the accuracy of Bayesian NMF for mutational signatures. The paper derives the full conditional for each signature column, which remains a truncated multivariate normal, and validates the model in simulations: with sixty samples and five true signatures it converges in hundreds to thousands fewer iterations than the independent model at equal accuracy, and with ten samples and eight true signatures it discovers more of them (six versus four by the paper's cosine-similarity threshold). The hierarchical version replaces the fixed covariance with a three-parameter structure—a base variance and two correlation parameters, one for mutation types sharing the same center base and one for those that do not—and shows on simulated data that the algorithm recovers the generating correlation parameters and, on a pancreas-endocrine cancer cohort, estimates near-zero correlations while still matching reference signatures as well as or better than the independent model.","pith_inferences":["If the in-sample matched covariance is the source of the gain, then the method's practical value depends on how transferable a reference-catalog covariance is across cancer types; this is directly testable by applying the MVN prior built from one cancer type's signatures to another.","The near-zero $\\rho_{\\text{same}}$ and $\\rho_{\\text{diff}}$ on the real dataset could mean the simple two-parameter covariance is too restrictive, or that the biological dependence is real but weaker than the reference catalog suggests; distinguishing these would require checking whether a richer covariance (e.g., per-center $\\rho$) improves real-data recovery.","The paper's convergence criterion is based on stabilization of the MAP, so 'fewer iterations to convergence' measures optimization progress rather than full posterior mixing; a user who needs credible intervals should also check trace convergence of the covariance parameters themselves."],"forward_implications":["Small sample cohorts (few samples per signature) can recover more true signatures when the prior accounts for mutation-type correlations.","The correlated model reaches the same MAP posterior in fewer MCMC iterations, though each iteration costs more, so wall-clock gains depend on implementation.","The hierarchical covariance parameters give a compact summary of how mutation types co-vary across cancers, with $\\rho_{\\text{same}}$ versus $\\rho_{\\text{diff}}$ interpretable as biological coupling.","Because only the prior changes, the formulation carries over to other NMF alphabets and to any non-negative matrix decomposition where columns are expected to be correlated."],"supporting_citations":[{"why":"Supplies the baseline Normal-Truncated Normal Bayesian NMF model, its convergence control, and the comparison method called nonMVN.","marker":"[Landy et al., 2025]"},{"why":"Provides the reference catalog whose correlation matrix becomes the external covariance prior, and the reference signatures used to score recovery.","marker":"[Tate et al., 2019]"},{"why":"Provides the truncated multivariate normal Gibbs sampler the paper uses to draw from the full conditional.","marker":"[Wilhelm and G, 2023]"},{"why":"Establishes Bayesian NMF with a Normal likelihood and truncated-normal priors, the framework the MVN model extends.","marker":"[Schmidt et al., 2009]"},{"why":"Gives the shrinkage estimator used to turn the reference signature sample covariance into a positive-definite prior.","marker":"[Ledoit and Wolf, 2003]"}],"fun_headline_variants":["Correlated priors speed up Bayesian NMF for signatures","Bayesian NMF with correlated mutation types finds more signatures","Modeling mutation-type correlations boosts signature discovery","Covariance-aware priors improve mutational signature recovery"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The simulations that show the accuracy gain generate data from a covariance structure and hand the same structure to the model as its prior, so the reported improvement is an in-sample test; in real applications the external covariance may not match the tumor type being analyzed.","fun_headline_variants_meta":{"raw":{"variants":["Correlated priors speed up Bayesian NMF for signatures","Bayesian NMF with correlated mutation types finds more signatures","Modeling mutation-type correlations boosts signature discovery","Covariance-aware priors improve mutational signature recovery"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000274,"raw_usage":{"total_tokens":1682,"prompt_tokens":1030,"completion_tokens":652,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":646,"completion_tokens_details":{"reasoning_tokens":600}},"tokens_in":646,"tokens_out":652,"duration_ms":6769,"temperature":1.0,"reasoning_tokens":600,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:48:47.127594+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate a test set from independent truncated normals (identity covariance) and run both models: if the covariance-aware model still recovers more signatures than the independent model, its advantage does not come from modelling true dependence; if it recovers fewer, the in-sample advantage of a matched prior is confirmed.","supporting_citations":[],"review_version":1}