{"id":"acc25eac-0f11-41cd-8a84-f396e35ec94f","arxiv_id":"2509.06891","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"An MPS-based tensor network computes mutual information from binarized single-cell expression to infer gene regulatory networks, demonstrated on six NF-κB pathway genes.","lead":"Tensor networks, a tool from quantum physics, are used to infer gene regulatory networks from single-cell RNA sequencing data. The method recovers known interactions among six B-cell genes and reports one statistically significant triple interaction.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported permutation p-values are internally inconsistent: Table II has p=0.0001 although with 1000 permutations the minimum is 1/1001, and symmetric QMI yields asymmetric p-values for the same pair. The claimed edge and triad significance is therefore not reliable.","rationale":"I agree with the reader that GMM binarization is a substantive modeling assumption and that a sensitivity analysis is missing. However, the single most load-bearing weakness is at the next step: even if binarization is accepted, the reported significance tests are internally inconsistent, and the central claim is entirely supported by those tests. An impossible p-value (0.0001 with 1000 permutations) and asymmetric p-values for a symmetric statistic mean the reader cannot tell which edges or triads are real. This is not an accusation of misconduct; typos or undisclosed procedural changes could explain it, but in either case the paper as written does not provide a reliable statistical basis for the claim. The concrete check is straightforward and would settle the issue: recompute from raw data and verify symmetry and granularity, then re-test the triad with multiple-testing control. If the corrected analysis preserves the same edges and the triad survives correction, the reader's conditional verdict could stand after a revision. If not, the central claim is unsupported and the appropriate verdict is unverified rather than conditional. I therefore recommend moving from CONDITIONAL to UNVERDICTED until the corrected statistical analysis and artifacts are provided.","tokens_in":16675,"tokens_out":16276,"duration_ms":149183,"concrete_test":"Recompute all pairwise QMI values and both p-values from the raw binarized gene-by-cell matrix (GSE126321/GSE158275) using the exact procedure in Section IV.2 with Eqs. (2)-(3) and 1000 permutations. Verify that every off-diagonal entry satisfies p(i,j)=p(j,i) and that each reported p is a multiple of 1/1001; identify which Table I/II entries change. Then recompute all C(6,3)=20 triadic QMI permutation p-values and apply Bonferroni (threshold 0.05/20=0.0025) or FDR control to the PRDM1-PAX5-IRF4 result. If the corrected pairwise table no longer yields the claimed known edges, or the triad does not survive correction, the central claim fails. If the code/data are not provided, request them; without them the significance claims remain unverifiable.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The central claim is that the MPS-QMI pipeline recovers known regulatory edges and a triadic mechanism from single-cell data. The evidence for every recovered edge and for the triad is the permutation-test p-value. Those p-values are internally inconsistent in a way that affects the reliability of the inference.\n\nFirst, QMI is symmetric: I(i,j)=I(j,i). A correct permutation null for a pair, shuffling one gene's labels across cells, has the same distribution regardless of which gene is shuffled because the test statistic is symmetric and the null depends only on the two marginal count vectors. Hence p(i,j) and p(j,i) must be equal. Table II reports PRDM1-PAX5=0.0010 and PAX5-PRDM1=0.0001; Table I reports PRDM1-REL=0.9900 and REL-PRDM1=0.9990. These are not rounding differences.\n\nSecond, Eqs. (2)-(3) with the stated Nperm=1000 give a minimum possible p-value of 1/(Nperm+1)=1/1001 approximately 0.0010. Table II's PAX5-PRDM1=0.0001 is impossible under the described procedure, indicating either a typo, a different but unreported number of permutations, or a different formula. Because the same pipeline is used for the headline triadic p=0.011, that value cannot be checked from the paper alone.\n\nThird, the triadic claim is selected from 20 possible triples with no multiple-testing correction; a p=0.011 chosen post hoc is the order of magnitude expected for one false positive under the null. Thus the statistical foundation of the central claim is not currently trustworthy.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a tensor-network (matrix product state, MPS) framework for inferring gene regulatory networks from single-cell RNA-seq data. The pipeline binarizes per-gene expression with two-component Gaussian mixture models, orders genes with a Hilbert curve after PCA, encodes the empirical distribution as a quantum state, builds an MPS, and computes quantum mutual information (QMI) between gene pairs and triples, with statistical significance assessed by permutation tests. The method is applied to six NF-κB pathway genes (IRF4, REL, PAX5, RELA, PRDM1, AICDA) in roughly 28,000 lymphoblastoid cells from GEO accessions GSE126321 and GSE158275. The authors report that the inferred pairwise edges and a PRDM1-PAX5-IRF4 triad are consistent with known biology, and argue that the MPS approach avoids the hardware and optimization difficulties of variational quantum algorithms.","tokens_in":16983,"tokens_out":4123,"duration_ms":38040,"significance":"If the statistical claims were sound, the contribution would be useful: it offers a classical, quantum-inspired method for higher-order dependency detection in gene expression, with polynomial MPS scaling and a permutation-testing framework, and it explicitly targets triadic regulatory interactions that pairwise methods miss. The paper is also honest in comparing against an existing quantum GRN method and in discussing scalability limits. However, the evidence for the central recovery claim currently rests on permutation p-values that are internally inconsistent, and the validation lacks benchmark comparisons and sensitivity analyses. The significance is therefore potential rather than demonstrated.","major_comments":[{"comment":"The reported permutation p-values are internally inconsistent in a way that affects the headline results. First, QMI is symmetric, so for a symmetric permutation null the p-value for (i,j) must equal the p-value for (j,i); yet Table I reports PRDM1-REL = 0.9900 and REL-PRDM1 = 0.9990, and Table II reports PRDM1-PAX5 = 0.0010 while PAX5-PRDM1 = 0.0001. These are not rounding artifacts. Second, with Nperm = 1000, Eqs. (2) and (3) give a minimum possible p-value of 1/(1001) ≈ 0.0010, so the value 0.0001 in Table II for PAX5-PRDM1 is impossible under the stated procedure. Because the same pipeline produces the triadic p = 0.011, the reader cannot verify any of the claimed significances from the paper alone. The authors should rerun the analysis, provide reproducible code and seeds, and correct the tables.","section":"Section III.E, Tables I and II"},{"comment":"The statistical validation uses a threshold of p < 0.1 on fifteen pairwise tests with no multiple-testing correction. At this threshold one expects, under the null, roughly one to two spurious significant pairs among the fifteen comparisons, so the recovered-edge list is not convincing without correction or a stated family-wise error control procedure. The biological interpretation of left-tailed p-values as evidence of 'mechanisms enforcing independence' is also introduced only after the results are seen; the manuscript should specify a pre-registered or clearly justified testing protocol.","section":"Section IV.2 and Figure 6"},{"comment":"The triadic claim is not adequately supported. The paper reports that the PRDM1-PAX5-IRF4 triad has p = 0.0110, but this triple is apparently selected after inspecting the results among the 20 possible triples of six genes, and no multiple-testing correction is applied. Under the null, a p-value of about 0.011 is close to the expected minimum of 20 uniform p-values, so the finding could easily be a false positive. In addition, the permutation procedure described in Section III.E is defined only for gene pairs; the manuscript does not explain how the null distribution for the triadic QMI in Eq. (4) is generated, what statistic is permuted, or how the p-value is computed. This must be specified and corrected before the triadic conclusion can be assessed.","section":"Section IV.3 and Eq. (4)"},{"comment":"The binarization step is load-bearing: all QMI values are computed from the two-component GMM assignment of each gene to active/inactive states. The BIC comparisons support K = 2 for these genes, but they do not establish that the inferred network is robust to the binarization choice. The authors should provide a sensitivity analysis varying the GMM threshold (e.g., using posterior-probability cutoffs other than the maximum), the number of components, the PCA dimensionality, the MPS bond dimension, and the Hilbert-curve ordering, and show that the significant edges and the triad survive these choices.","section":"Section III.A and Section IV.1"},{"comment":"There is no quantitative benchmark against standard GRN inference methods. The validation consists of matching known interactions from STRING and the literature, which is useful but does not demonstrate that the method outperforms or complements existing tools such as ARACNE, GENIE3, PIDC, or the quantum method of Ref. [24]. A comparison on a common benchmark (e.g., synthetic datasets with known ground truth or DREAM challenge data) with AUROC/AUPRC metrics would materially strengthen the central claim.","section":"Section IV.2 and Discussion"}],"minor_comments":[{"comment":"There are several typographical errors, including 'rigth' in Section IV.2, 'succesful' in the Discussion, '28.000' for the cell count, and 'one-tailed-left and-rigth'. These should be corrected.","section":"Throughout"},{"comment":"The text says PCA is applied to 'the gene expression matrix X or its binarized form Z' and it is not clear which input was used for the reported results. Please clarify.","section":"Section III.B"},{"comment":"The tables report p-values to varying numbers of decimals and contain an apparent typo in Table II ('0.92131' for IRF4 in the REL row). Please standardize the precision and recheck all entries.","section":"Tables I and II"},{"comment":"The QMI heatmap in Figure 5 would benefit from a colorbar and axis labels, and the text should state whether the displayed values are raw QMI or normalized in some way.","section":"Figure 5"},{"comment":"References [36] and [37] are the same paper; one should be removed or the two citations should be merged.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the journal's scope, but the statistical inconsistencies in the core evidence are serious. I would encourage the editor to require the authors to provide code and exact permutation-test details, and to rerun the analysis so that the p-values are internally consistent and reproducible. The absence of any benchmark against standard GRN methods also makes the contribution hard to position."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know upfront: this is a readable, honest methods note, but the headline statistical evidence is currently unreliable. The stress-test note holds. Table II lists PRDM1-PAX5 as 0.0010 and PAX5-PRDM1 as 0.0001, and with Nperm=1000 their own formula cannot produce 0.0001. Table I's PRDM1-REL (0.9900) versus REL-PRDM1 (0.9990) also points to different or unstable permutation runs. Since every edge and the triadic p=0.011 rest on these values, the central recovery claim cannot yet be audited.\n\nWhat is genuinely useful: the combination of GMM binarization, Hilbert-curve gene ordering, MPS decomposition, and permutation-tested QMI is new for GRN inference. The six-gene B-cell example is well chosen, and the authors check their recovered edges against STRING and published interactions. They also position the method honestly against variational quantum GRN inference, noting the classical advantages. That is fair and well referenced.\n\nThe soft spots are in proportion. The biggest conceptual one is that QMI is not just classical mutual information here, but not for the reason the reader gives. The state vector uses sqrt(p) amplitudes, so the reduced density matrices are non-diagonal and I(A;B) differs from Shannon mutual information on the binarized data. That extra 'quantum' content is an artifact of the amplitude encoding, and the paper never justifies it or tests how sensitive the results are to that choice. Second, the GMM binarization is load-bearing and is only defended with BIC; no sensitivity analysis is provided. Third, there is no benchmark against standard GRN methods, no multiple-testing correction for the twenty triads, and the p<0.1 threshold is lenient. No code or data are provided, which makes all of this hard to verify.\n\nWho gets value: people working on quantum-inspired tensor-network methods in genomics. It is not ready for a practitioner to adopt. With the permutation issue fixed, a benchmark, and code release, it could become a solid methods paper. I would send it to referees rather than desk reject, but only with the expectation of serious revision.","headline":"A likeable quantum-inspired GRN pipeline, but the permutation p-values are internally inconsistent, so the recovery claim is not yet supported.","tokens_in":17588,"tokens_out":7888,"would_cite":false,"duration_ms":74008,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A matrix-product-state pipeline using quantum mutual information and permutation testing recovers a known six-gene B-cell regulatory network and reports a significant PRDM1–PAX5–IRF4 triad.","keywords":["gene regulatory network inference","tensor networks","matrix product state","quantum mutual information","single-cell RNA sequencing","Gaussian mixture model","Hilbert curve","higher-order gene interactions"],"falsifier":"Re-run the full pipeline on the same 28,000+ cells with the binarization changed to three-component Gaussian mixtures, percentile thresholds, or raw continuous expression, and compare which edges and which triads remain statistically significant; if the reported PAX5–AICDA, PRDM1–PAX5, and PRDM1–PAX5–IRF4 signals do not survive any of these encodings, the network is an artifact of the two-state discretization rather than of the tensor-network machinery.","tokens_in":16424,"feed_emoji":"🧬","tokens_out":9340,"duration_ms":77006,"temperature":0.7,"pith_summary":"Single-cell transcriptomes are snapshots of which genes are on in each cell, and this paper asks whether those snapshots contain enough information to reconstruct who regulates whom. The authors' proposal is to turn each cell's gene activity pattern into a quantum-like state vector, compress it with a tensor network (a matrix product state), and score every pair and every triple of genes with a nonparametric measure, quantum mutual information, using permutation tests to separate signal from noise. Applied to six genes that control B-cell fate in more than 28,000 lymphoblastoid cells, the method recovers several previously documented regulatory edges, such as PAX5–AICDA and PRDM1–PAX5, and it reports one statistically significant three-gene module, PRDM1–PAX5–IRF4, that pairwise analysis would not reveal. The broader claim is that this quantum-inspired classical pipeline can capture higher-order, nonlinear gene dependencies that standard pairwise correlation or regression methods miss.","feed_headline":"Tensor networks recover a B-cell circuit and a three-gene module","feed_subtitle":"A quantum-inspired classical pipeline validates known links and surfaces a three-gene regulatory module.","key_machinery":"The load-bearing object is a Matrix Product State (MPS), a compact factorization of an exponentially large tensor into a chain of low-rank tensors whose bond dimensions limit how much correlation can be stored between neighboring genes. The pipeline fits a two-component Gaussian Mixture Model to each gene's expression, binarizes each cell's state to active or inactive, orders genes with a Hilbert space-filling curve so that nearby genes in a principal-component projection stay adjacent in the chain, and forms the empirical quantum state from the frequencies of observed binary patterns. Quantum mutual information is then obtained by contracting the MPS to single- and two-site reduced density matrices, with a separate contraction for the three-site formula $I(i;j;k) = S(i)+S(j)+S(k)-S(ij)-S(ik)-S(jk)+S(ijk)$; positive values indicate synergy and negative values redundancy. Permutation tests (1000 permutations) with one-tailed left and right comparisons turn these scores into p-values, so the network consists only of dependencies unlikely to arise by chance.","core_discovery":"The paper's claim, stated in the Discussion, is that the matrix-product-state plus quantum mutual information pipeline succeeds in recovering a gene regulatory network consisting of six pathway genes — IRF4, REL, PAX5, RELA, PRDM1, and AICDA — from single-cell RNA sequencing data of more than 28,000 lymphoblastoid cells. The recovered pairwise edges match known biology: PAX5 induces AICDA, PRDM1 represses PAX5, PRDM1 and PAX5 link to REL, and RELA–REL, PAX5–IRF4, PRDM1–IRF4, and PRDM1–AICDA appear as significant interactions at p < 0.1. The method also computes triadic quantum mutual information and finds the PRDM1–PAX5–IRF4 triad significant (p = 0.0110), which the authors read as a coordinated regulatory module in which IRF4 can act upstream of PRDM1 and PAX5. In the same analysis, the PRDM1–AICDA–IRF4 triad is not significant, which they interpret as evidence that IRF4's effect on AICDA is indirect, mediated through PRDM1 rather than direct.","pith_inferences":["A natural extension the paper does not run is to replace the GMM binarization with continuous expression values or an alternative discretization; if the significant edges survive, the tensor-network mechanism is the carrier of the result, and if they do not, the binarization is doing the work.","The significant PRDM1–PAX5–IRF4 triad could be a signature of a simple chain, such as IRF4 activating PRDM1 and PRDM1 repressing PAX5, rather than irreducible three-way synergy; a partial information decomposition would settle which interpretation is right.","The separate left-tailed p-values suggest a tool for finding mutually exclusive or compensatory gene pairs genome-wide, since a significantly small QMI marks pairs that are more independent than chance would predict.","The Hilbert-curve ordering is presented as preserving biological locality, but an ablation study comparing random gene orderings would show how much of the inferred signal actually depends on that locality rather than on the mutual information computation itself."],"forward_implications":["If the claim holds, gene regulatory network inference can be done classically with tensor networks, avoiding the connectivity, noise, and barren-plateau problems that hamper variational quantum algorithm approaches.","The method produces higher-order interaction scores, so regulatory motifs such as synergy and redundancy can be examined directly rather than assembled from pairwise edges.","Because the pipeline is nonparametric and permutation-based, it can report significance for small gene sets and would give a principled way to add genes, subject to bond-dimension growth.","The recovered interactions in the NF-κB and plasma-cell differentiation circuit strengthen the case that known edges such as PAX5–AICDA and PRDM1–PAX5 are present in these data and that IRF4–AICDA is indirect.","The reported PRDM1–PAX5–IRF4 triad gives a concrete candidate module for follow-up experiments or for decomposition into unique, redundant, and synergistic information."],"supporting_citations":[{"why":"The variational-quantum gene-regulatory-network baseline the paper compares against and the source of several previously reported gene-pair links.","marker":"[24]"},{"why":"The review that supplies the matrix-product-state formalism, truncation, and density-matrix renormalization background used for MPS contractions.","marker":"[26]"},{"why":"The finite-mixture-model reference supporting the Gaussian Mixture Model binarization of gene expression into active and inactive states.","marker":"[35]"},{"why":"The Hilbert space-filling curve reference used to reorder genes so that nearby genes stay adjacent in the one-dimensional MPS chain.","marker":"[36]"},{"why":"The Gene Expression Omnibus repository paper for the two scRNA-seq datasets, GSE126321 and GSE158275.","marker":"[41]"},{"why":"The Bayesian Information Criterion reference used to justify the two-component GMM over one- or three-component fits.","marker":"[42]"},{"why":"The study of the NF-κB c-Rel regulatory circuit that supplies the literature network whose edges the inferred GRN is checked against.","marker":"[43]"},{"why":"The study identifying the Aicda promoter, supporting the PAX5–AICDA edge recovered by the method.","marker":"[45]"},{"why":"The study of PRDM1/Blimp-1 in B-cell lymphomas, supporting the PRDM1–PAX5 repression edge recovered by the method.","marker":"[47]"},{"why":"The study of PAX5 as a plasma-cell fate inhibitor, supporting the PAX5–IRF4 cross-regulation edge recovered by the method.","marker":"[51]"}],"fun_headline_variants":["Tensor networks uncover gene circuitry in single cells","Quantum-inspired net recovers B-cell gene module","From 28k cells: tensor net maps gene interactions","Tensor model reveals triadic gene regulatory motif","Single-cell tensor network finds six-gene circuit"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that each gene's continuous expression is well captured by two biological states, active and inactive, and that the Gaussian-mixture cut used to binarize the data does not throw away the regulatory information being measured.","fun_headline_variants_meta":{"raw":{"variants":["Tensor networks uncover gene circuitry in single cells","Quantum-inspired net recovers B-cell gene module","From 28k cells: tensor net maps gene interactions","Tensor model reveals triadic gene regulatory motif","Single-cell tensor network finds six-gene circuit"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000152,"raw_usage":{"total_tokens":1210,"prompt_tokens":958,"completion_tokens":252,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":574,"completion_tokens_details":{"reasoning_tokens":181}},"tokens_in":574,"tokens_out":252,"duration_ms":3404,"temperature":1.0,"reasoning_tokens":181,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:13:29.326880+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the full pipeline on the same 28,000+ cells with the binarization changed to three-component Gaussian mixtures, percentile thresholds, or raw continuous expression, and compare which edges and which triads remain statistically significant; if the reported PAX5–AICDA, PRDM1–PAX5, and PRDM1–PAX5–IRF4 signals do not survive any of these encodings, the network is an artifact of the two-state discretization rather than of the tensor-network machinery.","supporting_citations":[{"cited_title":"A novel independence test for somatic alterations in cancer shows that biology drives mutual exclusivity but chance ex- plains most co-occurrence,","cited_arxiv_id":null,"evidence_quote":"The Bayesian Information Criterion reference used to justify the two-component GMM over one- or three-component fits."},{"cited_title":"Mutual exclusivity analysis identifies oncogenic network modules,","cited_arxiv_id":null,"evidence_quote":"The study of the NF-κB c-Rel regulatory circuit that supplies the literature network whose edges the inferred GRN is checked against."},{"cited_title":"Prdm1/blimp1: a tumor suppressor gene in b and t cell lymphomas,","cited_arxiv_id":null,"evidence_quote":"The study of PRDM1/Blimp-1 in B-cell lymphomas, supporting the PRDM1–PAX5 repression edge recovered by the method."},{"cited_title":"Loss of pax5 promotes plasma cell differentiation,","cited_arxiv_id":null,"evidence_quote":"The study of PAX5 as a plasma-cell fate inhibitor, supporting the PAX5–IRF4 cross-regulation edge recovered by the method."}],"review_version":1}