REVIEW 4 major objections 4 minor 59 references
Multiview Graph Fusion with Covariates
T0 review · 4 major / 4 minor · reviewed 2026-07-13 · grok-4.5
Pith's one-line read A hierarchical Bayesian model jointly learns multiple graphs on shared nodes while linking them to predictors, with proven predictive consistency and better node recovery than separate or tensor methods.
desk verdict Solid multiview graph-response Bayesian model with shared node selection, Hellinger consistency, and a usable fMRI illustration; the shared-ξ prior is a real modeling choice that simulations never stress-test against view-specific alternatives. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
A shared spike-and-slab prior on the stacked node-specific latent vectors across all graph views (equation 4), which induces joint selection of nodes associated with a predictor and couples the low-rank graph coefficient matrices through a common covariance.
What would settle it
Generate multiview graphs in which a non-empty set of nodes is truly active in only one view and inactive in the others; if the joint model still forces those nodes to be selected (or unselected) across all views and loses estimation accuracy relative to independent learning, the shared-indicator claim fails.
Extended reading notes
Core claim
Under mild growth and sparsity conditions, the posterior predictive density of the proposed multiview generalized linear model converges in Hellinger distance to the true data-generating density, while the shared hierarchical prior on node-specific latent vectors yields more accurate coefficient estimates and node selection than independent graph-on-predictors learning or predictor-dependent tensor learning.
Load-bearing premise
A node is either associated with a key predictor in every graph view or in none; view-specific associations break the shared selection mechanism.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a hierarchical Bayesian GLM for joint predictor-dependent learning of multiview undirected graphs on a common node set, allowing continuous or binary edge weights. Graph coefficients for key predictors are given a low-rank factorization through node-specific latent vectors, coupled across views by a joint spike-and-slab prior with shared inclusion indicators ξ_{p,k} and an inverse-Wishart covariance (Eq. 4). Auxiliary predictors enter with scalar coefficients. The authors prove Hellinger consistency of the posterior predictive density under growth and sparsity assumptions (Theorem 3.1), give a Gibbs sampler via standard full conditionals, report lower coefficient MSE and good node AUC versus independent learning (IL) and tensor learning (TL) in continuous-edge simulations (n=150, K=40, M=2), and apply the method to task-based fMRI functional connectivity (inhibition/initiation) related to MMSE, selecting 88 ROIs.
Significance. If the modeling assumptions hold, the work addresses a real methodological gap: joint multiview graph-on-predictors regression that respects symmetry, accommodates mixed edge types, performs node-level selection with uncertainty quantification, and supplies asymptotic theory for heterogeneous multiview graph responses. The fMRI application is scientifically relevant. Explicit full conditionals, transparent assumptions for the Hellinger result, and comparisons to natural competitors (IL, TL) are genuine strengths. The shared-sparsity prior is the central modeling device that underwrites both the theory and the claimed gains over IL; its empirical and theoretical support is therefore load-bearing for the contribution.
major comments (4)
- [§2.3, Eq. (4); §5] Section 2.3, Eq. (4): The shared indicators ξ_{p,k} force a node to be associated with a key predictor in every view or in none. Simulations (§5) generate data under exactly this shared-ξ mechanism (common ξ_k^{(0)}, correlated latents across M=2 views), so JL is correctly specified while IL is not; the reported MSE/AUC advantages are therefore expected under the matching generative model and do not establish robustness. A sensitivity study with view-specific true associations is needed before claiming inferential superiority over independent learning.
- [Assumption (A); §6] Assumption (A) requires R_n K_n ≺ n/log(n). The fMRI analysis uses K=200 and n=144, which strains or violates this regime for any nontrivial fitted rank R. The paper should discuss relevance of the asymptotic conditions to the application and supply finite-sample checks (e.g., smaller-K or larger-n experiments) that sit inside the stated growth regime.
- [Theorem 3.2; §3] Theorem 3.2 establishes that the posterior probability of over-selecting influential nodes vanishes only under continuous Gaussian edges with known unit variance. The abstract and model emphasize mixed binary/continuous views; the scope of node-selection guarantees should be clarified or the result extended, and the gap between Theorem 3.1 (predictive Hellinger) and node-level inference should be stated explicitly.
- [§5] Section 5 restricts all scenarios to continuous edges, M=2, n=150, K=40, and data generated under the same shared-sparsity structure as the prior. Performance under binary edges, larger M, misspecified rank, or view-specific sparsity is not assessed, limiting support for the general claims of the abstract and introduction.
minor comments (4)
- [Figure 1] Figure 1 is hard to read in grayscale; a clearer encoding of true activity vs posterior probability would help.
- [§2.2–2.3] Notation for stacked latents (β̃_{p,k}) and the distinction between fitted rank R and effective rank induced by λ is introduced densely; a short summary table of parameters would improve readability.
- [§6.1] The median-probability threshold 0.5 for ROI selection (§6.1) is standard but should be noted as a free choice; sensitivity to the threshold would strengthen the application.
- [§4] Appendix references in the main text (full conditionals, ROI list) are clear, but the main text could briefly state that binary-edge full conditionals follow the same Gibbs structure with logistic likelihood contributions.
Circularity Check
No significant circularity: asymptotic predictive consistency and simulation gains are standard Bayesian results under stated model class and matching DGP, not tautologies of fitted inputs or self-citation chains.
-
self citation load bearing
[§1 Novelty paragraph; also §2.2 low-rank motivation]
"Our theoretical exposition introduces several novel aspects over existing work in Bayesian predictor-dependent learning of multiple graphs. Firstly, the theoretical framework in this article addresses joint modeling with multiview graph responses, unlike scenarios of a single graph response addressed in prior literature [23, 19, 20]."
Authors cite their own single-graph papers for contrast and for the low-rank/transitivity construction. This is ordinary background, not load-bearing: the multiview shared-ξ prior (eq. 4), the joint posterior, and Theorems 3.1–3.2 are derived and proved in the present paper under new assumptions; no uniqueness theorem from the self-citations is invoked to force the multiview result.
full rationale
The central claims (Theorem 3.1 Hellinger consistency of the posterior predictive under Assumptions (A)–(G); finite-sample MSE/AUC superiority of JL vs IL/TL) do not reduce by construction to their inputs. The model (eqs. 1–4) is a hierarchical GLM with low-rank graph coefficients and a shared spike-and-slab on node latents; the theorems prove posterior concentration for that class under external growth/sparsity conditions on Kn, Rn, sn and the true low-rank coefficients (Assumptions A–G), using standard techniques (Hellinger balls, prior mass, testing) whose proofs appear self-contained in Appendix A and cite external results (Ghosal–van der Vaart, Jiang). Simulations generate data from the same shared-ξ / correlated-latent mechanism that defines the prior, so JL is correctly specified while IL is not; the reported gains are therefore expected under correct specification, not a fitted parameter renamed as a prediction. Mild self-citations of the authors’ prior single-graph work ([19,21,23,20]) supply background motivation and contrast, but the multiview joint prior, the shared-ξ construction, and the multiview consistency theorems are new and do not rest on an unverified uniqueness claim imported from those papers. The shared-ξ assumption is a modeling choice whose misspecification risk is real (as the skeptic notes) but is a correctness/robustness issue, not circularity of the derivation. Score 1 reflects only the ordinary background self-citation; the derivation chain itself is independent.
Assumptions & free parameters
free parameters (5)
- Fitted latent rank R (R_n)
- Spike-and-slab hyperparameter b_η
- Dirichlet weight ω for λ^(r)_{p,m}
- Inverse-Wishart ν and IG(a_σ,b_σ) for J_p and σ_m²
- Node-inclusion threshold 0.5 (median probability rule)
assumptions (5)
- domain assumption True and fitted graph coefficients admit low-rank factorizations with R_n ≥ R^*_{n,m} (Assumptions B–C, eq. 3).
- ad hoc to paper Node activity indicators ξ_{p,k} are shared across all M views (eq. 4).
- domain assumption R_n K_n ≺ n/log(n) and related sparsity growth b_η s_n/K_n ≻ n (Assumptions A, F).
- standard math Edges follow GLM densities with link derivatives satisfying Assumption (E); continuous case uses i.i.d. Gaussian errors for Thm 3.2.
- domain assumption Covariates are bounded |x|≤a_0 (Assumption G); no self-loops; undirected symmetry of each view.
invented entities (2)
-
Joint multiview node-latent vector β̃_{p,k} with shared spike-and-slab and IW covariance J_p across views
-
Discrete λ^{(r)}_{p,m} ∈ {-1,0,1} with Dirichlet(rω,1,1) for effective rank control
Cite this review
Pith. "Pith review of Multiview Graph Fusion with Covariates." pith.science (2026). https://pith.science/paper/O2FFV7VG
@misc{pith2026260322215,
author = {Pith},
title = {Pith review of: Multiview Graph Fusion with Covariates},
year = {2026},
howpublished = {\url{https://pith.science/paper/O2FFV7VG}},
note = {Machine review of arXiv:2603.22215}
}
read the original abstract
Joint modeling of multiview graphs with a common set of nodes between views and auxiliary predictors is an essential, yet less explored, area in statistical methodology. Traditional approaches often treat graphs in different views as independent or fail to adequately incorporate predictors, potentially missing complex dependencies within and across graph views and leading to reduced inferential accuracy. Motivated by such methodological shortcomings, we introduce an integrative Bayesian approach for joint learning of a multiview graph with vector-valued predictors. Our modeling framework assumes a common set of nodes for each graph view while allowing for diverse interconnections or edge weights between nodes across graph views, accommodating both binary and continuous valued edge weights. By adopting a hierarchical Bayesian modeling approach, our framework seamlessly integrates information from diverse graphs through carefully designed prior distributions on model parameters. This approach enables the estimation of crucial model parameters defining the relationship between these graph views and predictors, as well as offers predictive inference of the graph views. Crucially, the approach provides uncertainty quantification in all such inferences. Theoretical analysis establishes that the posterior predictive density for our model asymptotically converges to the true data-generating density, under mild assumptions on the true data-generating density and the growth of the number of graph nodes relative to the sample size. Simulation studies validate the inferential advantages of our approach over predictor-dependent tensor learning and independent learning of different graph views with predictors. We further illustrate model utility by analyzing functional connectivity graphs in neuroscience under cognitive control tasks, relating task-related brain connectivity with phenotypic measures.
Reference graph
Works this paper leans on
-
[1]
R., Snyder, A
Andrews-Hanna, J. R., Snyder, A. Z., Vincent, J. L., Lustig, C., Head, D., Raichle, M., and Buckner, R. L. (2007). Disruption of large-scale brain systems in advanced aging. Neuron,56(5), 924–935
2007
-
[2]
W., Murray, J
Anticevic, A., Cole, M. W., Murray, J. D., Corlett, P. R., Wang, X.-J., and Krystal, J. H. (2012). The role of default network deactivation in cognition and disease.Trends in Cognitive Sciences,16(12), 584–592
2012
-
[3]
Barbieri, M. M. and Berger, J. O. (2004). Optimal predictive model selection.The Annals of Statistics,32(3), 870–897
2004
-
[4]
and Sporns, O
Bullmore, E. and Sporns, O. (2009). Complex brain networks: graph theoretical analysis of structural and functional systems.Nature Reviews. Neuroscience,10(3), 186–198
2009
-
[5]
Cabeza, R., Albert, M., Belleville, S., Craik, F. I. M., Duarte, A., Grady, C. L., Linden- berger, U., Nyberg, L., Park, D. C., Reuter-Lorenz, P. A., Rugg, M. D., Steffener, J., and Rajah, M. N. (2018). Maintenance, reserve and compensation: the cognitive neuroscience of healthy ageing.Nature Reviews Neuroscience,19(11), 701–710
2018
-
[6]
Y., Park, D
Chan, M. Y., Park, D. C., Savalia, N. K., Petersen, S. E., and Wig, G. S. (2014). Decreased segregation of brain systems across the healthy adult lifespan.Proceedings of the National Academy of Sciences,111(46), E4997–E5006
2014
-
[7]
and Huang, J
Chen, L. and Huang, J. Z. (2012). Sparse reduced-rank regression for simultaneous di- mension reduction and variable selection.Journal of the American Statistical Association, 107(500), 1533–1545. 26
2012
-
[8]
Cheng, J., Levina, E., Wang, P., and Zhu, J. (2014). A sparse ising model with covariates. Biometrics,70(4), 943–953
2014
Show all 59 references
-
[9]
W., Raamana, P., Spring, R., and Strother, S
Churchill, N. W., Raamana, P., Spring, R., and Strother, S. C. (2017). Optimizing fmri preprocessing pipelines for block-design tasks as a function of age.NeuroImage,154, 240–254
2017
-
[10]
Coombes, B., Basu, S., Guha, S., and Schork, N. (2015). Weighted score tests imple- menting model-averaging schemes in detection of rare variants in case-control studies.Plos one,10(10), e0139355
2015
-
[11]
and Shulman, G
Corbetta, M. and Shulman, G. L. (2002). Control of goal-directed and stimulus-driven attention in the brain.Nature Reviews Neuroscience,3(3), 201–215
2002
-
[12]
Danaher, P., Wang, P., and Witten, D. M. (2014). The joint graphical lasso for inverse covariance estimation across multiple classes.Journal of the Royal Statistical Society Series B: Statistical Methodology,76(2), 373–397
2014
-
[13]
Dolcos, F., Denkova, E., and Dolcos, S. (2012). Neural correlates of emotional memories: A review of evidence from brain imaging studies.Psychologia,55(2), 80–111
2012
-
[14]
Mini-mental state
Folstein, M. F., Folstein, S. E., and McHugh, P. R. (1975). “Mini-mental state”: A practical method for grading the cognitive state of patients for the clinician.Journal of Psychiatric Research,12(3), 189–198
1975
-
[15]
Fosdick, B. K. and Hoff, P. D. (2015). Testing and modeling dependencies between a network and nodal attributes.Journal of the American Statistical Association,110(511), 1047–1056
2015
-
[16]
K., and Chen, K
Goh, G., Dey, D. K., and Chen, K. (2017). Bayesian sparse reduced rank multivariate regression.Journal of multivariate analysis,157, 14–28. 27
2017
-
[17]
and Murphy, T
Gollini, I. and Murphy, T. B. (2016). Joint modeling of multiple network views.Journal of Computational and Graphical Statistics,25(1), 246–265
2016
-
[18]
L., Protzner, A
Grady, C. L., Protzner, A. B., Kovacevic, N., Strother, S. C., Afshin-Pour, B., Wojtow- icz, M., Anderson, J. A. E., Churchill, N., and McIntosh, A. R. (2010). A multivariate analysis of age-related differences in default mode and task-positive networks across mul- tiple cogni...
2010
-
[19]
and Guhaniyogi, R
Guha, S. and Guhaniyogi, R. (2021). Bayesian generalized sparse symmetric tensor-on- vector regression.Technometrics,63(2), 160–170
2021
-
[20]
and Guhaniyogi, R
Guha, S. and Guhaniyogi, R. (2023). Covariate-dependent clustering of undirected networks with brain-imaging data. Technical report
2023
-
[21]
and Guhaniyogi, R
Guha, S. and Guhaniyogi, R. (2024). Covariate-dependent clustering of undirected networks with brain-imaging data.Technometrics, pages 1–23
2024
-
[22]
and Rodriguez, A
Guha, S. and Rodriguez, A. (2021). Bayesian regression with undirected network pre- dictors with an application to brain connectome data.Journal of the American Statistical Association,116(534), 581–593
2021
-
[23]
and Rodriguez, A
Guha, S. and Rodriguez, A. (2023). High-dimensional bayesian network classification with network global-local shrinkage priors.Bayesian Analysis,1(1), 1–30
2023
-
[24]
Guha, S., Rodriguez-Acosta, J., and Dinov, I. D. (2024). A bayesian multiplex graph classifier of functional brain connectivity across diverse tasks of cognitive control.Neu- roinformatics,22(4), 457–472
2024
-
[25]
and Rodriguez, A
Guhaniyogi, R. and Rodriguez, A. (2020). Joint modeling of longitudinal relational data and exogenous variables.Bayesian Analysis,15(2), 477–503
2020
-
[26]
and Spencer, D
Guhaniyogi, R. and Spencer, D. (2021). Bayesian tensor response regression with an application to brain activation studies.Bayesian Analysis,16(4), 1221–1249. 28
2021
-
[27]
Guhaniyogi, R., Qamar, S., and Dunson, D. B. (2017). Bayesian tensor regression. Journal of Machine Learning Research,18(79), 1–31
2017
-
[28]
Guo, J., Levina, E., Michailidis, G., and Zhu, J. (2011). Joint estimation of multiple graphical models.Biometrika,98(1), 1–15
2011
-
[29]
Lee, I., Sinha, D., Mai, Q., Zhang, X., and Bandyopadhyay, D. (2023). Bayesian regres- sion analysis of skewed tensor responses.Biometrics,79(3), 1814–1825
2023
-
[30]
Liu, H., Chen, X., Wasserman, L., and Lafferty, J. (2010). Graph-valued regression. Advances in Neural Information Processing Systems,23
2010
-
[31]
Lukemire, J., Kundu, S., Pagnoni, G., and Guo, Y. (2021). Bayesian joint modeling of multiple brain functional networks.Journal of the American Statistical Association, 116(534), 518–530
2021
-
[32]
S., Cummings, J
Mega, M. S., Cummings, J. L., Fiorello, T., and Gornbein, J. (1996). The spectrum of behavioral changes in alzheimer’s disease.Neurology,46(1), 130–135
1996
-
[33]
Niu, Y., Ni, Y., Pati, D., and Mallick, B. K. (2023). Covariate-assisted bayesian graph learning for heterogeneous data.Journal of the American Statistical Association, pages 1–15
2023
-
[34]
C., Polk, T
Park, D. C., Polk, T. A., Park, R., Minear, M., Savage, A., and Smith, M. R. (2004). Aging reduces neural specialization in ventral visual cortex.Proceedings of the National Academy of Sciences,101(35), 13091–13095
2004
-
[35]
Pessoa, L. (2008). On the relationship between emotion and cognition.Nature Reviews Neuroscience,9(2), 148–158
2008
-
[36]
C., and Vannucci, M
Peterson, C., Stingo, F. C., and Vannucci, M. (2015). Bayesian inference of multiple gaussian graphical models.Journal of the American Statistical Association,110(509), 159–174. 29
2015
-
[37]
and Kadri, H
Rabusseau, G. and Kadri, H. (2016). Low-rank regression with tensor responses.Ad- vances in Neural Information Processing Systems,29
2016
-
[38]
R., Baracchini, G., Nichol, D., Abdi, H., and Grady, C
Rieck, J. R., Baracchini, G., Nichol, D., Abdi, H., and Grady, C. L. (2021). Recon- figuration and dedifferentiation of functional networks during cognitive control across the adult lifespan.Neurobiology of Aging,106, 80–94
2021
-
[39]
Rodriguez-Acosta, J., Guha, S., Gailliot, S., and Williams, A. (2025). Supervised learn- ing with inter- and intra-dependence in multilayer networks with applications in security systems analysis.Technometrics,0(0), 1–14
2025
-
[40]
J., Levina, E., and Zhu, J
Rothman, A. J., Levina, E., and Zhu, J. (2010). Sparse multivariate regression with covariance estimation.Journal of Computational and Graphical Statistics,19(4), 947–962
2010
-
[41]
J., Bullmore, E
Rubia, K., Russell, T., Overmeyer, S., Brammer, M. J., Bullmore, E. T., Sharma, T., Simmons, A., Williams, S. C., Giampietro, V., Andrew, C. M., and Taylor, E. (2001). Map- ping motor inhibition: Conjunctive brain activations across different versions of go/no-go and stop task...
2001
-
[42]
Samanez-Larkin, G. R. and Knutson, B. (2015). Decision making in the ageing brain: changes in affective and motivational circuits.Nature Reviews Neuroscience,16(5), 278– 289
2015
-
[43]
M., Laumann, T
Schaefer, A., Kong, R., Gordon, E. M., Laumann, T. O., Zuo, X.-N., Holmes, A. J., Eickhoff, S. B., and Yeo, B. T. (2018). Local-global parcellation of the human cerebral cortex from intrinsic functional connectivity mri.Cerebral cortex,28(9), 3095–3114
2018
-
[44]
Shen, X., Zhang, H., Li, L., Yang, W., and Liu, L. (2022). Semi-supervised cross-modal hashing with multi-view graph representation.Information Sciences,604, 45–60
2022
-
[45]
Spencer, D., Guhaniyogi, R., and Prado, R. (2020). Joint bayesian estimation of voxel 30 activation and inter-regional connectivity in fmri experiments.psychometrika,85, 845– 869
2020
-
[46]
N., Stevens, W
Spreng, R. N., Stevens, W. D., Viviano, J. D., and Schacter, D. L. (2016). Attenuated anticorrelation between the default and dorsal attention networks with aging: evidence from task and rest.Neurobiology of Aging,45, 149–160
2016
-
[47]
Sun, D., Li, D., Ding, Z., Zhang, X., and Tang, J. (2022). A2ae: Towards adaptive multi-view graph representation learning via all-to-all graph autoencoder architecture. Applied Soft Computing,125, 109193
2022
-
[48]
and Logan, G
Verbruggen, F. and Logan, G. D. (2008). Response inhibition in the stop-signal paradigm.Trends in Cognitive Sciences,12(11), 418–424
2008
-
[49]
J., and Fink, G
Vossel, S., Geng, J. J., and Fink, G. R. (2014). Dorsal and ventral attention systems: Distinct neural circuits but collaborative roles.The Neuroscientist,20(2), 150–159. PMID: 23835449
2014
-
[50]
Wang, P., Robins, G., Pattison, P., and Lazega, E. (2013). Exponential random graph models for multilevel networks.Social networks,35(1), 96–115
2013
-
[51]
Wei, Y., Lei, F., Zhang, Y., Zhao, J., and Liu, K. (2023). Multi-view graph rep- resentation learning for answering hybrid numerical reasoning question.arXiv preprint arXiv:2305.03458
2023 arXiv
-
[52]
Xiao, S., Li, J., Lu, J., Huang, S., Zeng, B., and Wang, S. (2024). Graph neural networks for multi-view learning: a taxonomic review.Artificial Intelligence Review,57(12), 341. 31 Supplementary File: Multiview Graph Fusion with Covariates Abstract This supplementary material ...
2024
-
[54]
˜���� ��( � � � � �� � � ��� ��� � � � ��� � ��� ˜��������� � � � � � � �� � � ��� ��� � ), where� �����=� ������� � �� �����
-
[55]
Furthermore, let: �� = � �� ���� ���� � �� ,� = � �� � ����� � ����� � �� ,� � = � �� ��� �� � �������� �������� ��� �� � � �� , and ˜�� = � �� ˜��˜������ ˜��˜������ � ��
For the update of� � = (� � ����� � ���)� , define: � � = (� ���� ���������� �� ����� � �������� )� ,� � = (� ���� ���������� �� ����� � �������� )� , 11 ����= (����������� ���� ������������� ������������� ���� ����������)� , and ����= (����������� ���� ������������� ���������...
-
[56]
(� ��� ���� ���� ���� ���� ���)�� ����������(� � +�(� ��� � = 0)�1 +�(� ��� � = 1)�1 +�(� ��� � =�1))
-
[57]
14.��� �����(�+ #��:� � = 1�� �+��#��:� � = 1�)
(� ��� ���� ���� ���� ���� ���)�� ����������(� � +�(� ��� � = 0)�1 +�(� ��� � = 1)�1 +�(� ��� � =�1)). 14.��� �����(�+ #��:� � = 1�� �+��#��:� � = 1�). 13 �������� � ROI Names LH.VisCent.Striate.1 LH.VisCent.ExStr.4 LH.VisCent.ExStr.5 LH.VisPeri.ExStrInf.2 LH.VisPeri.ExStrInf....
-
[58]
and Van Der Vaart, A
Ghosal, S. and Van Der Vaart, A. (2007). Convergence rates of posterior distributions for noniid observations.������ �� ����������,��(1), 192–223
2007
-
[59]
Guhaniyogi, R., Qamar, S., and Dunson, D. B. (2017). Bayesian tensor regression. ������� �� ������� �������� ��������,��(79), 1–31
2017
-
[60]
Jiang, W. (2007). Bayesian variable selection for high dimensional generalized linear models: Convergence rates of the fitted densities.������ �� ����������,��(4), 1487 – 1511. 15
2007
Reviewed July 13, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.