Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Meta Fusion: A Unified Framework For Multimodality Fusion with Mutual Learning

T0 review · 3 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Mutual learning cuts generalization error, paper proves

desk verdict A well-written framework paper with a real but narrow theorem; the adaptive screening that defines the method is not the part that gets proved, and the lack of code makes the empirical wins hard to audit. read the letter →

arxiv 2507.20089 v1 pith:WO2UV3ZQ submitted 2025-07-27 cs.LG stat.MEstat.ML

classification cs.LGstat.MEstat.ML
keywords multimodalfusiondeepmutuallearningsoftinformationsharingensemblerepresentationgeneralizationerrortheory
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces Meta Fusion, a framework that treats multimodal fusion as a cohort of student models, each built from different combinations of latent representations of the available modalities, and trains them together by soft information sharing. The claim is that this approach automatically decides both when to fuse and what to fuse, with early, intermediate, and late fusion recovered as special cases. The central theoretical result is that aligning a weaker student's outputs with a stronger student's outputs lowers a specific component of the generalization error, the aleatoric variance, without inflating bias or epistemic variance, under stated conditions on the latent representations. If correct, this offers a principled justification for deep mutual learning in multimodal settings and a strategy that adapts to noisy or uninformative modalities.

What carries the argument

The load-bearing mechanism is the disagreement penalty in the combined loss $L = \|Y - V_I\theta_I\|^2 + \|Y - V_J\theta_J\|^2 + \rho\|V_I\theta_I - V_J\theta_J\|^2$. The proof of Theorem 1 relies on a signal-plus-noise model for fused representations ($V_I = VT_I + \epsilon_I$ with isotropic Gaussian noise, $T_I$ with orthogonal columns) and Lemma A4, which shows that the inner product term in $\Xi$ is strictly less than 1, forcing $\Xi<0$. The adaptive mutual learning step then uses K-Means clustering of initial validation losses to set divergence weights $d_{I,J} = 1$ only for top-performing peers.

What would settle it

Repeat the two-student deep linear network experiment with fused representations deliberately constructed to violate Assumption 1, for example by using an isotropic noise term of large variance or a non-orthogonal transformation matrix $T_I$, and check whether the derivative of the generalization error with respect to $\rho$ is no longer negative at $\rho=0$.

Watch

Extended reading notes

Core claim

The central claim is that soft information sharing between two student models reduces the generalization error of each student. Formally, for two deep linear networks with MSE loss and a disagreement penalty $\rho$, Theorem 1 decomposes student I's generalization error as $B^2 + V_a + V_e + \sigma^{*2}$ and shows that at $\rho=0$, the derivative of the bias term is 0, the derivative of the epistemic variance is $O_p(n^{-1/2})$, and the derivative of the aleatoric variance is negative: $\frac{d}{d\rho}V_a = \Xi + O_p(n^{-3/2})$ with $\Xi<0$. The paper further claims that a cohort built from all valid cross-modal pairings of latent representations unifies early, intermediate, and late fusion, and that its adaptive mutual learning step, where students learn only from top-performing peers identified by K-Means clustering, outperforms non-adaptive mutual learning.

Load-bearing premise

The proof that mutual learning reduces generalization error depends on each student's fused representation being a linear transformation of the oracle latent factors plus isotropic Gaussian noise, with the transformation having orthogonal columns, so the features are independent in the test-point calculation.

Editorial extensions

If this is right

  • A principled justification for deep mutual learning: the disagreement penalty acts by reducing intrinsic variance rather than bias or epistemic variance.
  • A unified framework that automatically interpolates between early, intermediate, and late fusion, with cooperative learning as a special case.
  • Adaptive mutual learning can mitigate negative knowledge transfer when some students are noisy, as supported by the ablation comparing learning from top versus low performers.
  • The framework is model-agnostic and extends to more than two modalities through cross-modal pairing of latent representations.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The proof's condition that each student's fused representation is a linear transformation of the oracle latent factors plus isotropic noise is restrictive; in practice deep network embeddings may not satisfy this, so the negative-derivative result may not hold for more complex representations.
  • The theoretical guarantee applies to a cohort of exactly two students with $d_{I,J}=1$, while the adaptive K-Means screening is the main methodological novelty and lies outside this proof; a natural test would be to check whether the benefit persists with more than two students.
  • The framework's focus on output-level sharing suggests it could be adapted to privacy-preserving or federated settings, as the paper itself hints and demonstrates with preliminary discussion.
  • The theorem's favorable regime requires that student features be mutually supportive; Corollary 1's condition that the angle between $v_I$ and $v_J$ is not too large suggests a practical diagnostic: monitor this alignment to decide when mutual learning will help.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes Meta Fusion, a multimodal fusion framework that constructs a cohort of student models from different combinations of latent representations, trains the cohort with an adaptive mutual-learning objective that aligns each student with top-performing peers identified by K-Means screening of initial validation losses, and then aggregates selected students via ensemble selection. The authors argue that the framework unifies early, intermediate, late, and cooperative fusion as special cases. The main theoretical result, Theorem 1, states that for two deep linear students under Gaussian signal-plus-noise assumptions, the generalization error decomposes into bias, aleatoric variance, epistemic variance, and an oracle term, and that the derivative of the aleatoric variance with respect to the mutual-learning strength rho is negative at rho=0 while the other derivatives vanish to leading order. The paper supports this with extensive simulation studies and two real-data applications: Alzheimer's disease detection on NACC data and neural decoding of hippocampal spike/LFP data.

Significance. If the theoretical claim holds, the paper would provide one of the first formal justifications for deep mutual learning in multimodal fusion, and the proposed framework is broad enough to be of practical interest. The proofs in Appendix A4 are detailed and internally coherent under the stated assumptions, and the paper is honest about the stylized setting used in the theory. The simulation study includes useful ablations, including a comparison of divergence weights that demonstrates the benefit of adaptive weighting, and the real-data applications are relevant. The main limitation is a scope mismatch: the theorem covers only two-student symmetric mutual learning without the adaptive K-Means screening, so the central novelty of Meta Fusion is not actually analyzed by the proof.

major comments (3)
  1. [Section 4.2, Theorem 1] The headline claim that soft information sharing reduces generalization error is proved only for a cohort of exactly two students with dI,J=1 and P={I,J}, and the derivative is evaluated only at rho=0. The adaptive K-Means screening in Eq. (5), the data-dependent zero/one divergence weights, and the ensemble selection in Algorithm 2 are entirely outside the proof. Since those components are what distinguish Meta Fusion from standard symmetric mutual learning, the theorem does not establish the abstract's claim about the proposed adaptive mechanism. This is a load-bearing gap: either the theorem needs to be extended to cover the adaptive weights, or the paper's claims need to be explicitly restricted to symmetric two-student mutual learning.
  2. [Assumptions 1-2 (Section 4.2)] Theorem 1 relies on Assumption 1 (VI = V TI + epsilonI with isotropic Gaussian noise and diagonal SigmaI) and Assumption 2 (orthogonal columns of TI). These assumptions are not satisfied by the empirical pipeline: PCA projections and raw identity features produce correlated, non-orthogonal columns, and deep network embeddings are nonlinear functions of the inputs. The negative sign of Xi in Lemma A4 is obtained by diagonalizing the feature covariance in the test-point calculation. Consequently, the theory provides no bound on negative transfer when the screening step selects an uninformative peer, and it does not guarantee the observed real-data behavior. The authors should either weaken these assumptions, provide a robustness analysis, or supply empirical diagnostics showing that the assumptions approximately hold in their simulations.
  3. [Sections 5-6 vs. Section 4] The theory is developed for MSE regression with deep linear networks, but the real-data applications are classification problems using cross-entropy loss and KL divergence, with nonlinear encoders. Theorem 1 cannot be directly invoked to explain the classification results in Section 6. The paper should acknowledge this transfer gap explicitly and, if the theoretical contribution is meant to support the real-data claims, provide a classification analogue or state clearly that the theory applies only to the regression setting.
minor comments (4)
  1. [Section 3.1.1] The dummy null extractor is defined twice with the same name: 'g0_x(X):=∅ and g0_x(X):=∅'; the second should presumably be g0_z(Z):=∅.
  2. [Section 4.1] 'For matries MI,MJ' should read 'matrices'.
  3. [Theorem 1, paragraph after the theorem] 'an increase in the disagreement penalty ρ does not effect the bias' should be 'does not affect the bias'.
  4. [Proof of Corollary 1] The proof states 'Assume without loss of generality that TI,TJ≥0' but also implicitly assumes theta≥0 when concluding ar heta*_I, ar heta*_J, theta*_I≥0; this additional condition on theta should be stated in the corollary.

Circularity Check

0 steps flagged · score 0.0 of 10

No meaningful circularity: Theorem 1 is a conditional derivation from explicitly stated assumptions, and the empirical comparisons use standard validation-based model selection.

full rationale

I walked the paper's derivation chain and found no circular step that reduces a prediction or first-principles result to its own inputs. Theorem 1 is a conditional mathematical statement: under Assumptions 1–2 (Eq. 7 and the orthogonal-column condition), the proof expands the global minimizer from Proposition 1, decomposes the MSE into bias, aleatoric variance, and epistemic variance, and evaluates derivatives at rho = 0. The negative sign of d/drho Va is obtained from Lemma A4, which is proved from the stated assumptions; it is not imported from a fitted value, from a prior result by the same authors, or from the conclusion itself. No fitted parameter is renamed as a prediction: rho, k_cls, k_top, and committee sizes are selected by validation, which is ordinary model selection rather than a circular fit. The paper's self-citations (e.g., [51] for supervised encoders, [57] for the neural decoding dataset) supply components or data and are not load-bearing for the theoretical claim. Appendix A2's equivalence between cooperative learning and a two-student special case of Meta Fusion is a structural identity between two loss objectives (A12) and (A14), not a circular reduction of a novel prediction to its input. The gap between Theorem 1, which treats exactly two students with dI,J = 1, and the full adaptive K-Means screening in Eq. (5) is a scope or correctness concern about how much of the method the theorem justifies; it is not evidence that the theorem presupposes the conclusion. Under the hard rules, this mismatch does not constitute circularity.

Assumptions & free parameters 6 free parameters · 6 assumptions · 0 invented entities

The theory rests on a specific Gaussian linear generative model and orthogonality assumptions that are not checked in the real-data experiments. The main methodological novelty (adaptive teacher selection) is outside the theorem, and the empirical pipeline has several user-chosen hyperparameters. There are no invented entities.

free parameters (6)
  • Mutual learning strength rho = chosen via cross-validation
    Balances task loss and divergence loss in Eq. (4); the theory analyzes only the derivative at rho=0, while simulations and applications use rho>0 selected on validation data.
  • Number of performance clusters k_cls = chosen via Silhouette or Elbow
    Used in the initial screening step (Section 3.1.2); no fixed rule is given for the experiments.
  • Number of top clusters k_top = user-specified (>=1)
    Defines the teacher set S_top in Eq. (5) that others learn from.
  • Committee selection hyperparameters p_prune, n_init, n_c = user-specified
    Control Algorithm 2's pruning and ensemble selection; their values are not reported.
  • Latent dimensions of feature extractors = manual grids (e.g., LFP dims 0 to 330, spike dims 0 to 54 in Figure A6)
    The cohort definition requires choosing extractor output dimensions; in the neural decoding experiment these are swept over a grid, and in simulations PCA dimensions are selected without stated criteria.
  • Number of extractors k_x, k_z = user-specified
    Determines cohort size (k_x+2)(k_z+2)-1; not data-driven in the experiments.
assumptions (6)
  • domain assumption Ground-truth labels follow a linear latent factor model Y = V theta with V rows i.i.d. N(0, I_p) (Section 4.2, Eq. 6).
    This is the data-generating model for the theory; real datasets need not satisfy it.
  • domain assumption Assumption 1: V_I = V T_I + epsilon_I with epsilon_I rows i.i.d. N(0, sigma_I^2 I) (Eq. 7).
    The fused representations are linear transformations of the oracle factors plus isotropic noise; this is the core premise for the variance decomposition.
  • domain assumption Assumption 2: T_I has orthogonal columns (Assumption 2, Section 4.2).
    Orthogonality eliminates cross-feature terms in the test-point expectation and makes the sign of Xi negative in Lemma A4.
  • ad hoc to paper Student models are deep linear networks and only two students, P={I,J}, with d_{I,J}=1, are analyzed (Section 4.2).
    The theorem covers standard mutual learning in a linear two-student setting; the proposed adaptive weighting and ensemble selection are not part of the proof.
  • standard math The power series expansion (I - rho E^{-1} F)^{-1} E^{-1} is valid and the higher-order terms H(rho) are independent of rho (Appendix A4.3, proof of Proposition 1).
    The proof relies on Neumann series and block matrix invertibility for small rho.
  • standard math Asymptotic approximations V_II = n Sigma_tilde_I + O_p(sqrt(n)) etc. from the central limit theorem (Eq. A37).
    Used to convert the matrix expressions into the deterministic limiting derivative Xi.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Meta Fusion: A Unified Framework For Multimodality Fusion with Mutual Learning." pith.science (2026). https://pith.science/paper/WO2UV3ZQ

@misc{pith2026250720089,
  author       = {Pith},
  title        = {Pith review of: Meta Fusion: A Unified Framework For Multimodality Fusion with Mutual Learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WO2UV3ZQ}},
  note         = {Machine review of arXiv:2507.20089}
}
read the original abstract

Developing effective multimodal data fusion strategies has become increasingly essential for improving the predictive power of statistical machine learning methods across a wide range of applications, from autonomous driving to medical diagnosis. Traditional fusion methods, including early, intermediate, and late fusion, integrate data at different stages, each offering distinct advantages and limitations. In this paper, we introduce Meta Fusion, a flexible and principled framework that unifies these existing strategies as special cases. Motivated by deep mutual learning and ensemble learning, Meta Fusion constructs a cohort of models based on various combinations of latent representations across modalities, and further boosts predictive performance through soft information sharing within the cohort. Our approach is model-agnostic in learning the latent representations, allowing it to flexibly adapt to the unique characteristics of each modality. Theoretically, our soft information sharing mechanism reduces the generalization error. Empirically, Meta Fusion consistently outperforms conventional fusion strategies in extensive simulation studies. We further validate our approach on real-world applications, including Alzheimer's disease detection and neural decoding.

Figures

Figures reproduced from arXiv: 2507.20089 by the authors.

Figure 1
Figure 1. Overview of classical multimodal fusion strategies. (a) Example of multimodal [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Overview of the Meta Fusion pipeline. mapping, g 0 x (X) := ∅ and g 0 x (X) := ∅, which excludes the corresponding modality. Given the extractors {g i x (X)}i∈{0,...,kx+1} and {g j z (Z)}j∈{0,...,kz+1}, we fuse their outputs to construct student models by using a cross-modality pairing strategy, illustrated in the left panel of [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Average accuracy of Meta Fusion and benchmarks on the NACC dataset. Error [PITH_FULL_IMAGE:figures/full_fig_p017_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Classification accuracy of Meta Fusion and top benchmarks (Spike-only and Late [PITH_FULL_IMAGE:figures/full_fig_p018_4.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DAIF: A Data-Driven Intermediate Fusion Framework for Multimodal Supervised Learning via Approximate Message Passing

    stat.ME 2026-08 conditional novelty 7.0 of 10

    DAIF adaptively selects fusion granularity via CKA clustering and clusterwise empirical Bayes AMP denoising, with consistency, state-evolution, and Bayes-optimality theorems.

Reference graph

Works this paper leans on

63 extracted references · 62 canonical work pages · cited by 1 Pith paper

  1. [1]

    Fuse- MODNet: Real-Time Camera and LiDAR Based Moving Object Detection for Ro- bust Low-Light Autonomous Driving

    H. Rashed, M. Ramzy, V. Vaquero, A. El Sallab, G. Sistu, and S. Yogamani. “Fuse- MODNet: Real-Time Camera and LiDAR Based Moving Object Detection for Ro- bust Low-Light Autonomous Driving”. In:Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops . Oct. 2019

  2. [2]

    Sensors and Sensor Fusion in Autonomous Vehicles

    J. Koci´ c, N. Joviˇ ci´ c, and V. Drndarevi´ c. “Sensors and Sensor Fusion in Autonomous Vehicles”. In: 2018 26th Telecommunications Forum (TELFOR). 2018, pp. 420–425

  3. [3]

    Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph

    A. Bagher Zadeh, P. P. Liang, S. Poria, E. Cambria, and L.-P. Morency. “Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph”. In: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Ed. by I. Gurevych and Y. Miyao. Melbourne, Australia: Association ...

  4. [4]

    Multimodal Sentiment Analysis: A Survey of Methods, Trends, and Challenges

    R. Das and T. D. Singh. “Multimodal Sentiment Analysis: A Survey of Methods, Trends, and Challenges”. In: 55.13s (July 2023)

  5. [5]

    Multimodal sentiment analysis: A systematic review of history, datasets, multimodal fusion methods, appli- cations, challenges and future directions

    A. Gandhi, K. Adhvaryu, S. Poria, E. Cambria, and A. Hussain. “Multimodal sentiment analysis: A systematic review of history, datasets, multimodal fusion methods, appli- cations, challenges and future directions”. In: Information Fusion 91 (2023), pp. 424– 444

  6. [6]

    Multimodal classification of Alzheimer’s disease and mild cognitive impairment

    D. Zhang, Y. Wang, L. Zhou, H. Yuan, and D. Shen. “Multimodal classification of Alzheimer’s disease and mild cognitive impairment”. In: NeuroImage 55.3 (2011), pp. 856–867

  7. [7]

    Multimodal deep learning for Alzheimer’s disease dementia assessment

    S. Qiu, M. I. Miller, P. S. Joshi, J. C. Lee, C. Xue, Y. Ni, Y. Wang, I. De Anda- Duran, P. H. Hwang, J. A. Cramer, B. C. Dwyer, H. Hao, M. C. Kaku, S. Kedar, P. H. Lee, A. Z. Mian, D. L. Murman, S. O’Shea, A. B. Paul, M.-H. Saint-Hilaire, E. Alton Sartor, A. R. Saxena, L. C. Shih, J. E. Small, M. J. Smith, A. Swaminathan, C. E. Takahashi, O. Taraschenko,...

  8. [8]

    Fusion of medical imaging and electronic health records using deep learning: a systematic review and implementation guidelines

    S.-C. Huang, A. Pareek, S. Seyyedi, I. Banerjee, and M. P. Lungren. “Fusion of medical imaging and electronic health records using deep learning: a systematic review and implementation guidelines”. In: npj Digital Medicine 3.1 (Oct. 2020)

Show all 63 references
  1. [9]

    Deep multimodal fusion for se- mantic image segmentation: A survey

    Y. Zhang, D. Sidib´ e, O. Morel, and F. M´ eriaudeau. “Deep multimodal fusion for se- mantic image segmentation: A survey”. In: Image and Vision Computing 105 (2021), p. 104042

  2. [10]

    Indoor Semantic Segmentation using depth information

    C. Couprie, C. Farabet, L. Najman, and Y. LeCun. “Indoor Semantic Segmentation using depth information”. In: arXiv: Computer Vision and Pattern Recognition (2013)

  3. [11]

    Multimodal End-to-End Autonomous Driving

    Y. Xiao, F. Codevilla, A. Gurram, O. Urfalioglu, and A. M. L´ opez. “Multimodal End-to-End Autonomous Driving”. In: Trans. Intell. Transport. Sys. 23.1 (Jan. 2022), pp. 537–547. 20

  4. [12]

    Fusion of deep learning models of MRI scans, Mini–Mental State Examination, and logical memory test enhances diagnosis of mild cognitive impairment

    S. Qiu, G. H. Chang, M. Panagia, D. M. Gopal, R. Au, and V. B. Kolachalama. “Fusion of deep learning models of MRI scans, Mini–Mental State Examination, and logical memory test enhances diagnosis of mild cognitive impairment”. In:Alzheimers Dement (Amst) 10.1 (Jan. 2018), pp. 737–749

  5. [13]

    Deep Learning Role in Early Diagnosis of Prostate Cancer

    I. Reda, A. Khalil, M. Elmogy, A. Abou El-Fetouh, A. Shalaby, M. Abou El-Ghar, A. Elmaghraby, M. Ghazal, and A. El-Baz. “Deep Learning Role in Early Diagnosis of Prostate Cancer”. In: Technology in Cancer Research Treatment 17 (Jan. 2018)

  6. [14]

    Deep learning of brain lesion patterns and user-defined clinical and MRI fea- tures for predicting conversion to multiple sclerosis from clinically isolated syndrome

    Y. Yoo, L. Y. W. Tang, D. K. B. Li, L. Metz, S. Kolind, A. L. Traboulsee, and R. C. Tam. “Deep learning of brain lesion patterns and user-defined clinical and MRI fea- tures for predicting conversion to multiple sclerosis from clinically isolated syndrome”. In: Computer Method...

  7. [15]

    Cooperative learning for mul- tiview analysis

    D. Y. Ding, S. Li, B. Narasimhan, and R. Tibshirani. “Cooperative learning for mul- tiview analysis”. In: Proceedings of the National Academy of Sciences 119.38 (2022), e2202113119

  8. [16]

    A Deep Learning Mammography-based Model for Improved Breast Cancer Risk Prediction

    A. Yala, C. Lehman, T. Schuster, T. Portnoi, and R. Barzilay. “A Deep Learning Mammography-based Model for Improved Breast Cancer Risk Prediction”. In: Radi- ology 292.1 (July 2019), pp. 60–66

  9. [17]

    Multi-Channel 3D Deep Feature Learning for Survival Time Prediction of Brain Tumor Patients Using Multi-Modal Neuroimages

    D. Nie, J. Lu, H. Zhang, E. Adeli, J. Wang, Z. Yu, L. Liu, Q. Wang, J. Wu, and D. Shen. “Multi-Channel 3D Deep Feature Learning for Survival Time Prediction of Brain Tumor Patients Using Multi-Modal Neuroimages”. In: Scientific Reports 9.1 (Jan. 2019)

  10. [18]

    An Efficient Multi-View Multimodal Data Processing Framework for Social Media Popularity Prediction

    Y. Tan, F. Liu, B. Li, Z. Zhang, and B. Zhang. “An Efficient Multi-View Multimodal Data Processing Framework for Social Media Popularity Prediction”. In: Proceedings of the 30th ACM International Conference on Multimedia . MM ’22. Lisboa, Portugal: Association for Computing Ma...

  11. [19]

    Multimodal Learning with Deep Boltzmann Machines

    N. Srivastava and R. R. Salakhutdinov. “Multimodal Learning with Deep Boltzmann Machines”. In: Advances in Neural Information Processing Systems. Ed. by F. Pereira, C. Burges, L. Bottou, and K. Weinberger. Vol. 25. Curran Associates, Inc., 2012

  12. [20]

    Multimodal generative models for scalable weakly-supervised learning

    M. Wu and N. Goodman. “Multimodal generative models for scalable weakly-supervised learning”. In: NIPS’18. Montr´ eal, Canada: Curran Associates Inc., 2018, pp. 5580–5590

  13. [21]

    Unity by Diversity: Improved Representation Learning in Multimodal VAEs

    T. M. Sutter, Y. Meng, N. Fortin, J. E. Vogt, B. Shahbaba, and S. Mandt. “Unity by Diversity: Improved Representation Learning in Multimodal VAEs”. In: Sixth Sympo- sium on Advances in Approximate Bayesian Inference - Non Archival Track . 2024

  14. [22]

    Deep Mutual Learning

    Y. Zhang, T. Xiang, T. M. Hospedales, and H. Lu. “Deep Mutual Learning”. In: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (2017), pp. 4320–4328

  15. [23]

    Bagging predictors

    L. Breiman. “Bagging predictors”. In: Machine Learning 24.2 (Aug. 1996), pp. 123– 140. 21

  16. [24]

    Ensemble selection from libraries of models

    R. Caruana, A. Niculescu-Mizil, G. Crew, and A. Ksikes. “Ensemble selection from libraries of models”. In: Proceedings of the Twenty-First International Conference on Machine Learning . ICML ’04. Banff, Alberta, Canada: Association for Computing Machinery, 2004, p. 18

  17. [25]

    Regression Shrinkage and Selection via the Lasso

    R. Tibshirani. “Regression Shrinkage and Selection via the Lasso”. In: Journal of the Royal Statistical Society. Series B (Methodological) 58.1 (1996), pp. 267–288

  18. [26]

    Model compression

    C. Bucilu˘ a, R. Caruana, and A. Niculescu-Mizil. “Model compression”. In: Proceedings of the 12th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. KDD ’06. Philadelphia, PA, USA: Association for Computing Machinery, 2006, pp. 535–541

  19. [27]

    Distilling the Knowledge in a Neural Network

    G. E. Hinton, O. Vinyals, and J. Dean. “Distilling the Knowledge in a Neural Network”. In: ArXiv abs/1503.02531 (2015)

  20. [28]

    Modality-Aware Mutual Learning for Multi-modal Medical Image Segmentation

    Y. Zhang, J. Yang, J. Tian, Z. Shi, C. Zhong, Y. Zhang, and Z. He. “Modality-Aware Mutual Learning for Multi-modal Medical Image Segmentation”. In: Medical Image Computing and Computer Assisted Intervention – MICCAI 2021: 24th International Conference, Strasbourg, France, Sept...

  21. [29]

    AM ³Net: Adaptive Mutual-Learning-Based Multimodal Data Fusion Network

    J. Wang, J. Li, Y. Shi, J. Lai, and X. Tan. “AM ³Net: Adaptive Mutual-Learning-Based Multimodal Data Fusion Network”. In: IEEE Transactions on Circuits and Systems for Video Technology 32.8 (2022), pp. 5411–5426

  22. [30]

    Multi-modal contrastive mutual learning and pseudo-label re-learning for semi-supervised medical image seg- mentation

    S. Zhang, J. Zhang, B. Tian, T. Lukasiewicz, and Z. Xu. “Multi-modal contrastive mutual learning and pseudo-label re-learning for semi-supervised medical image seg- mentation”. In: Medical Image Analysis 83 (2023), p. 102656

  23. [31]

    Graph neural networks with deep mutual learning for designing multi-modal recommendation systems

    J. Li, C. Yang, G. Ye, and Q. V. H. Nguyen. “Graph neural networks with deep mutual learning for designing multi-modal recommendation systems”. In:Information Sciences 654 (2024), p. 119815

  24. [32]

    Deep Residual Learning for Image Recognition

    K. He, X. Zhang, S. Ren, and J. Sun. “Deep Residual Learning for Image Recognition”. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2015), pp. 770–778

  25. [33]

    BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”. In:North American Chapter of the Association for Computational Linguistics . 2019

  26. [34]

    A unified theory of diversity in ensemble learning

    D. Wood, T. Mu, A. M. Webb, H. W. J. Reeve, M. Luj´ an, and G. Brown. “A unified theory of diversity in ensemble learning”. In: J. Mach. Learn. Res. 24.1 (Mar. 2024)

  27. [35]

    A review of feature set partitioning methods for multi-view ensemble learning

    A. Kumar and J. Yadav. “A review of feature set partitioning methods for multi-view ensemble learning”. In: Information Fusion 100 (2023), p. 101959

  28. [36]

    Silhouettes: A graphical aid to the interpretation and validation of cluster analysis

    P. J. Rousseeuw. “Silhouettes: A graphical aid to the interpretation and validation of cluster analysis”. In: Journal of Computational and Applied Mathematics 20 (1987), pp. 53–65. 22

  29. [37]

    The Determination of Cluster Number at k-Mean Using Elbow Method and Purity Evaluation on Headline News

    D. Marutho, S. Hendra Handaka, E. Wijaya, and Muljono. “The Determination of Cluster Number at k-Mean Using Elbow Method and Purity Evaluation on Headline News”. In: 2018 International Seminar on Application for Technology of Information and Communication. 2018, pp. 533–538

  30. [38]

    Getting the Most Out of Ensem- ble Selection

    R. Caruana, A. Munson, and A. Niculescu-Mizil. “Getting the Most Out of Ensem- ble Selection”. In: Sixth International Conference on Data Mining (ICDM’06) . 2006, pp. 828–833

  31. [39]

    K. P. Murphy. Machine Learning: A Probabilistic Perspective. The MIT Press, 2012

  32. [40]

    Exact solutions to the nonlinear dy- namics of learning in deep linear neural networks

    A. M. Saxe, J. L. McClelland, and S. Ganguli. “Exact solutions to the nonlinear dy- namics of learning in deep linear neural networks”. In: International Conference on Learning Representations (ICLR) (2014)

  33. [41]

    Deep Learning without Poor Local Minima

    K. Kawaguchi. “Deep Learning without Poor Local Minima”. In: Advances in Neural Information Processing Systems. Ed. by D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett. Vol. 29. Curran Associates, Inc., 2016

  34. [42]

    Identity Matters in Deep Learning

    M. Hardt and T. Ma. “Identity Matters in Deep Learning”. In: International Confer- ence on Learning Representations. 2017

  35. [43]

    Towards Understanding Knowledge Distillation

    M. Phuong and C. H. Lampert. “Towards Understanding Knowledge Distillation”. In: International Conference on Machine Learning . 2019

  36. [44]

    The Revised National Alzheimer’s Coordinating Center’s Neuropathology Form—Available Data and New Analyses

    L. M. Besser, W. A. Kukull, M. A. Teylan, E. H. Bigio, N. J. Cairns, J. K. Kofler, T. J. Montine, J. A. Schneider, and P. T. Nelson. “The Revised National Alzheimer’s Coordinating Center’s Neuropathology Form—Available Data and New Analyses”. In: Journal of Neuropathology & Ex...

  37. [45]

    Version 3 of the Alzheimer Disease Centers’ Neuropsychological Test Battery in the Uniform Data Set (UDS)

    S. Weintraub, L. Besser, H. H. Dodge, M. Teylan, S. Ferris, F. C. Goldstein, B. Giordani, J. Kramer, D. Loewenstein, D. Marson, D. Mungas, D. Salmon, K. Welsh- Bohmer, X.-H. Zhou, S. D. Shirk, A. Atri, W. A. Kukull, C. Phelps, and J. C. Morris. “Version 3 of the Alzheimer Dise...

  38. [46]

    “Mini-mental state

    M. F. Folstein, S. E. Folstein, and P. R. McHugh. ““Mini-mental state”: A practical method for grading the cognitive state of patients for the clinician”. In: Journal of Psychiatric Research 12.3 (1975), pp. 189–198

  39. [47]

    The Neuropsychological Profile of Alzheimer Disease

    S. Weintraub, A. H. Wicklund, and D. P. Salmon. “The Neuropsychological Profile of Alzheimer Disease”. In: Cold Spring Harbor Perspectives in Medicine 2.4 (Jan. 2012), a006171–a006171

  40. [48]

    Cognitive Tests to Detect Dementia: A Systematic Review and Meta-analysis

    K. K. F. Tsoi, J. Y. C. Chan, H. W. Hirai, S. Y. S. Wong, and T. C. Y. Kwok. “Cognitive Tests to Detect Dementia: A Systematic Review and Meta-analysis”. In: JAMA Internal Medicine 175.9 (Sept. 2015), p. 1450

  41. [49]

    Hypothetical model of dynamic biomarkers of the Alzheimer’s pathological cascade

    C. R. Jack, D. S. Knopman, W. J. Jagust, L. M. Shaw, P. S. Aisen, M. W. Weiner, R. C. Petersen, and J. Q. Trojanowski. “Hypothetical model of dynamic biomarkers of the Alzheimer’s pathological cascade”. In: The Lancet Neurology 9.1 (Jan. 2010), pp. 119–128. 23

  42. [50]

    Dementia prevention, intervention, and care: 2020 report of the Lancet Commission

    G. Livingston, J. Huntley, A. Sommerlad, D. Ames, C. Ballard, S. Banerjee, C. Brayne, A. Burns, J. Cohen-Mansfield, C. Cooper, S. G. Costafreda, A. Dias, N. Fox, L. N. Gitlin, R. Howard, H. C. Kales, M. Kivim¨ aki, E. B. Larson, A. Ogunniyi, V. Orgeta, K. Ritchie, K. Rockwood,...

  43. [51]

    Alzheimer’s disease detection using data fusion with a deep supervised encoder

    M. Trinh, R. Shahbaba, C. Stark, and Y. Ren. “Alzheimer’s disease detection using data fusion with a deep supervised encoder”. In: Frontiers in Dementia 3 (Feb. 2024)

  44. [52]

    Episodic and Semantic Memory

    E. Tulving. “Episodic and Semantic Memory”. In: Organization of Memory . Ed. by E. Tulving and W. Donaldson. New York: Academic Press, 1972, pp. 381–403

  45. [53]

    Neural basis of the perception and estimation of time

    H. Merchant, D. L. Harrington, and W. H. Meck. “Neural basis of the perception and estimation of time”. In: Annual review of neuroscience 36 (2013), pp. 313–336

  46. [54]

    Time cells in the hippocampus: a new dimension for mapping mem- ories

    H. Eichenbaum. “Time cells in the hippocampus: a new dimension for mapping mem- ories.” In: Nature Reviews Neuroscience 15 (2014), pp. 732–744

  47. [55]

    Mitzdorf et al

    U. Mitzdorf et al. Current source-density method and application in cat cerebral cor- tex: investigation of evoked potentials and EEG phenomena . American Physiological Society, 1985

  48. [56]

    Nonspatial sequence coding in CA1 neurons

    T. A. Allen, D. M. Salz, S. McKenzie, and N. J. Fortin. “Nonspatial sequence coding in CA1 neurons”. In: Journal of Neuroscience 36.5 (2016), pp. 1547–1563

  49. [57]

    Hippocampal ensembles represent sequential relation- ships among an extended sequence of nonspatial events

    B. Shahbaba, L. Li, F. Agostinelli, M. Saraf, K. W. Cooper, D. Haghverdian, G. A. Elias, P. Baldi, and N. J. Fortin. “Hippocampal ensembles represent sequential relation- ships among an extended sequence of nonspatial events”. In: Nature communications 13.1 (2022), p. 787

  50. [58]

    Federated Learning: Challenges, Methods, and Future Directions

    T. Li, A. K. Sahu, A. Talwalkar, and V. Smith. “Federated Learning: Challenges, Methods, and Future Directions”. In: IEEE Signal Processing Magazine 37.3 (2020), pp. 50–60

  51. [59]

    Inverses of 2 × 2 block matrices

    T.-T. Lu and S.-H. Shiou. “Inverses of 2 × 2 block matrices”. In: Computers & Math- ematics with Applications 43.1 (2002), pp. 119–129. 24 A1 Implementation Details A1.1 Extension to multiple modalities In this section, we describe how to build student cohorts for applications...

  52. [60]

    Generate latent variables: X∗∈ Rn×p∗ x, X∗ ij∼N(0, 1); Z∗∈ Rn×p∗ z, Z∗ ij∼N(0, 1); S∗∈ Rn×p∗ s, S∗ ij∼N(0, 1)

    Represent the latent dimensions for X-specific,Z-specific, and shared information as p∗ x, p∗ z, and p∗ s, respectively. Generate latent variables: X∗∈ Rn×p∗ x, X∗ ij∼N(0, 1); Z∗∈ Rn×p∗ z, Z∗ ij∼N(0, 1); S∗∈ Rn×p∗ s, S∗ ij∼N(0, 1). 25

  53. [61]

    Define the interaction terms using the Kronecker product: U∗∈ Rn×p∗ xp∗ z, U∗ i =X∗ i⊗Z∗ i, whereX∗ i andZ∗ i are the ith rows of X∗ andZ∗

  54. [62]

    For instance, setting ct = 0 removes that component’s influence on Y

    Generate the ground-truth label Y as: Y =cxβxfx(X∗) +czβzfz(Z∗) +csβsfs(S∗) +cuβuU∗, where βt (t∈{ x,z,s,u }) is the vector of coefficients, ft(·) applies an element-wise transformation (e.g., quadratic) to each latent component, and ct is the weight that controls the contribu...

  55. [63]

    Best Single,

    Let px andpz be the observed feature dimensions. The observed features are generated using a signal-plus-noise model based on the latent components: X = (1−rx)[X∗,S∗]Tx +rxϵx, Tx∈ R(p∗ x+p∗ s)×px; Z = (1−rz)[Z∗,S∗]Tz +rzϵz, Tz∈ R(p∗ z+p∗ s)×pz, where [·,·] denotes column-wise ...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.