REVIEW 2 cited by
On the Limitations of Multimodal VAEs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Multimodal variational autoencoders (VAEs) have shown promise as efficient generative models for weakly-supervised data. Yet, despite their advantage of weak supervision, they exhibit a gap in generative quality compared to unimodal VAEs, which are completely unsupervised. In an attempt to explain this gap, we uncover a fundamental limitation that applies to a large family of mixture-based multimodal VAEs. We prove that the sub-sampling of modalities enforces an undesirable upper bound on the multimodal ELBO and thereby limits the generative quality of the respective models. Empirically, we showcase the generative quality gap on both synthetic and real data and present the tradeoffs between different variants of multimodal VAEs. We find that none of the existing approaches fulfills all desired criteria of an effective multimodal generative model when applied on more complex datasets than those used in previous benchmarks. In summary, we identify, formalize, and validate fundamental limitations of VAE-based approaches for modeling weakly-supervised data and discuss implications for real-world applications.
Forward citations
Cited by 2 Pith papers
-
ShaLa: Multimodal Shared Latent Space Modelling
VIPER-R1 fine-tunes a vision-language model to read kinematic plots and propose symbolic equations, then refines them with symbolic regression, but its final metric is computed on the same data used for the refinement fit.
-
Weakly-Supervised Multimodal Learning on MIMIC-CXR
MMVM VAE representations outperform other multimodal VAEs and fully supervised baselines on MIMIC-CXR label prediction.
Discussion (0). Continue with ORCID to comment.