Pith. sign in

REVIEW 2 cited by

On the Limitations of Multimodal VAEs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.04121 v2 pith:RSGS7RD2 submitted 2021-10-08 cs.LG

classification cs.LG
keywords multimodalgenerativevaesdataqualityapproachesfundamentallimitations
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Multimodal variational autoencoders (VAEs) have shown promise as efficient generative models for weakly-supervised data. Yet, despite their advantage of weak supervision, they exhibit a gap in generative quality compared to unimodal VAEs, which are completely unsupervised. In an attempt to explain this gap, we uncover a fundamental limitation that applies to a large family of mixture-based multimodal VAEs. We prove that the sub-sampling of modalities enforces an undesirable upper bound on the multimodal ELBO and thereby limits the generative quality of the respective models. Empirically, we showcase the generative quality gap on both synthetic and real data and present the tradeoffs between different variants of multimodal VAEs. We find that none of the existing approaches fulfills all desired criteria of an effective multimodal generative model when applied on more complex datasets than those used in previous benchmarks. In summary, we identify, formalize, and validate fundamental limitations of VAE-based approaches for modeling weakly-supervised data and discuss implications for real-world applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ShaLa: Multimodal Shared Latent Space Modelling

    cs.LG 2025-08 reject novelty 5.0 of 10

    VIPER-R1 fine-tunes a vision-language model to read kinematic plots and propose symbolic equations, then refines them with symbolic regression, but its final metric is computed on the same data used for the refinement fit.

  2. Weakly-Supervised Multimodal Learning on MIMIC-CXR

    cs.LG 2024-11 conditional novelty 4.0 of 10

    MMVM VAE representations outperform other multimodal VAEs and fully supervised baselines on MIMIC-CXR label prediction.

Pith tools