Pith. sign in

REVIEW 3 cited by

Multimodal Generative Models for Scalable Weakly-Supervised Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1802.05335 v3 pith:XHE7D2ZO submitted 2018-02-14 cs.LG stat.ML

classification cs.LGstat.ML
keywords learningmodalitiesmvaegenerativeinferencejointlearnmissing
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multiple modalities often co-occur when describing natural phenomena. Learning a joint representation of these modalities should yield deeper and more useful representations. Previous generative approaches to multi-modal input either do not learn a joint distribution or require additional computation to handle missing data. Here, we introduce a multimodal variational autoencoder (MVAE) that uses a product-of-experts inference network and a sub-sampled training paradigm to solve the multi-modal inference problem. Notably, our model shares parameters to efficiently learn under any combination of missing modalities. We apply the MVAE on four datasets and match state-of-the-art performance using many fewer parameters. In addition, we show that the MVAE is directly applicable to weakly-supervised learning, and is robust to incomplete supervision. We then consider two case studies, one of learning image transformations---edge detection, colorization, segmentation---as a set of modalities, followed by one of machine translation between two languages. We find appealing results across this range of tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Variational Autoencoder Layer

    cs.LG 2026-06 unverdicted novelty 5.0 of 10

    VAEs can be recast as individual neural layers and trained without back-propagation via a multimodal ELBO, yet the resulting shallow classifiers reach only modest accuracy on standard image benchmarks.

  2. Predicting Pulmonary Hypertension in Newborns: A Multi-view VAE Approach

    eess.IV 2025-07 conditional novelty 4.0 of 10

    A multi-view VAE pretrained on newborn echocardiograms achieves higher balanced accuracy for PH severity grading than the supervised baseline, but falls behind it for binary PH detection.

  3. Latent Sensor Fusion: Multimedia Learning of Physiological Signals for Resource-Constrained Devices

    eess.SP 2025-07 conditional novelty 4.0 of 10

    A single image-pretrained VQ-VAE encoder, applied to spectrograms of six physiological signals, matches a modality-specific fusion baseline on WESAD stress classification while using less compute and memory.

Pith tools