Pith. sign in

REVIEW 6 cited by

Multimodal Machine Learning: A Survey and Taxonomy

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1705.09406 v2 pith:EDXTXKWQ submitted 2017-05-26 cs.LG

classification cs.LG
keywords multimodallearningmachinetaxonomyfieldfusionidentifymodalities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Our experience of the world is multimodal - we see objects, hear sounds, feel texture, smell odors, and taste flavors. Modality refers to the way in which something happens or is experienced and a research problem is characterized as multimodal when it includes multiple such modalities. In order for Artificial Intelligence to make progress in understanding the world around us, it needs to be able to interpret such multimodal signals together. Multimodal machine learning aims to build models that can process and relate information from multiple modalities. It is a vibrant multi-disciplinary field of increasing importance and with extraordinary potential. Instead of focusing on specific multimodal applications, this paper surveys the recent advances in multimodal machine learning itself and presents them in a common taxonomy. We go beyond the typical early and late fusion categorization and identify broader challenges that are faced by multimodal machine learning, namely: representation, translation, alignment, fusion, and co-learning. This new taxonomy will enable researchers to better understand the state of the field and identify directions for future research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How Much Reconstruction Does Quantum Machine Learning Need? Late Fusion of Independently Trained Quantum Subcircuits

    quant-ph 2026-08 conditional novelty 6.0 of 10

    Late fusion of independently trained quantum subcircuits matches exact circuit-cutting reconstruction accuracy on tested tasks while avoiding exponential sampling overhead, with a diagnostic indicating when reconstruc...

  2. Modeling Local, Global, and Cross-Modal Context in Multimodal 3D MRI

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    MICViT outperforms CNN and transformer baselines on brain age prediction from multimodal 3D MRI by combining modality-specific and cross-modal local/global attention across three heterogeneous datasets.

  3. Neural Conjugate Aggregation: Identifiable Unsupervised Multi-Sensor Regression under Heterogeneous Sensor Bias

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    NCAM is a hierarchical Bayesian model using neural networks and conjugate Gaussian inference to learn sensor-specific biases for unsupervised multi-source regression, with added conformal prediction for coverage guarantees.

  4. Time Series Foundation Models for Multivariate Financial Time Series Forecasting

    q-fin.GN 2025-07 reject novelty 6.0 of 10

    Pretrained TTM shows large transfer and sample-efficiency gains in three financial forecasting tasks relative to training from scratch, but methodological flaws including possible look-ahead bias weaken the quantitati...

  5. NAROCE: A Neural Algorithmic Reasoner Framework for Online Complex Event Detection

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A two-stage framework learns complex event rules from LLM-generated pseudo traces and then maps sensor embeddings into that rule space, matching a stronger baseline with half the labels on a synthetic benchmark.

  6. Foundation Models for Astrophysics

    astro-ph.IM 2026-08 conditional novelty 3.0 of 10

    Astronomical 'foundation models' largely reuse transformers and self-supervised pretraining, but evidence of transfer to new instruments, populations, or tasks remains rare; the paper argues such evidence, not archite...

Pith tools