Pith. sign in

REVIEW 7 cited by

M3-JEPA: Multimodal Alignment via Multi-gate MoE based on the Joint-Embedding Predictive Architecture

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.05929 v6 pith:HIQ6DGUX submitted 2024-09-09 cs.LG cs.AI

M3-JEPA: Multimodal Alignment via Multi-gate MoE based on the Joint-Embedding Predictive Architecture

classification cs.LG cs.AI
keywords m3-jepamultimodalframeworkspacetasksalignmentarchitecturedifferent
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Current multimodal learning strategies primarily optimize in the original token space. Such a framework is easy to incorporate with the backbone of pretrained language model, but might result in modality collapse. To alleviate such issues, we leverage the Joint-Embedding Predictive Architecture (JEPA) on the multimodal tasks, which converts the input embedding into the output embedding space by a predictor and then conducts the cross-modal alignment on the latent space. We implement this predictor by a Multi-Gate Mixture of Experts (MMoE) and name the framework as M3-JEPA, accordingly. The gating function disentangles the modality-specific and shared information and derives information-theoretic optimality. The framework is implemented with both contrastive and regularization loss, and solved by alternative gradient descent (AGD) between different multimodal tasks. By thoroughly designed experiments, we show that M3-JEPA can obtain state-of-the-art performance on different modalities and tasks, generalize to unseen datasets and domains, and is computationally efficient in both training and inference. Our observation suggests that M3-JEPA might become a new basis to self-supervised learning in the open world.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. BrainFIBRE: A Foundation Model via Information Decomposition for Brain Microstructure

    cs.CV 2026-07 conditional novelty 7.0

    BrainFIBRE pretrains a five-expert Mixture-of-Experts model on NODDI-derived microstructural maps and outperforms prior deep models on age, sex, cerebrovascular, neurodegenerative, and cognitive prediction.

  2. MoP-JEPA: Hard-Assigned Predictor Mixtures for Stochastic JEPA World Models

    cs.AI 2026-07 conditional novelty 6.0

    Hard-assigned predictor mixtures let JEPA world models enumerate discrete successor modes, raising verified planning success far above single-output and soft-mixture baselines on stochastic OGBench mazes.

  3. SARM2: Multi-Task Stage Aware Reward Modeling for Self Improving Robotic Manipulation

    cs.RO 2026-06 unverdicted novelty 6.0

    SARM2 presents RM, a multi-task stage-aware reward model achieving 80% lower value-estimation MSE, which when used in SPIRAL boosts manipulation task success from ~50% to near-perfect on several benchmarks.

  4. DART: A Vision-Language Foundation Model for Comprehensive Rope Condition Monitoring

    cs.CV 2026-05 unverdicted novelty 6.0

    DART is a cross-modal foundation model that delivers rope damage classification, severity regression, and few-shot recognition from a single frozen representation trained on 4270 images across 14 damage classes.

  5. MoP-JEPA: Hard-Assigned Predictor Mixtures for Stochastic JEPA World Models

    cs.AI 2026-07 conditional novelty 5.5

    Hard-assigned JEPA predictor mixtures quantize stochastic successor modes and enable verified planning where single-head and soft predictors fail.

  6. BrainFIBRE: A Foundation Model via Information Decomposition for Brain Microstructure

    cs.CV 2026-07 unverdicted novelty 5.0

    BrainFIBRE presents a foundation model for brain microstructure that applies self-supervised partial information decomposition on NODDI maps to disentangle unique, synergistic, and redundant information and reports st...

  7. Tackling Multimodal Learning Challenges with Mixture-of-Expert: A Survey

    cs.LG 2026-05 accept novelty 5.0

    A literature survey that categorizes how Mixture-of-Experts architectures address multimodal learning challenges and identifies open research gaps.