Pith. sign in

REVIEW 1 cited by

MMoE: Enhancing Multimodal Models with Mixtures of Multimodal Interaction Experts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.09580 v3 pith:SKSXPQ6C submitted 2023-11-16 cs.CL

classification cs.CL
keywords multimodalmodelsmmoeinteractionsdetectionexpertsexpressedhumor
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Advances in multimodal models have greatly improved how interactions relevant to various tasks are modeled. Today's multimodal models mainly focus on the correspondence between images and text, using this for tasks like image-text matching. However, this covers only a subset of real-world interactions. Novel interactions, such as sarcasm expressed through opposing spoken words and gestures or humor expressed through utterances and tone of voice, remain challenging. In this paper, we introduce an approach to enhance multimodal models, which we call Multimodal Mixtures of Experts (MMoE). The key idea in MMoE is to train separate expert models for each type of multimodal interaction, such as redundancy present in both modalities, uniqueness in one modality, or synergy that emerges when both modalities are fused. On a sarcasm detection task (MUStARD) and a humor detection task (URFUNNY), we obtain new state-of-the-art results. MMoE is also able to be applied to various types of models to gain improvement.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Discrete Diffusion Models for Language Generation

    cs.CL 2025-07 reject novelty 3.0 of 10

    This thesis reports an empirical D3PM versus autoregressive comparison on WikiText-103, but the claimed speed advantage is unsupported because the speed numbers duplicate NLL values.

Pith tools