Pith. sign in

REVIEW 5 cited by

Supervised Multimodal Bitransformers for Classifying Images and Text

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.02950 v2 pith:JFK4VRHC submitted 2019-09-06 cs.CL cs.CVcs.LGstat.ML

Supervised Multimodal Bitransformers for Classifying Images and Text

classification cs.CL cs.CVcs.LGstat.ML
keywords multimodalclassificationimagesinformationperformancesupervisedtaskstext
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Self-supervised bidirectional transformer models such as BERT have led to dramatic improvements in a wide variety of textual classification tasks. The modern digital world is increasingly multimodal, however, and textual information is often accompanied by other modalities such as images. We introduce a supervised multimodal bitransformer model that fuses information from text and image encoders, and obtain state-of-the-art performance on various multimodal classification benchmark tasks, outperforming strong baselines, including on hard test sets specifically designed to measure multimodal performance.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. HuggingFace's Transformers: State-of-the-art Natural Language Processing

    cs.CL 2019-10 accept novelty 6.0

    Hugging Face releases an open-source Python library that supplies a unified API and pretrained weights for major Transformer architectures used in natural language processing.

  2. EVL-MCoT: Enhanced Vision-Language Multi-CoT for Harmful Meme Detection

    cs.CV 2026-07 conditional novelty 5.0

    EVL-MCoT combines multiple hateful and benign chain-of-thought explanations with prototype-guided vision-language fusion, reporting state-of-the-art harmful meme detection accuracy on HatefulMemes and MultiOFF.

  3. LongMoE: Longitudinal Multimodal Learning via Trajectory-Aware Mixture-of-Experts

    cs.LG 2026-06 unverdicted novelty 5.0

    LongMoE is a multimodal framework combining context-aware imputation, frequency-domain attentional tokenization, trajectory encoding, and context-conditioned sparse MoE routing to jointly handle modality missingness a...

  4. Fall into a Pit, Gain in a Wit: Cognitive-Guided Harmful Meme Detection via Misjudgment Risk Pattern Retrieval

    cs.LG 2025-10 unverdicted novelty 5.0

    PatMD improves harmful meme detection by retrieving misjudgment risk patterns to guide MLLMs, reporting 8.30% average F1 and 7.71% accuracy gains on 6,626 memes across 5 tasks.

  5. Connecting online criminal behavior with machine learning: Using authorship attribution to analyze and link potential online traffickers

    cs.CL 2026-04 unverdicted novelty 4.0

    Machine learning can link potential online traffickers by identifying persistent writing and image patterns in anonymous advertisements despite identity changes.