Pith. sign in

REVIEW 6 cited by

UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.15409 v4 pith:XXJH3VV4 submitted 2020-12-31 cs.CL

classification cs.CL
keywords single-modaldatamulti-modaltasksunimotextualunderstandingvisual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existed pre-training methods either focus on single-modal tasks or multi-modal tasks, and cannot effectively adapt to each other. They can only utilize single-modal data (i.e. text or image) or limited multi-modal data (i.e. image-text pairs). In this work, we propose a unified-modal pre-training architecture, namely UNIMO, which can effectively adapt to both single-modal and multi-modal understanding and generation tasks. Large scale of free text corpus and image collections can be utilized to improve the capability of visual and textual understanding, and cross-modal contrastive learning (CMCL) is leveraged to align the textual and visual information into a unified semantic space over a corpus of image-text pairs. As the non-paired single-modal data is very rich, our model can utilize much larger scale of data to learn more generalizable representations. Moreover, the textual knowledge and visual knowledge can enhance each other in the unified semantic space. The experimental results show that UNIMO significantly improves the performance of several single-modal and multi-modal downstream tasks. Our code and pre-trained models are public at the UNIMO project page https://unimo-ptm.github.io/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. scBIT: Integrating Single-cell Transcriptomic Data into fMRI-based Prediction for Alzheimer's Disease Diagnosis

    q-bio.QM 2025-02 reject novelty 6.0 of 10

    A new cross-modal deep learning model, scBIT, pairs single-cell gene expression with fMRI to improve Alzheimer's diagnostic accuracy, though the reported gains may be inflated by target-label leakage.

  2. Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A vision-language pretraining method generates token-level negative image samples from a visual dictionary and combines them with textual negatives to improve fine-grained understanding.

  3. Describe Anything Model for Visual Question Answering on Text-rich Images

    cs.CV 2025-07 conditional novelty 4.0 of 10

    DAM-QA aggregates answers from full-image and sliding-window views of the Describe Anything Model with a weighted vote, improving text-rich VQA on some benchmarks but not all.

  4. From Screens to Scenes: A Survey of Embodied AI in Healthcare

    cs.AI 2025-01 conditional novelty 4.0 of 10

    A survey of embodied AI in healthcare, organizing 35 tasks into four application domains and proposing a five-level intelligence scale.

  5. Visual question answering: from early developments to recent advances -- a survey

    cs.CV 2025-01 conditional novelty 2.0 of 10

    A survey that classifies VQA architectures by encoder, fusion, and decoder, reviews datasets and metrics, and discusses applications and future directions.

  6. Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey

    cs.CL 2024-11 conditional novelty 1.0 of 10

    A survey that organizes VQA methods from feature extraction through MLLM reasoning, datasets, and metrics, without introducing new experimental results.

Pith tools