Pith. sign in

ARMOR: Empowering Multimodal Understanding Model with Interleaved Multimodal Generation Capability

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Unified multimodal understanding and generation have recently received much attention in the area of vision and language. Existing UniMs are designed to simultaneously learn both multimodal understanding and generation capabilities, demanding substantial computational resources, and often struggle to generate interleaved text-image. We present ARMOR, a resource-efficient and pure autoregressive framework that achieves both understanding and generation by fine-tuning existing multimodal large language models (MLLMs). Specifically, ARMOR extends existing MLLMs from three perspectives: (1) For model architecture, an asymmetric encoder-decoder architecture with a forward-switching mechanism is introduced to unify embedding space integrating textual and visual modalities for enabling natural text-image interleaved generation with minimal computational overhead. (2) For training data, a meticulously curated, high-quality interleaved dataset is collected for fine-tuning MLLMs. (3) For the training algorithm, we propose a ``what or how to generate'' algorithm to empower existing MLLMs with multimodal generation capabilities while preserving their multimodal understanding capabilities, through three progressive training stages based on the collected dataset. Experimental results demonstrate that ARMOR upgrades existing MLLMs to UniMs with promising image generation capabilities, using limited training resources. Our code will be released soon at https://github.com/finyorko/armor.

citation-role summary

background 1

citation-polarity summary

fields

cs.CV 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

IA-T2I: Internet-Augmented Text-to-Image Generation

cs.CV · 2025-05-21 · conditional · novelty 6.0

IA-T2I uses active retrieval, hierarchical image selection, and self-reflection to supply internet reference images to T2I models, improving generation accuracy on uncertain-knowledge prompts.

citing papers explorer

Showing 1 of 1 citing paper.

  • IA-T2I: Internet-Augmented Text-to-Image Generation cs.CV · 2025-05-21 · conditional · none · ref 16 · internal anchor

    IA-T2I uses active retrieval, hierarchical image selection, and self-reflection to supply internet reference images to T2I models, improving generation accuracy on uncertain-knowledge prompts.