REVIEW 10 cited by
VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In this work, we introduce Vision-Language Generative Pre-trained Transformer (VL-GPT), a transformer model proficient at concurrently perceiving and generating visual and linguistic data. VL-GPT achieves a unified pre-training approach for both image and text modalities by employing a straightforward auto-regressive objective, thereby enabling the model to process image and text as seamlessly as a language model processes text. To accomplish this, we initially propose a novel image tokenizer-detokenizer framework for visual data, specifically designed to transform raw images into a sequence of continuous embeddings and reconstruct them accordingly. In combination with the existing text tokenizer and detokenizer, this framework allows for the encoding of interleaved image-text data into a multimodal sequence, which can subsequently be fed into the transformer model. Consequently, VL-GPT can perform large-scale pre-training on multimodal corpora utilizing a unified auto-regressive objective (i.e., next-token prediction). Upon completion of pre-training, VL-GPT exhibits remarkable zero-shot and few-shot performance across a diverse range of vision and language understanding and generation tasks, including image captioning, visual question answering, text-to-image generation, and more. Additionally, the pre-trained model retrains in-context learning capabilities when provided with multimodal prompts. We further conduct instruction tuning on our VL-GPT, highlighting its exceptional potential for multimodal assistance. The source code and model weights shall be released.
Forward citations
Cited by 10 Pith papers
-
STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs
Treating image-editing as a shared target between understanding and generation, with sequential reinforcement learning, improves a unified multimodal model's caption-image consistency and several benchmark scores.
-
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs
DisCo assigns each visual token to a unique concept from the caption and aligns its attention across frames, improving video MLLM accuracy and token efficiency.
-
Olympus: A Universal Task Router for Computer Vision Tasks
Olympus is a trained MLLM router that delegates 20 vision tasks to specialist models and supports chain-of-action execution of up to five tasks per instruction.
-
Exploring Multi-Grained Concept Annotations for Multimodal Large Language Models
Multi-grained concept annotations, when integrated into interleaved image-text documents, improve MLLM comprehension and generation beyond caption-only pretraining at much smaller data scale.
-
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation
Diffusion-conditioned denoising trains a video tokenizer whose features support both video question answering and text-to-video generation in a single LLM.
-
EventGPT: Event Stream Understanding with Multimodal Large Language Models
EventGPT adapts a LLaVA-style MLLM to event camera streams via three-stage training (image-language, event-language, instruction tuning) and outperforms RGB-based MLLMs on its own benchmark.
-
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
PIIP is a multi-branch, parameter-inverted image pyramid that uses smaller pretrained networks for high-resolution inputs and larger networks for low-resolution inputs, improving efficiency across perception and multi...
-
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
SynerGen-VL introduces token folding and vision expert FFNs to train a 2.4B encoder-free MLLM that matches larger unified models like Emu3 on multiple image understanding and generation benchmarks.
-
Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation
A unified image understanding and generation model with decoupled visual encoders achieves competitive benchmark scores on both tasks.
-
Visual Large Language Models for Generalized and Specialized Applications
This paper reviews and taxonomizes VLLM applications into vision-to-text, vision-to-action, and text-to-vision, adding ethics and future-work discussion.
Discussion (0). Continue with ORCID to comment.