Pith. sign in

REVIEW 10 cited by

VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.09251 v1 pith:6Y4H76CN submitted 2023-12-14 cs.CV

classification cs.CV
keywords modelvl-gptimagemultimodaltexttransformerdatageneration
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this work, we introduce Vision-Language Generative Pre-trained Transformer (VL-GPT), a transformer model proficient at concurrently perceiving and generating visual and linguistic data. VL-GPT achieves a unified pre-training approach for both image and text modalities by employing a straightforward auto-regressive objective, thereby enabling the model to process image and text as seamlessly as a language model processes text. To accomplish this, we initially propose a novel image tokenizer-detokenizer framework for visual data, specifically designed to transform raw images into a sequence of continuous embeddings and reconstruct them accordingly. In combination with the existing text tokenizer and detokenizer, this framework allows for the encoding of interleaved image-text data into a multimodal sequence, which can subsequently be fed into the transformer model. Consequently, VL-GPT can perform large-scale pre-training on multimodal corpora utilizing a unified auto-regressive objective (i.e., next-token prediction). Upon completion of pre-training, VL-GPT exhibits remarkable zero-shot and few-shot performance across a diverse range of vision and language understanding and generation tasks, including image captioning, visual question answering, text-to-image generation, and more. Additionally, the pre-trained model retrains in-context learning capabilities when provided with multimodal prompts. We further conduct instruction tuning on our VL-GPT, highlighting its exceptional potential for multimodal assistance. The source code and model weights shall be released.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Treating image-editing as a shared target between understanding and generation, with sequential reinforcement learning, improves a unified multimodal model's caption-image consistency and several benchmark scores.

  2. DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs

    cs.CV 2025-07 conditional novelty 6.0 of 10

    DisCo assigns each visual token to a unique concept from the caption and aligns its attention across frames, improving video MLLM accuracy and token efficiency.

  3. Olympus: A Universal Task Router for Computer Vision Tasks

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Olympus is a trained MLLM router that delegates 20 vision tasks to specialist models and supports chain-of-action execution of up to five tasks per instruction.

  4. Exploring Multi-Grained Concept Annotations for Multimodal Large Language Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Multi-grained concept annotations, when integrated into interleaved image-text documents, improve MLLM comprehension and generation beyond caption-only pretraining at much smaller data scale.

  5. Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Diffusion-conditioned denoising trains a video tokenizer whose features support both video question answering and text-to-video generation in a single LLM.

  6. EventGPT: Event Stream Understanding with Multimodal Large Language Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    EventGPT adapts a LLaVA-style MLLM to event camera streams via three-stage training (image-language, event-language, instruction tuning) and outperforms RGB-based MLLMs on its own benchmark.

  7. Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding

    cs.CV 2025-01 conditional novelty 5.0 of 10

    PIIP is a multi-branch, parameter-inverted image pyramid that uses smaller pretrained networks for high-resolution inputs and larger networks for low-resolution inputs, improving efficiency across perception and multi...

  8. SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding

    cs.CV 2024-12 conditional novelty 5.0 of 10

    SynerGen-VL introduces token folding and vision expert FFNs to train a 2.4B encoder-free MLLM that matches larger unified models like Emu3 on multiple image understanding and generation benchmarks.

  9. Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A unified image understanding and generation model with decoupled visual encoders achieves competitive benchmark scores on both tasks.

  10. Visual Large Language Models for Generalized and Specialized Applications

    cs.CV 2025-01 conditional novelty 3.0 of 10

    This paper reviews and taxonomizes VLLM applications into vision-to-text, vision-to-action, and text-to-vision, adding ethics and future-work discussion.

Pith tools