Pith. sign in

REVIEW 2 cited by

VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.00522 v2 pith:M757D4R2 submitted 2024-03-01 cs.CV

classification cs.CV
keywords visionllamavisiontasksgenerationimagellamamanyprocess
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models are built on top of a transformer-based architecture to process textual inputs. For example, the LLaMA stands out among many open-source implementations. Can the same transformer be used to process 2D images? In this paper, we answer this question by unveiling a LLaMA-like vision transformer in plain and pyramid forms, termed VisionLLaMA, which is tailored for this purpose. VisionLLaMA is a unified and generic modelling framework for solving most vision tasks. We extensively evaluate its effectiveness using typical pre-training paradigms in a good portion of downstream tasks of image perception and especially image generation. In many cases, VisionLLaMA have exhibited substantial gains over the previous state-of-the-art vision transformers. We believe that VisionLLaMA can serve as a strong new baseline model for vision generation and understanding. Our code is released at https://github.com/Meituan-AutoML/VisionLLaMA.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. PixNerd: Pixel Neural Field Diffusion

    cs.CV 2025-07 conditional novelty 6.0 of 10

    PixNerd is a single-stage pixel-space diffusion transformer that uses predicted neural field weights to decode large patches, reaching 2.15 FID on ImageNet 256 without a VAE.

  2. GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Adding LLM-style GEGLU, RMSNorm, and rotary position embeddings to CoCa's vision encoder reduced contrastive loss, perplexity, and CoCa loss on one pretraining and three fine-tuning datasets, compared with an internal...

Pith tools