REVIEW 14 cited by
Scalable Pre-training of Large Autoregressive Image Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
This paper introduces AIM, a collection of vision models pre-trained with an autoregressive objective. These models are inspired by their textual counterparts, i.e., Large Language Models (LLMs), and exhibit similar scaling properties. Specifically, we highlight two key findings: (1) the performance of the visual features scale with both the model capacity and the quantity of data, (2) the value of the objective function correlates with the performance of the model on downstream tasks. We illustrate the practical implication of these findings by pre-training a 7 billion parameter AIM on 2 billion images, that achieves 84.0% on ImageNet-1k with a frozen trunk. Interestingly, even at this scale, we observe no sign of saturation in performance, suggesting that AIM potentially represents a new frontier for training large-scale vision models. The pre-training of AIM is similar to the pre-training of LLMs, and does not require any image-specific strategy to stabilize the training at scale.
Forward citations
Cited by 14 Pith papers
-
DifFoundMAD: Foundation Models meet Differential Morphing Attack Detection
DifFoundMAD improves differential morphing attack detection by replacing traditional embeddings with those from vision foundation models and applying class-balanced lightweight fine-tuning, cutting high-security error...
-
Hita: Holistic Tokenizer for Autoregressive Image Generation
Hita's holistic-to-local tokenization lets vanilla autoregressive image models generate global tokens first, improving FID, convergence, and enabling zero-shot style transfer and inpainting.
-
Visual Pre-Training on Unlabeled Images using Reinforcement Learning
Casting image-crop consistency as temporal-difference value learning improves visual representations on unlabeled web, scene, and video data.
-
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
TokLIP semanticizes VQ image tokens with a causal CLIP-style encoder, improving multimodal comprehension while preserving autoregressive image generation.
-
Multimodal Autoregressive Pre-training of Large Vision Encoders
AIMV2 pre-trains vision encoders by autoregressively predicting both image patches and text tokens, beating CLIP and SigLIP on many recognition and multimodal benchmarks.
-
Patient-Specific Autoregressive Models for Organ Motion Prediction in Radiotherapy
Auto-RMP, an autoregressive VQGAN plus causal transformer model, predicts future 4D CT phases from prior phases and reports higher lung and heart motion accuracy than DAM and DiffuseRT on public and private datasets.
-
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput
Flash-VL 2B is a 2-billion-parameter vision-language model with higher measured throughput than similar 2B models and slightly better average benchmark scores, thanks to a new image-tiling method.
-
An Empirical Study of Autoregressive Pre-training from Videos
Autoregressive next-token prediction on video and image tokens yields competitive visual representations across recognition, tracking, and robotics benchmarks, with scaling laws that are slower than those of language models.
-
Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types
On WiWa, factorised multi-task bird species and call-type classification with adaptive loss balancing improves call-type recognition most consistently, while preferred weighting and adaptation depth depend on backbone...
-
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
OpenVision 2 shows that a caption-only generative objective can match contrastive learning for multimodal vision encoders at lower training cost, scaling to 1B parameters.
-
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement
A three-stage coarse-to-fine training recipe for vision backbones produces consistent benchmark gains for lightweight multimodal LLMs.
-
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
A comprehensive survey that frames multimodal understanding and generation as next token prediction and proposes a five-part taxonomy.
-
MAL: Cluster-Masked and Multi-Task Pretraining for Enhanced xLSTM Vision Performance
A cluster-masked autoregressive plus multi-task pretraining framework lifts xLSTM vision backbones to 83.4% top-1 accuracy on ImageNet-1K, roughly one point above the ViL-B baseline.
-
Foundation Models for Astrophysics
Astronomical 'foundation models' largely reuse transformers and self-supervised pretraining, but evidence of transfer to new instruments, populations, or tasks remains rare; the paper argues such evidence, not archite...
Discussion (0). Continue with ORCID to comment.