Pith. sign in

REVIEW 7 cited by

Scalable Pre-training of Large Autoregressive Image Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.08541 v1 pith:62AEYVZ3 submitted 2024-01-16 cs.CV

classification cs.CV
keywords modelspre-trainingperformancescaleautoregressivebillionfindingslarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces AIM, a collection of vision models pre-trained with an autoregressive objective. These models are inspired by their textual counterparts, i.e., Large Language Models (LLMs), and exhibit similar scaling properties. Specifically, we highlight two key findings: (1) the performance of the visual features scale with both the model capacity and the quantity of data, (2) the value of the objective function correlates with the performance of the model on downstream tasks. We illustrate the practical implication of these findings by pre-training a 7 billion parameter AIM on 2 billion images, that achieves 84.0% on ImageNet-1k with a frozen trunk. Interestingly, even at this scale, we observe no sign of saturation in performance, suggesting that AIM potentially represents a new frontier for training large-scale vision models. The pre-training of AIM is similar to the pre-training of LLMs, and does not require any image-specific strategy to stabilize the training at scale.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DifFoundMAD: Foundation Models meet Differential Morphing Attack Detection

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    DifFoundMAD improves differential morphing attack detection by replacing traditional embeddings with those from vision foundation models and applying class-balanced lightweight fine-tuning, cutting high-security error...

  2. Hita: Holistic Tokenizer for Autoregressive Image Generation

    cs.CV 2025-07 conditional novelty 7.0 of 10

    Hita's holistic-to-local tokenization lets vanilla autoregressive image models generate global tokens first, improving FID, convergence, and enabling zero-shot style transfer and inpainting.

  3. Visual Pre-Training on Unlabeled Images using Reinforcement Learning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Casting image-crop consistency as temporal-difference value learning improves visual representations on unlabeled web, scene, and video data.

  4. Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types

    cs.SD 2026-07 conditional novelty 4.5 of 10

    On WiWa, factorised multi-task bird species and call-type classification with adaptive loss balancing improves call-type recognition most consistently, while preferred weighting and adaptation depth depend on backbone...

  5. OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning

    cs.CV 2025-09 conditional novelty 4.0 of 10

    OpenVision 2 shows that a caption-only generative objective can match contrastive learning for multimodal vision encoders at lower training cost, scaling to 1B parameters.

  6. SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A three-stage coarse-to-fine training recipe for vision backbones produces consistent benchmark gains for lightweight multimodal LLMs.

  7. Foundation Models for Astrophysics

    astro-ph.IM 2026-08 conditional novelty 3.0 of 10

    Astronomical 'foundation models' largely reuse transformers and self-supervised pretraining, but evidence of transfer to new instruments, populations, or tasks remains rare; the paper argues such evidence, not archite...

Pith tools